Your Plant Floor: Can Diagnose Issues in Seconds, But Fix in Days

Edge devices can spot plant problems fast but rarely fix them. Keep MQTT as the device interface and add a NATS backbone for local storage, two-way commands, and autonomous recovery.


Industry Article October 05, 2026 by Synadia

At first glance, everything in the plant looks fine. No alarms, no buzzers, no LEDs on control panels or text notifications indicating any sort of problem. Each individual system is working the way it was designed. All is quiet and humming along smoothly.

Over the next month, however, scrap rates climb. The quality assurance department raises rework requests and excursion flags. Technicians and process engineers scratch their heads, while the usable throughput drops.

Sound familiar?

The latest edge devices excel at collecting data, and there is plenty being collected. The plant can see everything but does nothing about it. Simply collecting data is not enough to keep the plant efficient. Occasionally dropped communication links, gaps in archival data, and other such clues get missed. Each process looks functional while falling apart behind the scenes.

The solution lies in examining the control architecture itself and assigning roles based on what each component is designed for, rather than forcing it into a role where it will fail silently. One of the best ways to do this is to use MQTT as it was designed. Then, choose a robust backbone that keeps auditable trails, sends only necessary data to the cloud, keeps the rest onsite, and moves commands to the hardware, all while surviving connection outages.

The State of the Industry

Message Queuing Telemetry Transport (MQTT) is a protocol designed to move data to and from devices. Transport is in the name for a reason; it transports data. MQTT is a publish-and-subscribe system, where a broker receives all published data points from sensors, filters them, and then sends them to devices that have subscribed to particular data. MQTT is designed to work with unreliable networks, and is also low power and low bandwidth, making it great for battery-powered sensors.

Synadia recently sponsored a survey of 500 edge practitioners to take the pulse on the current state of industry. In this survey, Synadia defined four operating tiers: observer, responder, predictor, and autonomous. The observer tier says a problem can be detected, but a human must fix the problem. The responder tier diagnoses the problem, but a human must fix the problem manually. The predictor tier anticipates a problem and notifies a human, who must still fix the problem manually. In the autonomous tier, the system predicts, detects, identifies, and fixes the problem itself. Across these four tiers are five dimensions: security, act, recover, correct, and foundation.

The results of the Synadia survey, conducted by Kelsey White Research / Rep Data in March 2026, are shown in Fig. 01. Only 10% of respondents were scored in the autonomous tier. This means edge devices are detecting problems, but fixing very few of them.

The real problem is shown in Figure 1: it is easy to detect problems, but much harder to act upon them.

Figure 1. Synadia's classification of survey data across four tiers and five dimensions.

This survey also found that 86% of edge practitioners say that decision-making intelligence will continue to push toward the edge. Another question showed that only 20% of respondents say that their systems have reached closed-loop remediation. Furthermore, only 13% of teams that experience connectivity challenges say their systems can recover automatically when connection is lost.

Recognizing the Role of MQTT

The problem is not that MQTT is wrong; it is that it was designed as an interface, not as the backbone of the automation system. It is not designed to archive historical data, play back old data, or perform any sort of data analysis. Due to network disconnections, sessions are tied to a single broker. If that broker is unreachable, there is no fallback. Historical data is not reliable for audit trails or troubleshooting. The specifications for MQTT do not include storing data during a communication link drop.

Instead of treating connection disruptions as an anomaly, the architecture should be more robust. Dropped links are a fact of life, and the system should be able to locally store, buffer, and manage data during these events.

Figure 2 shows Site 01 operation is ideal; the buffer is draining and communication is good. However, communication links drop as shown in Site 02. Site 02 is operating as well: data is buffering and nothing is lost, even though communication has dropped. Designing for these inevitable events is essential to keep a complete, accurate record of data for troubleshooting and auditing purposes.

Figure 2. Two sites, one with a good connection and one without, at this snapshot in time.

Instead of using MQTT as the communications backbone, a server runs at every site, regardless of whether the wide area link is connected or not. Data is collected and stored locally, rather than every piece of data being sent to the cloud. The devices themselves still use MQTT to communicate, but the architecture above them changes.

Two-way communication along links is also necessary to push edge devices toward fully autonomous operation. Publish-and-subscribe models, like those used in MQTT, are great at sending data, but lack acknowledgment. MQTT's Quality of Service (QoS) levels only guarantee “one hop”, and do not guarantee the whole path will be completed. In other words, a command may be sent, but MQTT does not provide a way to see if the action was actually taken. Problems cannot be fixed unless commands, authorizations, acknowledgments, and other policies can be sent back to the edge devices.

Integrating MQTT and a More Robust NATS Backbone

MQTT by itself is limited, serving only as a device interface. The solution is not to abandon MQTT, but to use it in conjunction with other systems that can more effectively use the data that is being collected.

Synadia has identified five key architectural features that help drive true autonomous operation at the edges. Based on the survey data, Synadia listed: secure by identity, bidirectional by design, resilient by default, self-recovering/reconciling, and “one fabric” or not another point tool as important factors in a more robust backbone.

From there, secure by identity means each device has its own identity that does not rely on shared credentials. Bidirectional by design means that one reliable path can handle telemetry, commands, policy, and acknowledgment. Resilient by default means that the system can use a store-and-forward method, even during communication outages. Self-recovery and self-reconciliation mean that a system can course correct after an interruption without needing a human to manually recover the device. The last feature, “one fabric”, not another point tool means that all devices and services share a mechanism to communicate between edge and cloud.

In Figure 3, each of these architectural features was scored out of 20 possible points. This chart represents the median score for each feature.

Figure 3. While security and resilience were higher, bidirectionality and self-correction scores were lower.

NATS, by Synadia, is an open-source platform that delivers all of these architectural features. Device-level communication is still performed using MQTT. Above the device level, NATS leaf nodes are servers that work locally, regardless of link connections. Communication links are bidirectional, so commands can travel back down to devices. Above the NATS leaf nodes, there is a uniform fabric that binds all of these systems together and allows for communication with the cloud, sending only the proper data there. Figure 4 highlights the NATS architecture.

Figure 4. The NATS architecture, with MQTT at the device level and leaf nodes operating locally beneath a uniform fabric.

With over 450 million downloads, NATS has demonstrated its capability at numerous facilities. Synadia also created the Synadia Platform, which is a commercial offering that pairs with NATS to help teams automate edge devices at scale.

One such example is MachineMetrics, which implemented NATS.io and found increased flexibility and reduced operational overhead as a result. Their facility had many machines that sampled data in the kHz range with often unreliable data connections. NATS was used to sort and send only relevant data to the cloud, while allowing playback of historical data as needed. Data is processed locally, so that problems can be handled immediately.

The Bottom Line

So, is the solution to abandon MQTT? Absolutely not. It is very good at its role in automation. The solution is to recognize its role and use it for that, instead of trying to get it to do everything else. A communications and data management backbone that is robust enough to support occasional dropped links is essential for troubleshooting and data auditing purposes. Design for drops, rather than letting MQTT alone miss them.

Check out the Edge Autonomy Gap Report to see what edge professionals are saying about autonomous operations in their facilities. From there, implementing NATS is the next step, and Synadia is the first logical place to look. The Synadia Platform allows users to secure, operate, and scale to meet the needs of virtually any business that is interested in edge automation.