Skip to content
Contact us
Technician testing the network cabling between the PLC and the switches of a communications cabinet, with the caption «PLC y SCADA, sin señal» (PLC and SCADA, no signal)
27 May 2026PLC and SCADA

PLC–SCADA communication failures: symptoms, causes and how to diagnose them

PLC and SCADA · Diagnosis

A communication failure between PLC and SCADA is tracked down layer by layer: first the physical layer (cable, connectors, electrical noise), then the network (switches, loops, IP addresses), then the protocol (timeouts and error codes in Modbus, PROFINET or OPC UA) and finally the application (how many variables are requested, how often and against which clock).

This guide is for people who have the problem in front of them. If you need some background first, we explain what a PLC is, the controller that runs the machine, and what a SCADA system is, the software that supervises and records the plant.

Before restarting anything: note what time it fails, which devices lose their data and what the system says (data quality, error code, driver message). Restarting the server or the switch usually brings communication back, and the restart wipes the switch counters and whatever the driver does not save to disk, which is exactly what showed why the connection had been lost.

Symptoms of a communication loss and what they indicate

It rarely starts with a complete outage. The usual signs are small:

  • Frozen variables. The value does not move even though the process does. Many drivers keep the last value read when the connection drops, and if the screen does not show the data quality, the operator sees a believable figure that is no longer real.
  • “Bad” or “Uncertain” quality. In OPC UA every value travels with a status code: Good is a good-quality value, Uncertain one of doubtful quality and Bad one that cannot be used [4]. The specific code usually tells you more than the alarm.
  • Intermittent communication alarms. They appear and clear by themselves. If they coincide with a large motor starting, a shift change or a washdown, the time is your best clue.
  • Gaps in the historical data. The trend joins two points with a straight line or records are missing for a batch, and without an upstream buffer that gap can no longer be filled.
  • Writes that do not arrive. A setpoint or a recipe leaves the SCADA and the PLC does not execute it, or executes it late.
  • Absurd but stable values. A huge or negative number where there should be a temperature. The data arrives, but it is interpreted wrongly: we cover this in the application layer.

Diagnostic table: symptom, probable cause and check

SymptomProbable causeWhat to check
All variables of one device frozen, with no alarmConnection down and a driver that holds the last value; OPC UA subscription closed or not recovered on reconnectionTag quality and connection status in the driver; port link on the switch
“Bad” quality on a group of variablesDevice not responding, wrongly mapped address, gateway that cannot reach the slaveThe code: in Modbus, 02 is an illegal address and 0B a slave behind the gateway that does not respond [1]
Alarm that comes and goes, more so when motors are startingElectrical noise, cable shield wrongly connected, data cables next to power cablesCRC/FCS error counters on the port, and whether they rise when the motors start
A whole segment drops and does not come back until a cable is removedLoop and broadcast stormActivity LED flashing non-stop; a new patch cable that closes a loop; broadcast counters if the switch is managed
Several devices drop for a few seconds and come back by themselvesThe redundancy (STP or ring) reconfigures after a breakTopology changes in the switch log; in PROFINET rings, watchdog compared with the reconfiguration time [2]
One device stops communicating when another is switched onDuplicate IP addressThe MAC address for that IP keeps changing; arping or a capture with two replies
Occasional PROFINET station failuresWatchdog too short for the real network, ring that reconfiguresController diagnostic buffer; update time and accepted cycles without data [2]
Slow screens and timeouts at peak hoursToo many variables, polling faster than necessary, scattered addressesRequests per cycle and response time, in the driver or in a capture
Gaps in the historical data or events out of orderOutage with no buffer before the failed segment; unsynchronised clocksHistorian service log; PLC time compared with server time

Physical layer: cable, connectors and electrical noise

It is the first one to rule out and the one most often skipped. Typical causes: RJ45 connectors fitted on site that work loose with vibration, cable entries into field boxes that are not sealed, moisture that stays inside after washdowns and cable shields left unconnected or connected differently from what the bus manufacturer specifies.

Electrical noise has a recognisable pattern: errors rise when a large motor starts or a drive ramps up. If the data cable shares trunking with the drive outputs, the problem will not be fixed by touching the software.

How to check it: port error counters (frames with a bad CRC or FCS) on the switch or in the PLC’s port diagnostics. If they are rising, use a cable tester and swap the suspect patch cable for a factory-made one. If the errors follow the patch cable, the fault is in the cable.

Green multicore cables running along wire mesh cable trays above the banks of stainless steel pipework in a food plant
Multicore cables on wire mesh trays, above the process pipework of a food plant.

Network layer: switches, loops, duplicate IPs and VLANs

An unmanaged switch in a plant cabinet works, but it gives no per-port counters and cannot copy traffic so that it can be captured.

Loops. A patch cable joining two ports on the same switch, or two switches joined by two paths without a protocol to manage the redundancy, creates a loop in which broadcast frames go round and round endlessly. It tends to happen after an extension or a rushed repair: the whole segment drops at once and does not come back until the cable closing the loop is removed. With redundancy (STP or ring), a cut cable does not bring the segment down, although there may be brief drops while the network reconfigures.

Duplicate IPs. A spare configured with the IP of a device that is still connected, an outside technician’s laptop, a new screen with its factory IP. Two devices take turns communicating: the MAC address for that IP changes from one query to the next, and an arping or a capture shows two devices answering to the same IP.

Segmentation. The NIST guide to operational technology security describes segmenting the plant network into zones, physically with separate switches or logically with VLANs [6]. As well as protecting the network, it stops broadcast traffic from the office reaching the PLCs.

Protocol layer: Modbus, PROFINET and OPC UA

Modbus: a timeout and an exception say different things

The Modbus specification separates two cases that get confused on screen. If the request does not reach the slave, or arrives with a parity or CRC error, there is no response and the client ends up logging a timeout. If it arrives correctly but the device cannot handle it, it returns an exception response: the function code plus 80 hexadecimal and a code giving the reason [1].

In plain terms: a timeout means no response arrived. It points to the cable, the network, the gateway, a device that is switched off or parameters that do not match (slave address, baud rate, parity), because the slave does not answer a frame that is not addressed to it or one that arrives with errors [1]. An exception says that the device received the request and why it is not handling it: 01 and 02 are usually configuration issues, 04 is a device fault and 0A and 0B point to the gateway or whatever is behind it. The most useful codes:

CodeNameWhat it usually means
01Illegal FunctionThe device does not support that function, or is not in a state to handle it [1]
02Illegal Data AddressThe request goes beyond the map: on a device with 100 registers, reading 5 starting at 96 fails because register 100 does not exist [1]
04Server Device FailureUnrecoverable error in the device while executing the request [1]
0AGateway Path UnavailableGateway misconfigured or overloaded [1]
0BGateway Target Device Failed to RespondThe gateway asked and the slave did not answer; normally it is not on the network [1]

Increasing the driver timeout can clear the alarm without removing the cause: the data still arrives late, and on a serial line where the master polls the slaves one at a time, each wait delays the requests queued behind it.

PROFINET: the watchdog time

PROFINET IO links the PLC to its distributed I/O and its drives. The SCADA usually talks to the PLC over S7 or OPC UA, so this fault reaches it indirectly: as a station failure alarm, if the PLC reports it, or as readings that are no longer valid.

The PROFINET watchdog is the time the IO controller or IO device will accept without receiving data. If it is exceeded, the device applies substitute values and the controller logs it as a station failure. In Siemens software it is set as an integer multiple of the update time [2].

The controller’s diagnostic buffer tells you which device dropped and when. In rings, MRP, the redundancy protocol defined in IEC 62439-2, has a typical reconfiguration time of 200 ms and supports up to 50 devices per ring, and Siemens calls for a watchdog of 256 ms or more when several rings are coupled [2]. A watchdog shorter than the reconfiguration time turns a break that the ring should absorb into a stoppage.

OPC UA: subscriptions that get lost

With OPC UA, the SCADA usually subscribes to the variables. The server sends the changes at each publishing interval and, if several cycles in a row pass without changes, a keep-alive message. If too many cycles pass without the client requesting publications, the lifetime counter runs out and the subscription is closed [3].

Subscriptions are designed to survive a loss of the connection and the session [3], but there is a typical scenario: the connection comes back and a group of variables stays frozen, because the outage lasted longer than the subscription lifetime or because the client opened a new session without transferring the subscription or creating it again. The specification provides a retransmission queue for recovering lost notifications with the Republish service [3]; whether the client uses it depends on how it is configured.

Application layer: variables, polling and clocks

Excessive polling. Requesting every variable every second loads the PLC’s communications processor and the driver. A tank temperature does not need the update rate of a weighing scale in the middle of a weighing.

Scattered addresses. A Modbus register read accepts from 1 to 125 contiguous registers [1]. If the variables are spread across the PLC memory, many small requests are needed per cycle; grouped in a contiguous exchange area, far fewer are enough.

Heartbeat. A counter that the PLC increments and the SCADA monitors: if it stops changing, the data is stale even though the connection looks open. It also works the other way round: the PLC monitors a heartbeat from the SCADA and, if it stops, it no longer accepts remote setpoints and switches to a defined safe state.

Clocks. With a different clock in the PLC, the server and the historian, events come out of order. NTP is the standard protocol for synchronising clocks over a network [7], and NIST points out that time synchronisation is needed to correlate events and logs [6].

Data types. Absurd values are usually a signed integer read as unsigned, or a floating-point value that travels in two registers and is reassembled with the word order swapped. You confirm it by comparing the online value in the PLC with the value in the SCADA.

How to diagnose a communication failure between PLC and SCADA, step by step

  1. Narrow down the scope. One device, one line or the whole plant? All the time, on and off, at certain hours? If everything drops at once, look at the network or the server; if one device drops, look at its cable, its port or its configuration.
  2. Read what the system is already telling you. Variable quality, driver error (timeout or exception, and which one), the PLC’s diagnostic buffer and the server events.
  3. Look at the counters. On the managed switch: CRC errors, discarded frames, link drops and broadcasts per port. On the PLC: port and station diagnostics. Reset them to zero, wait for the fault to recur and compare.
  4. Capture the traffic. Wireshark, the free network analyser, decodes Modbus/TCP, PROFINET IO and OPC UA [8]. On a switched network, a laptop plugged into any port only sees its own traffic and broadcast traffic: you need a mirror port (port mirroring or SPAN) on a managed switch, or a network tap. The mirror port must be at least as fast as the port it copies and it does not forward damaged frames, so CRC errors show up in the counters, not in the capture [5].
  5. Isolate it, carefully. Connect a laptop directly to the device with a read-only test client, through a second free port or during a shutdown if it has to be taken off the network. Beforehand, check which interlocks between PLCs run over that network: disconnecting it cuts them. If it then communicates well for hours, the problem is in the network or the load; if it fails in the same way, it is in the device or its configuration.
  6. Change one thing at a time. If you change three at once and the fault disappears, you will not know which one it was.

Solutions to stop it happening again

Finding the cause fixes today’s fault. The next one is prevented through the architecture:

MeasureWhat it preventsWhen it pays off
Separate control network, with its own switches or VLANs [6]Office traffic reaching the PLCs; a loop in an office bringing down the plantIf control and office share switches
Industrial-grade managed switchesFaults that leave no trace: they provide per-port counters and allow captures [5]Where PLCs and servers come together
Ring with MRP and a matching watchdog [2]A cut cable leaving half a segment without communicationLines where a stoppage costs more than the ring
Polling by criticality and contiguous exchange areas [1]Overloading of the PLC and the driver; timeouts at peak hoursSlow screens or timeouts with no physical errors
Timestamped buffer before the segment that drops (in the PLC or in the line-side data collection device) that is flushed on reconnectionGaps in historical data and incomplete batches after an outageWhen those records serve as traceability
NTP on PLCs, servers and historian [7]Events out of order and batches with inconsistent timesAlways
Heartbeat and safe state programmed in the PLCActing on stale data; setpoints arriving at the wrong timeWhere setpoints or recipes depend on it

What changes in an agri-food factory

With frequent washdowns, moisture gets in through connectors and cable glands that are not watertight, and faults appear after cleaning. In feed mills, dust builds up in cabinets and switches that are not properly closed. In a winery, the network is under full load during the grape harvest, when you can least afford to stop: network changes are tested outside the harvest season.

In any food factory, gaps in the historical data leave the batch record incomplete. What the law requires is covered in our article on food traceability.

Stainless steel electrical cabinet with ventilation grilles beside a process line with stainless steel pipework and filters
Line-side stainless steel cabinet in a food plant. With frequent washdowns, the sealing of the cabinet and its cable entries needs watching.

If the cause is the age of the system (PLCs with no spare parts, a SCADA on an unsupported operating system, drivers that nobody can update), the diagnosis ends in a migration plan. How to do it without stopping production is covered in the guide to SCADA and PLC modernisation.

How we work on it at ER Ingeniería

We have been working in electrical installations and automation since 1981, and we have automated 45 agri-food factories, including feed mills, wineries and food plants. When a plant loses communication, we diagnose it layer by layer and on site at the factory. That is what our winery and feed mill automation service covers, from the electrical panel to the software.

For the data we have SuitER, our industrial software: its SuitER Server module connects the PLC network with the management network, collects the data in real time and stores it in a database. For plants that are already running, there is our SatER technical service.

Is your SCADA losing communication with the PLCs?

Tell us which protocol your plant uses and since when you have been noticing the symptoms. With that we will tell you where we would start looking.

Talk to our team

Or call us on 967 140 850

Frequently asked questions about PLC and SCADA communication failures

Why does the SCADA lose communication with the PLC?

The causes fall into four layers: physical (cable, connectors, noise from drives and motors), network (unmanaged switches, loops, duplicate IPs), protocol (timeouts, Modbus exception codes, PROFINET watchdog, OPC UA subscriptions) and application (too many variables, polling too fast, unsynchronised clocks). They are checked in that order.

What is the difference between a timeout and an exception in Modbus?

A timeout means no response arrived: the request was lost, arrived damaged or did not match the device (different slave address, baud rate or parity). An exception is a response: the device received the request and says why it is not handling it. 02, for example, is a register address that does not exist in its map; 04, a fault in the device itself; and 0B, a slave behind a gateway that does not respond.

What does it mean when a variable has Bad or Uncertain quality in the SCADA?

It is the quality status that accompanies the value. In OPC UA, Good is a good-quality value, Uncertain one of doubtful quality and Bad one that cannot be used. The specific code that accompanies Bad usually indicates the cause better than the generic communication alarm.

What is the watchdog in PROFINET?

It is the time the IO controller or IO device will accept without receiving data. If it is exceeded, the device applies substitute values on its outputs and the controller logs it as a station failure. In Siemens software it is set as a multiple of the update time.

Is there any point in increasing the driver timeout?

It can clear the alarm without removing the cause. The data arrives later, and on a serial line each wait delays the requests queued behind it. It is best increased only when you know why it is needed.

How do you capture traffic between the PLC and the SCADA with Wireshark?

With a managed switch that has a mirror port (port mirroring or SPAN) or with a network tap: a laptop plugged into any port of a switch only sees its own traffic and broadcast traffic. Frames with a CRC error do not appear in a mirror-port capture, so those errors are checked in the switch counters.

Why are there gaps in the SCADA historical data?

The usual reason is that the data was lost on the segment that went down and nobody stored it before that segment. It is prevented with a timestamped buffer before that segment, in the PLC or in the device that collects the data, which is flushed on reconnection; time offsets are prevented by synchronising all devices with NTP.

Sources

  1. Modbus Organization. (2012). MODBUS Application Protocol Specification V1.1b3 (section 7 and function 03). modbus.org (PDF)
  2. Siemens AG. (2022). SIMATIC PROFINET with STEP 7: Function Manual (11/2022, A5E03444486-AM; sections 3.1 and 6.4). cache.industry.siemens.com (PDF)
  3. OPC Foundation. (n.d.). OPC 10000-4: OPC Unified Architecture Part 4: Services (v1.05.07, 5.14.1.1). reference.opcfoundation.org
  4. OPC Foundation. (n.d.). OPC Unified Architecture Part 4: Services (v1.04, 7.34.1 StatusCode). reference.opcfoundation.org
  5. Wireshark Foundation. (n.d.). CaptureSetup/Ethernet. wiki.wireshark.org
  6. Stouffer, K. et al. (2023). Guide to Operational Technology (OT) Security (NIST SP 800-82r3; sections 6.2.1.3 and 6.2.12). csrc.nist.gov
  7. Mills, D., Martin, J., Burbank, J. and Kasch, W. (2010). RFC 5905: Network Time Protocol Version 4: Protocol and Algorithms Specification. IETF. rfc-editor.org
  8. Wireshark Foundation. (n.d.). Display Filter Reference: Modbus/TCP (and those for PROFINET IO and OpcUa Binary Protocol). wireshark.org
Call us 967 140 850 Request a quote