LearnSCADA

Part IV · Chapter 16

Troubleshooting Modbus

Modbus is one of the most diagnosable protocols you will ever meet, because it almost always tells you which kind of problem you have. The trick is to follow its signposts instead of poking at random.

By the end of this chapter you can

  • Turn "it doesn't work" into a named cause with a two-level method: which outcome, then which cause.
  • Pick the right tool: a poll tool, a protocol analyzer, a multimeter, or the device's own counters.
  • Recognize the common fault signatures, and run an ordered checklist when the bus is silent.
  • Use isolate, simplify, expand when the checklist runs out.

The method: start from the reply

Every diagnosis begins with the three-outcome question from Chapter 3: when you send a request, what comes back? The answer drops you into one of three worlds.

  • Data means the path works end to end. If the data looks wrong, the problem is interpretation, not communication.
  • An exception means the device is alive and listening but refuses this request. The problem is in what you asked.
  • Silence means the request and reply aren't completing at all. The problem is in the connection.

Almost every Modbus fault lives in exactly one of these three. Naming which one, first, is the most valuable troubleshooting move you can make.

The master troubleshooting tree. One request at the top branches three ways. Data that looks wrong leads to interpretation: signed values, scale factor, word order, off by one. An exception leads to what you asked: 01 no such function, 02 wrong address, 03 bad quantity or value, 0B device behind a gateway. Silence leads to the connection checklist: serial port, permission, settings, unit address and wiring, or TCP IP address, port, firewall and unit ID. A marker walks down each branch in turn. Send one request. What comes back? Data …that looks wrong Exception read the code Silence work the checklist Interpretation • signed read as unsigned • wrong scale factor • word order (32-bit) • off by one • wrong table What you asked 01 no such function 02 wrong address 03 quantity / value 0B device behind gateway is silent The connection serial: port, permission, settings, unit, wiring TCP: IP, port, firewall, unit ID behind a gateway
Figure 16.1 The master troubleshooting tree. The kind of reply routes you to the right area; the exception code or the symptom narrows it from there.

From each branch, a second step narrows further. Data that looks wrong: suspect decoding, such as a missed signed conversion, the wrong scale factor, or, most often, the word-order trap from Chapter 14. An exception: read its code. 01 means the device lacks that function, 02 means your address is wrong, 03 means your quantity or value broke a rule. Silence: work the connection checklist below. This two-level method (which world, then which cause) resolves most problems without any special equipment.

The toolbox

Some faults need you to see what's happening. You don't need every tool, but knowing they exist saves hours:

  • A poll tool or scanner sends individual requests so you can test one address at a time and read the exact reply. Your own pymodbus client counts; so do Modbus Poll, ModScan, QModMaster, and the command-line mbpoll.
  • A protocol analyzer. For Modbus TCP, Wireshark decodes frames off the network and lays out the MBAP header and PDU. For serial, a line monitor or serial-port sniffer (or a second USB-to-RS-485 adapter just listening) shows the actual bytes on the bus.
  • A multimeter for the physical layer: continuity, voltages, and whether the A and B lines are reversed or shorted on an RS-485 run.
  • The diagnostic counters from Chapter 9, read with function 0x08, which let a serial device report its own CRC-error and message tallies.
ToolAnswers the questionBest for
Poll tool"What does this one request get back?"Simplifying to one known read
Protocol analyzer"Did my request leave? Did anything answer? What exactly?"Turning guesswork into fact
Multimeter"Is the wire itself sound?"New installs, reversed or shorted pairs
Diagnostic counters (0x08)"What has the device itself seen?"Intermittent faults, noise

Of these, the protocol analyzer is the one that turns guesswork into certainty. Watching the real frames, you see at once whether your request even left the machine, whether the device answered, and exactly what it said.

A serial line monitor printing frames in hex, one line at a time. A request to unit 1 draws a normal data reply. A request to unit 2 draws nothing: silence. A request to unit 1 for address 500 draws exception 83 02. Notes beside each line name the outcome. line monitor · /dev/ttyUSB0 · 9600 8E1 TX 01 03 00 00 00 04 44 09 RX 01 03 08 01 F4 00 00 41 48 00 00 75 FE TX 02 03 00 00 00 04 44 3A … nothing for 1000 ms TX 01 03 01 F4 00 04 04 07 RX 01 83 02 C0 F1 data ✓ silence: unit 2? exception 02
Figure 16.3 A line monitor shows the bytes actually on the bus. Seeing the real frames, or seeing that a request draws no answer, replaces speculation with fact.

For Modbus TCP the analyzer is almost free: Wireshark understands Modbus out of the box. Two display filters do most of the work:

WIRESHARK — display filters
mbtcp                 # only Modbus TCP frames (decoded on port 502)
tcp.port == 502       # everything on the Modbus port, incl. connection setup

Wireshark decodes Modbus on port 502 by default. For the practice server on 5020, use Decode As… and pick Modbus/TCP for that port, or just filter tcp.port == 5020 and read the raw bytes. On Windows, capture localhost traffic with the Npcap loopback adapter.

For serial, a line monitor takes a little more setup, but one that prints each frame in hex is worth its weight when a bus misbehaves. And your own client makes a fine poll tool: one request, the exact reply, nothing else in the way.

PYTHON — probe.py (a one-shot poll tool)
from pymodbus.client import ModbusTcpClient

client = ModbusTcpClient("127.0.0.1", port=5020, timeout=2)
if not client.connect():
    raise SystemExit("no connection: check IP, port, firewall")

rr = client.read_input_registers(0, count=1)   # one known register
print(rr)               # data, or the exception with its code
client.close()

Fault signatures

Experience compresses into recognizable signatures: a symptom that points almost always at one cause. These are worth memorizing:

SymptomMost likely cause
Exception 02 on a clean requestAddress off by one, or wrong table
Exception 03 on a large readQuantity over the limit (125 registers)
A 32-bit value is wildly wrongWrong word order (Chapter 14)
A negative reading shows as ~65535Signed value read as unsigned
Total silence on a new serial busBaud/parity mismatch, or A/B reversed
Silence from one device onlyWrong unit address, or that node is down
Intermittent CRC errorsNoise, missing termination, bad connection
TCP connection refused or times outWrong IP/port, or a firewall on 502

The last row splits usefully. Refused comes back instantly: the host is up but nothing is listening on that port (wrong port, server not running), or a firewall is actively rejecting. A connect that times out usually means a wrong IP, a host that's down, or a firewall silently dropping the packets.

Diagnose it

A checklist for silence

Silence is the hardest outcome because, by definition, the protocol tells you nothing. So it gets an explicit checklist, run in order, from the most common and cheapest checks to the rarest.

Serial bus

  1. Port name. Is it really /dev/ttyUSB0 (or COM3)? ls /dev/ttyUSB* or dmesg | grep ttyUSB on Linux.
  2. Permission to open it. "Permission denied" almost always means the user isn't in the dialout group, not a wiring fault.
  3. Serial settings match the device exactly, baud rate and parity first, then stop bits.
  4. Unit address names a device that is actually present (and no two devices share it).
  5. Wiring. Try an A/B swap; check termination at the two far ends only, and the common/ground.

TCP connection

  1. IP address and reachability. Can you ping it? No ping points to cabling, subnet, VLAN, or route.
  2. Port, often 502 (5020 for this course's practice server).
  3. Firewall. Nothing blocking the port in between. In PowerShell, Test-NetConnection 192.168.1.50 -Port 502 tests reachability and the port in one line; look for TcpTestSucceeded : True.
  4. Unit identifier, if a gateway is involved: it must name the serial device behind it.
TERMINAL — PowerShell
Test-NetConnection 192.168.1.50 -Port 502

Run them in sequence rather than jumping around. Most silent buses fail on one of the first two or three items: a wrong baud rate, a missing dialout membership, a reversed pair. Finding it early saves you from chasing exotic causes that almost never apply.

The isolate-simplify-expand workflow

When the checklist doesn't crack it, fall back on a workflow that works for nearly any communication problem:

  1. Isolate by removing variables: connect to just one device, run a loopback test with the Diagnostics function, or point a known-good scanner at the device to rule your own code in or out.
  2. Simplify to the smallest possible success: read a single register whose correct value you already know, rather than running your full polling loop.
  3. Expand only once that works: add registers, devices, and your application logic back one layer at a time, watching for the step that breaks.
Isolate, simplify, expand. In the first panel, a client connected to five devices drops four of them, leaving one. In the second, a single read of one known register succeeds. In the third, layers are added back one at a time: more registers, a second device, then the application logic, which fails and is marked as the step that broke. 1 · Isolate 2 · Simplify 3 · Expand Client one device only read IR 0, count 1 500 ✓ value you already know (level 50.0 %) 1 register ✓ + all 3 registers ✓ + coil and alarm bit ✓ + app logic ✗ the step you just added
Figure 16.5 Isolate, simplify, expand. Strip the problem down to one known-good read, get that working, then add complexity back one layer at a time until the failing step shows itself.

This defeats the most common troubleshooting mistake: changing several things at once and losing track of which change mattered. Growing outward by single steps, you always know which addition broke things, because it's the one you just made. It feels slower; in practice it's far faster than poking at a complex setup hoping something shifts.

Check your understanding

1. A request returns data, but a flow rate reads as 4.6 × 1018. Which world are you in?

Data came back, so communication works end to end. A wildly wrong 32-bit value is the word-order signature.

2. A brand-new serial bus gives total silence from every device. What do you check first?

Silence gets the ordered checklist, cheapest and most common first. A firewall doesn't apply to a serial bus.

3. Which tool turns "did my request even go out?" into a fact?

Watching the real frames shows whether the request left, whether anything answered, and exactly what it said.

4. Your full polling loop fails mysteriously. What's the isolate-simplify-expand first move?

Get the smallest possible success first, then add layers back one at a time so the failing step reveals itself.