Part IV · Chapter 16
Troubleshooting Modbus
Modbus is one of the most diagnosable protocols you will ever meet, because it almost always tells you which kind of problem you have. The trick is to follow its signposts instead of poking at random.
By the end of this chapter you can
- Turn "it doesn't work" into a named cause with a two-level method: which outcome, then which cause.
- Pick the right tool: a poll tool, a protocol analyzer, a multimeter, or the device's own counters.
- Recognize the common fault signatures, and run an ordered checklist when the bus is silent.
- Use isolate, simplify, expand when the checklist runs out.
The method: start from the reply
Every diagnosis begins with the three-outcome question from Chapter 3: when you send a request, what comes back? The answer drops you into one of three worlds.
- Data means the path works end to end. If the data looks wrong, the problem is interpretation, not communication.
- An exception means the device is alive and listening but refuses this request. The problem is in what you asked.
- Silence means the request and reply aren't completing at all. The problem is in the connection.
Almost every Modbus fault lives in exactly one of these three. Naming which one, first, is the most valuable troubleshooting move you can make.
From each branch, a second step narrows further. Data that looks wrong: suspect decoding, such as a missed signed conversion, the wrong scale factor, or, most often, the word-order trap from Chapter 14. An exception: read its code. 01 means the device lacks that function, 02 means your address is wrong, 03 means your quantity or value broke a rule. Silence: work the connection checklist below. This two-level method (which world, then which cause) resolves most problems without any special equipment.
The toolbox
Some faults need you to see what's happening. You don't need every tool, but knowing they exist saves hours:
- A poll tool or scanner sends individual requests so you can test one address at a time and read the exact reply. Your own pymodbus client counts; so do Modbus Poll, ModScan, QModMaster, and the command-line
mbpoll. - A protocol analyzer. For Modbus TCP, Wireshark decodes frames off the network and lays out the MBAP header and PDU. For serial, a line monitor or serial-port sniffer (or a second USB-to-RS-485 adapter just listening) shows the actual bytes on the bus.
- A multimeter for the physical layer: continuity, voltages, and whether the A and B lines are reversed or shorted on an RS-485 run.
- The diagnostic counters from Chapter 9, read with function
0x08, which let a serial device report its own CRC-error and message tallies.
| Tool | Answers the question | Best for |
|---|---|---|
| Poll tool | "What does this one request get back?" | Simplifying to one known read |
| Protocol analyzer | "Did my request leave? Did anything answer? What exactly?" | Turning guesswork into fact |
| Multimeter | "Is the wire itself sound?" | New installs, reversed or shorted pairs |
Diagnostic counters (0x08) | "What has the device itself seen?" | Intermittent faults, noise |
Of these, the protocol analyzer is the one that turns guesswork into certainty. Watching the real frames, you see at once whether your request even left the machine, whether the device answered, and exactly what it said.
For Modbus TCP the analyzer is almost free: Wireshark understands Modbus out of the box. Two display filters do most of the work:
mbtcp # only Modbus TCP frames (decoded on port 502)
tcp.port == 502 # everything on the Modbus port, incl. connection setup
Wireshark decodes Modbus on port 502 by default. For the practice server on 5020, use Decode As… and pick Modbus/TCP for that port, or just filter tcp.port == 5020 and read the raw bytes. On Windows, capture localhost traffic with the Npcap loopback adapter.
For serial, a line monitor takes a little more setup, but one that prints each frame in hex is worth its weight when a bus misbehaves. And your own client makes a fine poll tool: one request, the exact reply, nothing else in the way.
from pymodbus.client import ModbusTcpClient
client = ModbusTcpClient("127.0.0.1", port=5020, timeout=2)
if not client.connect():
raise SystemExit("no connection: check IP, port, firewall")
rr = client.read_input_registers(0, count=1) # one known register
print(rr) # data, or the exception with its code
client.close()
Fault signatures
Experience compresses into recognizable signatures: a symptom that points almost always at one cause. These are worth memorizing:
| Symptom | Most likely cause |
|---|---|
| Exception 02 on a clean request | Address off by one, or wrong table |
| Exception 03 on a large read | Quantity over the limit (125 registers) |
| A 32-bit value is wildly wrong | Wrong word order (Chapter 14) |
| A negative reading shows as ~65535 | Signed value read as unsigned |
| Total silence on a new serial bus | Baud/parity mismatch, or A/B reversed |
| Silence from one device only | Wrong unit address, or that node is down |
| Intermittent CRC errors | Noise, missing termination, bad connection |
| TCP connection refused or times out | Wrong IP/port, or a firewall on 502 |
The last row splits usefully. Refused comes back instantly: the host is up but nothing is listening on that port (wrong port, server not running), or a firewall is actively rejecting. A connect that times out usually means a wrong IP, a host that's down, or a firewall silently dropping the packets.
Diagnose it
A checklist for silence
Silence is the hardest outcome because, by definition, the protocol tells you nothing. So it gets an explicit checklist, run in order, from the most common and cheapest checks to the rarest.
Serial bus
- Port name. Is it really
/dev/ttyUSB0(orCOM3)?ls /dev/ttyUSB*ordmesg | grep ttyUSBon Linux. - Permission to open it. "Permission denied" almost always means the user isn't in the
dialoutgroup, not a wiring fault. - Serial settings match the device exactly, baud rate and parity first, then stop bits.
- Unit address names a device that is actually present (and no two devices share it).
- Wiring. Try an A/B swap; check termination at the two far ends only, and the common/ground.
TCP connection
- IP address and reachability. Can you ping it? No ping points to cabling, subnet, VLAN, or route.
- Port, often 502 (5020 for this course's practice server).
- Firewall. Nothing blocking the port in between. In PowerShell,
Test-NetConnection 192.168.1.50 -Port 502tests reachability and the port in one line; look forTcpTestSucceeded : True. - Unit identifier, if a gateway is involved: it must name the serial device behind it.
Test-NetConnection 192.168.1.50 -Port 502
Run them in sequence rather than jumping around. Most silent buses fail on one of the first two or three items: a wrong baud rate, a missing dialout membership, a reversed pair. Finding it early saves you from chasing exotic causes that almost never apply.
The isolate-simplify-expand workflow
When the checklist doesn't crack it, fall back on a workflow that works for nearly any communication problem:
- Isolate by removing variables: connect to just one device, run a loopback test with the Diagnostics function, or point a known-good scanner at the device to rule your own code in or out.
- Simplify to the smallest possible success: read a single register whose correct value you already know, rather than running your full polling loop.
- Expand only once that works: add registers, devices, and your application logic back one layer at a time, watching for the step that breaks.
This defeats the most common troubleshooting mistake: changing several things at once and losing track of which change mattered. Growing outward by single steps, you always know which addition broke things, because it's the one you just made. It feels slower; in practice it's far faster than poking at a complex setup hoping something shifts.
Check your understanding
1. A request returns data, but a flow rate reads as 4.6 × 1018. Which world are you in?
Data came back, so communication works end to end. A wildly wrong 32-bit value is the word-order signature.
2. A brand-new serial bus gives total silence from every device. What do you check first?
Silence gets the ordered checklist, cheapest and most common first. A firewall doesn't apply to a serial bus.
3. Which tool turns "did my request even go out?" into a fact?
Watching the real frames shows whether the request left, whether anything answered, and exactly what it said.
4. Your full polling loop fails mysteriously. What's the isolate-simplify-expand first move?
Get the smallest possible success first, then add layers back one at a time so the failing step reveals itself.