Part IV · Chapter 17
Performance, Polling, and Scaling
A Modbus system can be perfectly correct and still too slow. Performance isn't about clever tricks; it's about knowing where the time goes and spending it wisely.
By the end of this chapter you can
- Budget a poll cycle from its parts: transmit, turnaround, response, and serial gaps.
- Explain why baud rate sets a hard ceiling, and what actually lifts it.
- Speed up a loop with batching and tiered polling.
- Keep one dead device from stalling everything, and know when to scale out.
Where the time goes
Every transaction has parts: the client transmits the request, the device processes it and turns its line around, the device transmits the response, and on serial there are the mandatory silent gaps that frame each message. A single read takes anywhere from a few milliseconds on Ethernet to tens of milliseconds on a modest serial bus. Multiply by the transactions in your loop and you have the cycle time: how long it takes to refresh your whole picture once.
The crucial insight: the number of transactions usually matters more than the amount of data. Each transaction carries fixed overhead (addressing, framing, turnaround, gaps) that you pay whether you read one register or a hundred. In the read above, the four registers' actual data is 8 bytes; everything else is overhead. Forty single-register reads pay that overhead forty times; one read of forty contiguous registers pays it once. To go faster, send fewer, larger requests.
The serial ceiling
Serial Modbus has a hard speed limit set by the baud rate. Each byte travels as a character of 10 or 11 bits (start bit, 8 data bits, parity or a second stop bit, stop bit). So at 9600 baud with 8N1 (10-bit characters) the bus carries 960 bytes per second; with 8E1 or 8N2 (11 bits), about 873.
A simple read (request, response, and two gaps) is a few dozen character times plus the device's own delay. With 10–20 ms of device turnaround that works out to very roughly 20 to 30 transactions per second at 9600 baud, and often fewer. No amount of clever code makes a 9600-baud bus carry more than 9600 baud.
There are only two ways to live within that ceiling: do less work (batching and tiering, next) or do it faster. Faster means raising the baud rate (19200 or 38400 doubles or quadruples the byte ceiling, if every device on the bus supports it) or, where the system allows, moving to Modbus TCP, which isn't bound by baud rate and can pipeline many requests at once. Notice in the figure that transactions per second grow less than the byte rate: above 19200 baud the inter-frame gap is fixed at 1.75 ms, and the device's turnaround doesn't shrink at all.
Batching: read in blocks
The highest-value optimization is batching: reading a contiguous span of registers in one request instead of fetching them one by one. Need registers 0 through 19? Send one request for twenty and slice the result. You pay the per-transaction overhead once instead of twenty times, often cutting cycle time by an order of magnitude. The limits: 125 registers per read, and the registers must be contiguous.
| Approach | Transactions | Overhead paid | At 9600 8N1, 10 ms turnaround |
|---|---|---|---|
| 20 single-register reads | 20 | 20× | 20 × 32.9 ≈ 658 ms |
| 1 read of 20 registers | 1 | 1× | ≈ 72 ms |
Batching needs the data to sit in a contiguous block, which is partly why thoughtful device makers group related registers together. When the values you need are scattered with gaps, it's a judgment call: one big read that also pulls registers you don't need, or several smaller reads. Usually the single read with some waste still wins, because overhead dominates; but a very large gap can tip it the other way. Measure when it matters.
rr = client.read_holding_registers(0, count=20) # registers 0-19, one transaction
if not rr.isError():
level, setpoint = rr.registers[0], rr.registers[1]
flow_regs = rr.registers[4:6] # a two-register float
Tiered polling
Not all data needs the same freshness. A pump's pressure may need reading several times a second; its run-hours counter changes meaningfully once a minute; its firmware version never changes. Tiered polling gives each point a rate that matches its need: a fast tier every cycle, a medium tier every few cycles, a slow tier rarely.
Tiering pairs naturally with batching: within each tier, group the contiguous reads. Read the fast block every cycle, the medium block occasionally, the slow block seldom, each as one batched request. It takes a little planning, but it's the difference between a network that comfortably serves a hundred points and one that chokes on twenty.
The slow-device problem
On a serial bus transactions happen one at a time, so one slow or unresponsive device stalls the entire loop. If a device has gone offline and your timeout is two seconds, every cycle wastes two full seconds on it; with three attempts (two retries), six seconds. One dead node can cripple the refresh rate of a dozen healthy ones. It's the most common reason a system that "worked yesterday" feels broken today.
The cure has two parts. First, tune the timeout and retries to your bus: long enough that a healthy-but-slow device isn't falsely declared dead, and no longer, because every extra second is paid on every failed poll. Second, back off from failing devices: when one misses several polls in a row, stop polling it every cycle and check it only occasionally, so a dead node costs one slow probe now and then instead of a full timeout every cycle. When it answers again, return it to the normal rotation.
Poll-cycle calculator (Modbus RTU, FC03 reads)
Model: each read is FC03 (request 8 bytes, response 5 + 2N bytes), each frame followed by a 3.5-character gap (fixed 1.75 ms above 19200 baud). A dead device costs one request plus a full timeout per attempt; once a device fails, the client skips its remaining reads that cycle. Backoff probes a dead device once (no retries) every 10th cycle; the cycle time shown is the average.
Scaling out
When one bus is genuinely full even after batching and tiering, scale out rather than up.
- On serial: split devices across several buses, each with its own adapter and its own polling loop, so the loops run in parallel and each bus carries less traffic.
- On TCP: open several connections and let them work concurrently, using the pipelining that the transaction identifier from Chapter 7 makes possible.
The principle is the same in both worlds: parallelism buys throughput that a single serial conversation, bound to its one-at-a-time discipline, simply can't offer. Four buses of ten devices refresh four times faster than one bus of forty.
Check your understanding
1. You need holding registers 0–39 from one device. What's fastest?
Each transaction pays fixed overhead. One contiguous read of 40 (under the 125 limit) pays it once.
2. Roughly how many bytes per second does a 9600-baud bus carry with 8N1 framing?
Each byte travels as a 10-bit character (start, 8 data, stop), so 9600 ÷ 10 = 960 bytes/s. With 11-bit characters (8E1, 8N2) it's about 873.
3. One of twelve devices goes offline; timeout 2 s, three attempts. What happens to the cycle?
Serial is one transaction at a time, so the client waits out every attempt. Backoff is the cure.
4. A firmware-version register never changes. How should it be polled?
Tiered polling gives each point the freshness it needs and no more, freeing the bus for fast, control-critical values.