LearnSCADA

Part IV · Chapter 17

Performance, Polling, and Scaling

A Modbus system can be perfectly correct and still too slow. Performance isn't about clever tricks; it's about knowing where the time goes and spending it wisely.

By the end of this chapter you can

  • Budget a poll cycle from its parts: transmit, turnaround, response, and serial gaps.
  • Explain why baud rate sets a hard ceiling, and what actually lifts it.
  • Speed up a loop with batching and tiered polling.
  • Keep one dead device from stalling everything, and know when to scale out.

Where the time goes

Every transaction has parts: the client transmits the request, the device processes it and turns its line around, the device transmits the response, and on serial there are the mandatory silent gaps that frame each message. A single read takes anywhere from a few milliseconds on Ethernet to tens of milliseconds on a modest serial bus. Multiply by the transactions in your loop and you have the cycle time: how long it takes to refresh your whole picture once.

One FC03 read of 4 registers at 9600 baud 8N1, drawn to scale as a bar that fills left to right: request 8 bytes, 8.3 milliseconds; a 3.5-character gap, 3.6 ms; device turnaround, 10 ms; response 13 bytes, 13.5 ms; another gap, 3.6 ms. Total about 39 milliseconds. ONE FC03 READ · 4 REGISTERS · 9600 BAUD 8N1 request turnaround response gapgap 8 B · 8.3 ms 3.6 ms 10 ms 13 B · 13.5 ms 3.6 ms ≈ 39 ms per transaction × 50 transactions = a 2-second cycle
Figure 17.1 Where a transaction's time goes: request, the device's turnaround, response, and on serial the silent gaps that frame each message. The sum, times the number of transactions, is your cycle time.

The crucial insight: the number of transactions usually matters more than the amount of data. Each transaction carries fixed overhead (addressing, framing, turnaround, gaps) that you pay whether you read one register or a hundred. In the read above, the four registers' actual data is 8 bytes; everything else is overhead. Forty single-register reads pay that overhead forty times; one read of forty contiguous registers pays it once. To go faster, send fewer, larger requests.

The serial ceiling

Serial Modbus has a hard speed limit set by the baud rate. Each byte travels as a character of 10 or 11 bits (start bit, 8 data bits, parity or a second stop bit, stop bit). So at 9600 baud with 8N1 (10-bit characters) the bus carries 960 bytes per second; with 8E1 or 8N2 (11 bits), about 873.

A simple read (request, response, and two gaps) is a few dozen character times plus the device's own delay. With 10–20 ms of device turnaround that works out to very roughly 20 to 30 transactions per second at 9600 baud, and often fewer. No amount of clever code makes a 9600-baud bus carry more than 9600 baud.

Bars grow to show the byte ceiling at four baud rates with 8N1: 9600 baud carries 960 bytes per second, about 44 single-register reads per second at most; 19200 carries 1920, about 87 reads; 38400 carries 3840, about 135 reads; 115200 carries 11520, about 208 reads. Reads per second assume no device delay. BYTES PER SECOND (8N1) MAX READS/S* 9600 19200 38400 115200 960≈ 44 1 920≈ 87 3 840≈ 135 11 520≈ 208 * one-register FC03 read (8 + 7 bytes) plus two gaps, zero device delay. Real devices add turnaround.
Figure 17.2 Baud rate sets a hard ceiling. At 9600 baud the bus moves about 960 bytes per second; raising the baud rate is the only way to lift the ceiling itself.

There are only two ways to live within that ceiling: do less work (batching and tiering, next) or do it faster. Faster means raising the baud rate (19200 or 38400 doubles or quadruples the byte ceiling, if every device on the bus supports it) or, where the system allows, moving to Modbus TCP, which isn't bound by baud rate and can pipeline many requests at once. Notice in the figure that transactions per second grow less than the byte rate: above 19200 baud the inter-frame gap is fixed at 1.75 ms, and the device's turnaround doesn't shrink at all.

Batching: read in blocks

The highest-value optimization is batching: reading a contiguous span of registers in one request instead of fetching them one by one. Need registers 0 through 19? Send one request for twenty and slice the result. You pay the per-transaction overhead once instead of twenty times, often cutting cycle time by an order of magnitude. The limits: 125 registers per read, and the registers must be contiguous.

ApproachTransactionsOverhead paidAt 9600 8N1, 10 ms turnaround
20 single-register reads2020×20 × 32.9 ≈ 658 ms
1 read of 20 registers11×≈ 72 ms
Two timelines drawn to the same scale fill left to right. The top lane is twenty single-register reads, each mostly overhead with a sliver of data, finishing at about 658 milliseconds. The bottom lane is one read of twenty registers, finishing at about 72 milliseconds, roughly nine times sooner. 20 single-register reads 1 read of 20 registers done · 72 ms ✓ done · 658 ms overhead: framing, gaps, turnaround register data
Figure 17.3 Batching. Twenty small reads each pay the full overhead; one read of the same twenty contiguous registers pays it once and slices the result locally, here about nine times faster.

Batching needs the data to sit in a contiguous block, which is partly why thoughtful device makers group related registers together. When the values you need are scattered with gaps, it's a judgment call: one big read that also pulls registers you don't need, or several smaller reads. Usually the single read with some waste still wins, because overhead dominates; but a very large gap can tip it the other way. Measure when it matters.

PYTHON — one batched read, sliced locally
rr = client.read_holding_registers(0, count=20)   # registers 0-19, one transaction
if not rr.isError():
    level, setpoint = rr.registers[0], rr.registers[1]
    flow_regs = rr.registers[4:6]                   # a two-register float

Tiered polling

Not all data needs the same freshness. A pump's pressure may need reading several times a second; its run-hours counter changes meaningfully once a minute; its firmware version never changes. Tiered polling gives each point a rate that matches its need: a fast tier every cycle, a medium tier every few cycles, a slow tier rarely.

A grid of twelve poll cycles. The fast tier, pressure, is read in every cycle. The medium tier, run hours, is read in cycles 1, 5 and 9. The slow tier, firmware, is read only in cycle 1. A highlight steps across the cycles. Twelve cycles cost 16 requests instead of 36. 123456 789101112 CYCLE → Fastpressure Mediumrun hours Slowfirmware version 12 cycles: 12 + 3 + 1 = 16 requests instead of 36
Figure 17.4 Tiered polling. Fast, control-critical points every cycle; medium points every few cycles; slow or static points rarely. Each value gets the freshness it needs and no more.

Tiering pairs naturally with batching: within each tier, group the contiguous reads. Read the fast block every cycle, the medium block occasionally, the slow block seldom, each as one batched request. It takes a little planning, but it's the difference between a network that comfortably serves a hundred points and one that chokes on twenty.

The slow-device problem

On a serial bus transactions happen one at a time, so one slow or unresponsive device stalls the entire loop. If a device has gone offline and your timeout is two seconds, every cycle wastes two full seconds on it; with three attempts (two retries), six seconds. One dead node can cripple the refresh rate of a dozen healthy ones. It's the most common reason a system that "worked yesterday" feels broken today.

Two timelines of the same 6.2 seconds fill left to right. Without backoff, five healthy reads take a sliver of time, then three 2-second timeouts on the dead device fill the rest: one cycle completes. With backoff, short cycles repeat back to back and the dead device is probed only once every tenth cycle: about 25 cycles complete in the same time. No backoff: device 4 dead, 2 s timeout × 3 attempts Backoff: skip device 4, probe it every 10th cycle timeoutretry · timeoutretry · timeout one probe · timeout 1 cycle in 6.2 s ≈ 25 cycles in the same 6.2 s healthy reads (5 devices × 33 ms) waiting on the dead device
Figure 17.5 One slow device stalls the loop. While the client waits out the timeout, every other device goes unread. Backing off, temporarily skipping a failing device, keeps one bad node from starving the rest.

The cure has two parts. First, tune the timeout and retries to your bus: long enough that a healthy-but-slow device isn't falsely declared dead, and no longer, because every extra second is paid on every failed poll. Second, back off from failing devices: when one misses several polls in a row, stop polling it every cycle and check it only occasionally, so a dead node costs one slow probe now and then instead of a full timeout every cycle. When it answers again, return it to the normal rotation.

Poll-cycle calculator (Modbus RTU, FC03 reads)

–bytes per healthy transaction
–ms per healthy transaction
–cycle time
–transactions per second
Where the cycle's time goes

Model: each read is FC03 (request 8 bytes, response 5 + 2N bytes), each frame followed by a 3.5-character gap (fixed 1.75 ms above 19200 baud). A dead device costs one request plus a full timeout per attempt; once a device fails, the client skips its remaining reads that cycle. Backoff probes a dead device once (no retries) every 10th cycle; the cycle time shown is the average.

Scaling out

When one bus is genuinely full even after batching and tiering, scale out rather than up.

  • On serial: split devices across several buses, each with its own adapter and its own polling loop, so the loops run in parallel and each bus carries less traffic.
  • On TCP: open several connections and let them work concurrently, using the pipelining that the transaction identifier from Chapter 7 makes possible.

The principle is the same in both worlds: parallelism buys throughput that a single serial conversation, bound to its one-at-a-time discipline, simply can't offer. Four buses of ten devices refresh four times faster than one bus of forty.

Check your understanding

1. You need holding registers 0–39 from one device. What's fastest?

Each transaction pays fixed overhead. One contiguous read of 40 (under the 125 limit) pays it once.

2. Roughly how many bytes per second does a 9600-baud bus carry with 8N1 framing?

Each byte travels as a 10-bit character (start, 8 data, stop), so 9600 ÷ 10 = 960 bytes/s. With 11-bit characters (8E1, 8N2) it's about 873.

3. One of twelve devices goes offline; timeout 2 s, three attempts. What happens to the cycle?

Serial is one transaction at a time, so the client waits out every attempt. Backoff is the cure.

4. A firmware-version register never changes. How should it be polled?

Tiered polling gives each point the freshness it needs and no more, freeing the bus for fast, control-critical values.