How we found and fixed transactions lost behind a nonce gap

Under steady load on a four-validator network we found that a sender submitting transactions back to back could lose most of them. In one 90-second run only 14 of 291 transfers were confirmed. Consensus was never at risk: blocks, state roots and balances were identical on every node. Users' transactions were what got lost.

How we found it

We moved the network tests into the repository (tools/quorum-net). Every run starts four validators from genesis and keeps its parameters, the logs of every node and a summary. The first run on the clean build passed with no errors. Only a series of repeats showed that the defect was intermittent and older than the work in progress. The logs gave the exact picture: the transaction with nonce N had not yet reached the block proposer, while N+1, N+2… were already in a block, failed on their nonce and were dropped from the pool.

First cause: how transactions were picked for a block

The mempool sorted a sender's transactions by nonce but never checked that they continued from the committed one. We looked at how go-ethereum (its pending and queued sets), Cosmos EVM, Aptos and Polygon Bor, which hit almost the same bug, handle this. They share one rule, and we adopted it: only a contiguous run of nonces goes into a block, and a transaction behind a gap waits in the pool. Every run since has had zero nonce errors.

Second cause, uncovered by the fix: losses in transit

Now a validator that never received transaction N waited honestly, and its blocks came out empty. Measurements showed each peer missing 2–5 transactions per run. Every message opened its own TCP connection, and our DoS protection, a limit on connections per address, cut some of them. The fix follows Bitcoin, which re-announces until delivery, and Solana, which re-sends every two seconds: a node re-sends its pending transactions in one batch until they are included in a block.

Third step: persistent connections between peers

The same limit was cutting consensus messages too, so about half of the heights were decided only in the second round, with stalls of up to 12 seconds. Following CometBFT and libp2p, each node now keeps one long-lived connection per peer. The DoS protection is unchanged; honest peers simply no longer need to open hundreds of connections.

Result

At the same limits: 296 of 296 transfers confirmed, the longest stall 1.8 seconds instead of 12, and no connection dropped. Next comes a handshake signed with ML-DSA between peers, so that limits apply to a validator's identity rather than to an IP address.

All news