Vendor benchmarks stop at the socket

Server capacity under synthetic load is the axis everyone publishes. What fraction of messages reaches a handset on a real carrier network is the other one, and it is measured device-side or not at all.

EMQX publishes a benchmark of 5,000,000 concurrent MQTT connections on a single node, at a 2.93 millisecond average connect response time.1emqx.com · Aug 10, 2023 It ran on one 64-core machine with 128 GB of RAM, clients arriving at 5,000 per second, held for thirty minutes, inside a cloud provider’s own network. What it measures is capacity.

Every commercial number in this market sits on the same axis. AWS IoT Core caps a single connection at 100 messages per second or 512 KB per second.2docs.aws.amazon.com · read Aug 14, 2026 Azure IoT Hub is sized and billed in messages per day, from 400,000 on S1 to 300,000,000 on S3.3azure.microsoft.com · read Aug 14, 2026 Connections held, messages accepted, messages billed. All of it counts what the server did.

There is a second axis, and no vendor publishes anything on it. It is the fraction of messages that reach a handset: across a carrier’s NAT, into a radio that is trying to sleep, on an operating system working against any process that wants to stay awake. Nothing downstream of the socket belongs to the vendor, so no vendor benchmark reaches it.

A team relying on Firebase Cloud Messaging and Apple Push Notification service alone measured delivery at 75 to 85 percent across user segments, in Courier: Reimagining How We Send Push Notifications4gojek.io · Oct 23, 2023.

That figure is the delivery rate of two push notification services, measured at the device. No broker appears in it, EMQX included. A server capacity benchmark and a device-side delivery rate measure different systems, and neither is evidence about the other. The grade on it is stated: a single self-published source, not independently verifiable.

So a capacity number and a delivery complaint can both be accurate in the same meeting, and the meeting is still no closer to knowing which layer to open.

What sits between the two is the part nobody sells. In Doze mode, Android will not fire an app’s alarm more than once every 9 minutes,5gojek.io · Oct 1, 2021 which bounds how often a client can be woken to prove its socket is still there, whatever the protocol keepalive is set to. That restriction is what sent the same team looking for a keepalive long enough to need fewer alarms, and the interval a network will tolerate varies from carrier to carrier. Four strategies were compared for discovering it per network at runtime, binary, exponential, composite and linear, over test intervals such as 2, 4, 8, 16, 32 minutes and 4, 8, 12, 16, 20 minutes, bounded at 4 minutes below and 120 above, in Adaptive Heartbeats For Our Information Superhighway6gojek.io · Nov 16, 2021. That post is an algorithm comparison and reports no numeric outcome.

As of a search on Aug 14, 2026, not one of HiveMQ, EMQX, Cedalo, Ably or PubNub publishes anything on adaptive keepalive discovery per carrier network, on Doze mode alarm restrictions, or on choosing a transport by battery cost. Finding the value a given radio will tolerate is left to whoever is holding the phone. Worth re-running before anyone leans on it.

The engineering that closes the gap is glue, and it is dull to describe. Send over the long-lived MQTT connection where one exists, fall back to the push services otherwise, at QoS 1, with client-side deduplication, because at-least-once means the same message will sometimes arrive twice. Measured device-side on the same applications, against the same 75 to 85 percent baseline: 95 percent or better on the consumer app, and 99 percent or better on the driver and merchant apps.4gojek.io · Oct 23, 2023 The four points between those two figures are the Doze problem showing up in a delivery rate. A consumer app spends most of its life backgrounded, radio asleep and wakeups throttled, which is where the missing messages go. A driver or merchant app is open for the length of a shift, so the socket stays live and there is far less to recover from. A later migration of driver bid notifications onto that path reported 99.9 percent or better against 97 to 99 percent before, with about a 22 percent reduction in P99 bid acceptance latency, rolled out to tier 3 cities first, then tier 2, then tier 1.7gojek.io · Dec 8, 2023 Both carry the same grade as the baseline they improve on: stated, single-source, not independently verifiable, and for the bid migration the baseline definition is not published either.

All three are delivery rates measured on the device, on production applications, published alongside the baselines they replaced. The work behind one is specific and unglamorous: which transport wins while the socket is up, how long a client waits before it stops believing the socket, what happens to the message that arrives twice, and how anyone knows afterwards that it arrived at all. Only the client can answer the last one, and only if someone built it to. That gets decided years before anybody asks what the delivery rate is.