A thousand well-behaved chats are not a well-behaved fleet

  • performance
  • telegram
  • node
  • postgres

My Wordle Mini App draws a live scoreboard into every group chat that plays it — one photo per chat, edited in place as people guess, refreshed at most once every ten seconds. Ten seconds is generous per chat and I had never questioned it. What I wanted to know was what the puzzle rollover looks like if the bot is in a thousand groups and ten people in each open the app at the same minute. So I built a rig: the real compiled server against the real database, real signed initData, a local stub standing in for Telegram that enforces its actual ceilings (30 sends a second, 20 a minute per chat), and the process pinned to two cores so the numbers meant something for a VPS.

The first run answered in nine seconds. The server process was gone.

Not out of memory — 134 MB resident when it died, no stack, no log line. That is a native crash, and the only native thing in the process is the canvas library that draws the board. Twenty lines of script, no application code: create N canvases of the size a board actually is, draw into them, encode them to PNG concurrently. Twelve survive. Twenty segfault. The lockfile pinned @napi-rs/canvas at 1.0.2 and npm ci installs the lockfile, so that was exactly what production had been running. 1.0.7 takes eighty.

Which raises the question of why ten boards were rendering at the same instant in the first place, and that is the actual story. The ten-second debounce is scoped to a chat. Every chat’s leading edge fires the moment the word changes. A thousand chats each doing one modest, well-behaved thing at the same instant is a thousand simultaneous renders, and nothing anywhere was counting them.

With the crash fixed, the same run got to 11 GB resident and stalled the event loop for 23 seconds. That produced the symptom I had not predicted: 8,285 requests refused at the TCP level. Not slow — refused. The process was alive and healthy by every internal metric; the accept queue had simply overflowed while nothing drained it. Telegram refused 2,597 calls on top of that, and the refusals fed themselves, because the board’s recovery from a rejected edit is delete-and-repost — two more calls than the one that just failed.

The measurement that made the shape obvious was an A/B: the same ten thousand players, the same sixty thousand guesses, but scored to themselves so no shared board is drawn. Without the board, a guess was 7 ms at the median and 14 ms at p99. With boards in a hundred chats — a tenth of the target — the same guess was 2,080 ms and 22,740 ms. The API was free. The picture was the entire cost.

So: a semaphore of three around board rendering, taken before the job is handed to the worker thread rather than inside it, because a queued job’s avatars get structured-cloned across the boundary and a thousand waiting jobs is a thousand copies. A single send queue in front of every outgoing message, installed as a grammY transformer so nothing added later can route around it, where a 429 pauses everyone for its retry_after — the limit is on the bot, not on the chat, so per-chat backoff is the one thing that cannot work. The refresh interval stopped being a constant and started scaling with the fleet. And the live board dropped from @2x to @1x: 169 KB and 178 ms became 76 KB and 48 ms, for an image Telegram downscales to about 1280 px before anyone sees it anyway.

The costs are real: at a thousand active chats the interval lands near 67 seconds, so a busy board can sit a minute behind. JPEG turned out to be worse than useless — the board is flat colour blocks, PNG’s best case, and q80 came out larger, 223 KB against 169.

Then the bottleneck moved somewhere I did not expect. With the boards under control, the API still had a p90 of 2.9 seconds, and the event-loop metric was flat the whole time. The database pool was the answer: the driver defaults to ten connections, and a pool is a queue that requests stand in outside the event loop, so a saturated one looks exactly like a healthy server answering slowly. Ten connections to thirty took p90 from 2,855 ms to 203 ms on the same two cores. One number in a config nobody had ever set.

The last one was avatars — reads, correctly exempt from the send limiter, and therefore unbounded. A thousand first boards asked for ten thousand profile photos at once: about 25,000 HTTP round trips, 6,074 open handles, roughly 13 seconds of latency tail with the event loop perfectly calm. Collapsing concurrent requests for the same user and capping how many are open cost nothing, because an avatar is cosmetic and cached for a day.

The same scenario now: every request served, 387 MB peak instead of 11 GB, no event-loop stall longer than 248 ms instead of 23.7 seconds, zero rate-limit rejections instead of 2,597, eighteen refused connections instead of 8,285, and nothing in the error log at all instead of 1,547 lines.

What is still wrong is that the interval counts chats while the cost scales with cards. A four-card board renders in 19 ms; a full forty-card one takes 171 ms and 212 KB. A handful of large groups eats the budget of dozens of small ones, and nothing in the arithmetic knows that yet.

← all posts