Lost connection while connecting
New connections fail with:
1 | |
(or your driver's equivalent: "connection reset", "EOF during handshake"). Existing connections keep working, and the failures stop on their own shortly afterwards.
Why it happens#
Each source IP address may open at most 300 new connections in any 10 seconds to a cluster endpoint. Past that, new connections are closed as soon as they're accepted, before the database has sent anything. So the client sees a dropped connection rather than an access error.
The window slides: connections opened shortly before a burst count toward it. Refused attempts count too, so a client that retries immediately stays blocked for as long as it keeps trying.
Everything behind one NAT gateway or egress address shares the limit. That includes many containers or serverless functions.
Fix#
- Use a connection pool and keep connections open, instead of opening one per request or per job.
- Back off on this error with a randomised delay of at least a second, rather than retrying at once. See Retrying transactions.
- Stagger start-up when many instances start together, for example after a deploy, so they don't all fill their pools in the same few seconds.
- If your workload really needs a higher rate from one address, contact ScaiLabs support.