Retrying transactions
On a synchronously replicated cluster a transaction can fail at COMMIT even though every statement succeeded. Applications must retry such transactions.
Which errors to retry#
| Error | Meaning | Action |
|---|---|---|
1213 deadlock |
a conflicting transaction won certification | retry the whole transaction |
1205 lock wait timeout |
a row stayed locked too long | retry, then investigate long transactions |
1047 not ready |
the server is rejoining the cluster | reconnect and retry |
2006, 2013 connection lost |
failover or maintenance; while connecting, also the connection rate limit | reconnect and retry, backing off at least a second when it happens at connect time |
The pattern#
Retry the whole transaction, from BEGIN, a bounded number of times with a short, randomised pause.
python
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 | |
Make the work repeatable#
A retried transaction runs twice from the application's point of view. Keep side effects — sending mail, calling another service — outside the transaction, or make them idempotent.
Reduce conflicts#
Keep transactions short, touch rows in a consistent order, and avoid hot counter rows that every request updates.