Goals
- A handler or plan failure that might succeed on a second try (CouchDB unavailable, a network blip, a write conflict with a concurrent edit) is retried instead of being recorded as permanent, and keeps being retried until it works.
- The cool-down ramps: the first retry is immediate, then one minute, two, and so on up to five minutes, then every five minutes indefinitely.
- Every handler becomes idempotent, so re-running a batch converges instead of reporting spurious guard failures.
Handler contract
Today a handler returns the operations that failed and any thrown error fails the whole batch. The contract becomes:
- Return the operations that failed permanently: missing doc, guard mismatch, a row error other than
conflict, an operation with no id. These are recorded in failed_operations and never retried.
- Throw for anything that stopped the batch being processed as a whole: a rejected read or write, or rows that came back
conflict. The batch is re-run on the next attempt with the cursor unmoved. Consider using a new named RetryableError so we can distinguish between expected retryable errors and unexpected errors (that should just fail the action.).
| Handler |
Already applied looks like |
Treat as |
set-parent |
doc.parent deep-equals op.parent |
success, skip the write |
set-contact |
doc.contact deep-equals op.contact (both absent counts) |
success, skip the write |
delete |
no row doc in medic |
success (already the case) |
delete-user |
deleteUser rejects with a 404 |
success |
Handler errors that we can retry
- Any error coming back from a PouchDB call with a status of
408, 409, or 5xx - issue with Couch, not with the operation. Need to retry until Couch improves.
- Any
bulkDocs entry that comes back with error: 'conflict' - simultaneous edit. Try again with the updated _rev.
Goals
Handler contract
Today a handler returns the operations that failed and any thrown error fails the whole batch. The contract becomes:
conflict, an operation with no id. These are recorded infailed_operationsand never retried.conflict. The batch is re-run on the next attempt with the cursor unmoved. Consider using a new namedRetryableErrorso we can distinguish between expected retryable errors and unexpected errors (that should just fail the action.).set-parentdoc.parentdeep-equalsop.parentset-contactdoc.contactdeep-equalsop.contact(both absent counts)deletemedicdelete-userdeleteUserrejects with a 404Handler errors that we can retry
408,409, or5xx- issue with Couch, not with the operation. Need to retry until Couch improves.bulkDocsentry that comes back witherror: 'conflict'- simultaneous edit. Try again with the updated_rev.