Sagas
A transaction is the right tool when every write is in one database. It is no help when the steps are in four places: charge a card, reserve stock, book a courier, send the receipt. A saga is the answer, and grit generate workflow writes one.
The problem it solves
The card does not roll back when the courier refuses. The process holding all of this in its head is exactly the thing that crashes, usually between the charge and the reservation, and what is left is a customer charged for nothing and no record of why.
A saga gives each step a Do and an Undo. The steps run in order, and when one fails for good the completed ones are undone newest first. Where the run got to is a row in saga_runs, not a stack frame, so a process that dies resumes rather than losing the thread.
Generating one
grit generate workflow Checkout --steps "charge,reserve_stock,book_courier,send_receipt"
That writes internal/sagas/checkout.go with the steps stubbed out and registers it. Each Do returns an error until you write it, so a half-built saga fails loudly rather than reporting success and doing nothing: the first you would otherwise hear of it is a customer saying the parcel never came.
func Checkout() saga.Definition {return saga.Definition{Name: "checkout",Steps: []saga.Step{{Name: "charge",Do: func(ctx context.Context, r *saga.Run) error {var in CheckoutInputif err := r.Input(&in); err != nil {return err}charge, err := payments.Charge(ctx, in.OrderID, payments.Idempotency(r.IdempotencyKey()))if err != nil {if payments.Declined(err) {return saga.Fatal(err) // retrying will not fix a declined card}return err}r.Set("charge_id", charge.ID)return nil},Undo: func(ctx context.Context, r *saga.Run) error {id := r.GetString("charge_id")if id == "" {return nil // the charge never landed, so there is nothing to refund}return payments.Refund(ctx, id)},},// ...},}}
Starting a run
run, err := saga.Start(ctx, db, "checkout", CheckoutInput{OrderID: order.ID},saga.Key("order:"+order.ID))
It records the run and returns at once. A runner advances it, and every replica runs one: a run is claimed before it is touched, so two of them cannot execute the same step and charge a card twice. The key makes the start idempotent, which is what you want behind a retried request or a redelivered webhook.
Start it in the same transaction as the write that justifies it and the two commit together, the way outbox.Enqueue does.
The three rules
- Every Do must be idempotent. A crash between the side effect and the record of it is not preventable, so a resumed run will sometimes repeat a step. Pass
r.IdempotencyKey()to whatever you are calling: it is stable for a given run and step, so a provider that deduplicates on it will not charge twice. - Every Undo must be idempotent too, and must tolerate a
Dothat never finished. Compensation runs because something went wrong, and that includes not knowing whether the charge landed. Check before you reverse. - Steps you cannot undo go last. An email cannot be unsent, so
Undo: nilis a legitimate thing to write. The consequence is the ordering: a step after it that fails will compensate everything before it and leave that email sent. Sending the receipt before the parcel is booked is a bug you only find in production.
r.Set and r.Get, because by the time the refund runs the process that made the charge is usually gone. What a step writes is saved even when that step then fails, which is deliberate: a charge id obtained just before a timeout is exactly what the refund needs.What the statuses mean
running: working through the steps.compensating: a step failed for good and the completed ones are being undone.done: every step completed.compensated: a step failed and everything before it was undone. The world is back where it started, which is the good outcome of a bad run.stuck: a compensation itself failed past its attempts. This is the one that needs a person. Something happened that could not be taken back, and the run says so instead of carrying on with a status that reads as resolved.
A compensation is never skipped, because skipping one is how money goes missing. A failed step is retried with exponential backoff and jitter; saga.Fatal(err) skips the retries for a failure that retrying will not fix.
Finding a stuck run
The admin has a screen at /system/sagas: every run, filtered by status, with its steps in order and the error that stopped each one. A stuck run gets a Retry button, which is the only state where a person pressing something is the right answer. It resumes compensating from the step it stopped on, because the thing that could not be undone is still not undone and whoever fixed the refund API wants that same undo attempted again.
Without the screen the only way to find a stuck run is SQL against saga_runs, which means nobody finds one until a customer complains.
Not the same as a workflow field
Workflows turn a status column into a state machine: the states are one record's own, and the transitions are usually a person clicking a button. A saga is the other thing: work that crosses systems, with nobody watching, that has to either finish or be undone. A project often wants both, and they do not overlap.
