04 Cut over — instructor notes
Run the window as a timed exercise, let the sequence step fail before it is fixed, and stop a participant improvising the runbook they spent module 02 writing.
Budget 35 to 50 minutes. The window itself is under a minute; everything else is the discipline around it.
The exercise is executing a document, not migrating a database
Participants wrote runbook/cutover.md in module 02 and have not executed it. The point of this
module is finding out what their own document was missing under time pressure, so the two rules
worth stating before anyone starts are: execute it as written, and note changes rather than
making them. A participant who improves the runbook mid-window has run a different exercise and
has nothing to amend afterwards.
Expect the sequence step to be the omission. It is the one nobody writes on the first pass.
The sequence step: let it fail, on a timer
A participant whose runbook omits the sequence fix will repoint the writer and watch it fail on every insert. Let that happen. It is instant, cheap, fully reversible, and it is the workshop's central lesson arriving as a consequence rather than as a bullet point. A room where half the participants hit it and half did not is the ideal outcome, because the two halves can then talk to each other.
Cap it at about three minutes. When the writer's log is full of duplicate-key errors and the
participant has read one, point at sql/04_sequences.sql and have them verify rather than trust
it: last_value from pg_sequences against max(order_id), equal, before the writer goes back
in.
Then ask the room the question that makes it transferable: which check in a data-focused cutover checklist would have caught this? None of them. Row counts match, checksums match, lag is zero, every panel renders. The failure is in an object that holds state and is not table data, and there are four of them in this schema.
Where the time goes
| Beat | Attention | Wall clock |
|---|---|---|
| Quiesce writes, confirm level | 5 min | seconds |
| Fix the sequences, including the failing insert | 15 min | 3 min |
| Repoint the writer and the dashboard | 10 min | window is PROVISIONAL: 38 seconds |
| Reconcile, stop replicating | 10 min | 5 min |
| Amend the runbook | 10 min | none |
Common failures
- Sequences not fixed before the writer restarts.
duplicate key value violates unique constraint "orders_pkey", repeatedly. It is the lesson, so let it happen; the recovery issql/04_sequences.sqland a writer restart. - The reconcile check reads as a failure.
diffreports a difference once the writer is on the target, which is expected. The check below it is the verdict, and it is the one to point at. default_transaction_read_onlyleft on after an abort. A participant who aborts and rolls back to RDS meets a database that refuses every write, and it looks like a permissions problem.- A participant improves the runbook mid-execution. Ask them to note the change and apply it in Step 2 instead.
End checkpoint
Do not start module 06 until, for every participant:
- the writer is committing against
$TARGET_HOSTwithfailedat zero; - every sequence's
last_valueequals its table'smax(id); - the reconcile check prints
reconciled; pg_replication_slotson RDS returns no rows; andrunbook/cutover.mdhas been amended with what their first draft missed.
03 Replicate — instructor notes
Pace the replication leg, run the failure injection as a synchronized room-wide beat, and know exactly which mistakes here strand a participant for half an hour.
05 Stream the analytical copy — instructor notes
Hold the line on the module claiming no measured improvement, spend the initial load on the engine and ordering-key decision, and catch the pipe that was pointed at RDS.