02 Plan the cutover — instructor notes
Facilitate the runbook exercise while the seed finishes, nudge without giving away the sequence step, and keep the model runbook closed until module 04 has been executed.
Budget 30 to 45 minutes, and run it inside module 01's seed rather than after it. This module provisions nothing, which is exactly why it is here: it is the only module whose cost is attention rather than wall clock, so it is the one that can absorb a wait.
It is also the module where the workshop's central insight is either discovered or handed over, and that depends entirely on how you facilitate it.
The one rule
Do not name the sequence step, and do not let the reference track name it either. The learner page says there is a step nobody thinks of and refuses to say which. Module 03 then either confirms the participant found it or shows them the duplicate-key storm. Either outcome teaches; being told the answer here teaches nothing, and it is unrecoverable — you cannot un-tell it.
That means two concrete things for the room:
- The model runbook stays closed until module 04 Step 2. Say that out loud, with the reason, rather than hoping nobody clicks. A participant who reads it now loses the exercise and gains a document they did not write.
- When you nudge, nudge at the level of the question, never the answer. The productive prompts
are in the learner page already: enumerate every object in
sql/01_schema.sqlthat holds state and is not table data, then look at the column type of all four primary keys. A participant who works that through has found it themselves; a participant who is told "sequences" has not.
Expect roughly half a room to miss it. That is the designed distribution, not a facilitation failure.
What to demo, what to let them do
- Demo the catalog views, not the decisions. Five minutes on your own instance is enough:
pg_publicationandpg_publication_tables, thenpg_replication_slotsandpg_stat_replicationon the source, thenpg_current_wal_lsn()andpg_wal_lsn_diff(). Show what each one prints and stop there. Say that all of these are source-side, which is why the gate they build reads the same whether the consumer is a hand-made subscription or a managed pipe. Participants cannot choose a lag signal if they have never seen the columns; they will not learn anything if you choose it for them. - Let them write alone first. Ten minutes of silence with the skeleton, no discussion. A room that discusses before writing converges on one runbook, and the six decisions stop being six decisions.
- Then pair them for review, and give the reviewer a script: read the other person's runbook as if it were 11pm and you had not written it, and mark every step whose verification column is empty. Peer review finds unverifiable steps far more reliably than the author does.
- Do not review every runbook yourself. In a room of twelve that is the whole module. Sample three, aloud, with the author's permission, and use them to make the general points.
Questions worth asking
Ask these to the room while they write, one at a time, spaced out. Each one is a prompt toward a decision rather than a quiz:
- "What is the difference between the row counts matching and replication having caught up, and can the first be true while the second is false?" It can, on a database taking writes, which is the whole reason a lag signal is a separate decision from verification.
- "Is a lag of exactly zero a state that ever occurs while the writer is running?" No — and that answer decides the order of "quiesce" and "wait for lag" in the window. This is the single highest-value question on the page, because a runbook that waits for zero before quiescing waits forever.
- "
FOR ALL TABLESon a source that gains a table next week: feature, or unreviewed change?" Both answers are defensible; the point is that it is a decision and an explicit list is what a decision looks like. - "After an abort, is the replication slot still on the source?" It is. A participant whose abort procedure does not say so has written a rollback that starts an incident a few hours later, and module 03 Step 6 will show them exactly which incident.
- "For how long after cutover is rolling back to RDS clean?" Until the first write lands on the target that RDS will never see — which is seconds. That reframes rollback from a procedure into a point of no return with something else on the far side of it.
Common failures
- The blank page. A participant stares at the skeleton for ten minutes. Give them one section to start with — section 3, the window — because the six numbered steps are the most concrete part and the rest follows from them.
TODOleft as prose. "Verify replication looks healthy" is aTODOwith confidence. Push for a command and an expected value, per row, or the runbook fails its own module 04 test.- Every step marked reversible. Ask which step is the irreversible-feeling one and why. If the answer is "none", the window has not been thought through.
- Abort and rollback collapsed into one section. They differ by whether the application has been repointed, which is the cheap-versus-expensive line. Make them separate the two out loud.
- The participant who has done a cutover before finds the sequence step in four minutes. Do not let them tell the room. Give them the harder half instead: write the abort procedure that leaves no slot behind on the source, and the point-of-no-return paragraph. Then have them review two neighbours' runbooks.
- The seed finished and nobody noticed. Watch the clock yourself. The moment module 01's seed log ends with ANALYZE, send participants back to module 01 Step 5 for the baseline and let them finish the runbook after it. The baseline is on the critical path; the runbook is not.
End checkpoint
Every participant has runbook/cutover.md with:
- no
TODOremaining; - a numbered window whose every row names a command and a verification, and answers whether the step is reversible;
- an expected write downtime as a number, and a maximum they would tolerate;
- abort criteria that say what state an abort leaves the source's replication slot in; and
- a rollback section naming an explicit point of no return.
Nobody needs to have found the sequence step. What everyone needs is a document concrete enough to execute under time pressure, because module 04 asks them to execute it as written — including its omissions.
Before module 03, confirm module 01's end checkpoint separately: the seed finished and
results/before.json exists from a run that exited 0. A runbook is not a substitute for a
baseline.