07 Validate and tear down — instructor notes
Run the quiz as a real attempt, review the questions a room reliably misses, and facilitate the teardown as a checklist nobody leaves without completing.
Budget 35 to 45 minutes: 20 to 30 for the fourteen questions and their review, and 15 for the teardown. The teardown is the half that cannot be dropped when the session overruns — it is the workshop's only ongoing cost and it is a daily one in two separate accounts.
The quiz
Fourteen questions, graded client-side and instantly. Nothing is sent anywhere and you will not see a score, so do not plan a review around results you do not have. The pass mark the bank carries is 10 of 14.
Facilitation:
- Open-book but individual. The learner pages are the point of an open-book format; a neighbour's answer is not. Say both halves.
- One real attempt first. Answers and explanations stay hidden until every question is answered and submitted, which is deliberate — it makes the first pass a measurement rather than a browse.
- Tell them to read the explanations on the questions they got right too. Several name the failure mode behind the distractor, which is information the correct answer alone does not carry.
- Then review by show of hands, theme by theme, since you have no scores. Ask "who changed their mind on this one while answering?" rather than "who got it right" — it surfaces the reasoning without making anyone report a failure.
The questions a room reliably splits on, and what each one is really testing:
| Theme | What trips people | Where to send them |
|---|---|---|
| Covering both legs | Anyone who learned ClickPipes before the Managed Postgres destination shipped will pick the option saying it only lands data in ClickHouse | module 01's table of the two legs |
confirmed_flush_lsn catching pg_current_wal_lsn() | Every option is a true statement about replication; only one answers what the LSN comparison establishes | module 03 Step 5 |
| What reconciliation proves | "Every line matches" feels total; the query hashes one expression per table at one instant | sql/05_reconcile.sql, and module 04's reconcile step |
| Choosing the ordering key | Read patterns and dedup correctness pull in different directions, and the question asks what should drive the choice | module 05 Step 3 |
What IMPORT FOREIGN SCHEMA changes | Nothing, which does not feel like a plausible answer to a question about a migration step | module 06 Step 1 |
| What PostgresBench substantiates | The temptation is to let a transactional number carry an analytical claim | module 04's citation and its narrowing |
Use the rationales page for the review, and open it only after the room has submitted. It is grouped by theme and written for someone who got a question wrong, so it works as a discussion script.
If a question's explanation does not land for someone, send them to the module rather than to the reference page. The reference page adds reasoning; the module has the context.
Teardown
Run this as a facilitated checklist with the whole room together, out loud, in order. Do not publish it as homework: the participants who most need it are the ones who have run out of time and energy, and an RDS instance left up costs roughly USD 4 a day until somebody destroys it.
The order is not the reverse of the build order, and two dependencies decide it:
- The migration pipe holds a slot on RDS. Delete it while RDS is still reachable and the remote slot goes with it. Destroy RDS first and the pipe has nobody to talk to.
- The analytics pipe holds a slot on Managed Postgres. Delete the pipe before the instance.
Both dependencies are the same shape, which is worth saying to the room: a pipe's slot lives on its source, so every pipe must be deleted before the thing it was reading from.
The sequence: stop the local processes, delete the migration pipe and confirm no slot remains on
RDS, delete the analytics pipe and confirm no slot remains on Managed Postgres, terraform destroy,
delete the Managed Postgres instance, delete the ClickHouse service, then keep the artifacts and
delete .env.local.
Verify rather than assume, and do it twice for AWS: terraform state list prints nothing, and both
aws rds describe-db-instances and aws rds describe-db-snapshots --snapshot-type manual name
nothing belonging to the workshop. A destroy that failed partway is the expensive case, and Terraform
reporting success is not the same as AWS holding nothing.
Point out what the destroy reverts that matters most: the wide port 5432 allowlist opened in module 03. In production that allowlist is the exception you make for a migration and revert the day it finishes, and this is that day.
Common failures
-
terraform destroyerrors on a dependency. Usually a security group still referenced, or the instance still deleting. Re-run it; RDS deletion takes several minutes and the second run finishes what the first started. -
RDS was destroyed before the pipe was deleted. The pipe now cannot reach the source to remove the slot it created, and it will report errors until it is deleted. Delete it from the console; there is nothing to clean up on the source, because the source is gone.
The orphaned slot on a destroyed instance is harmless — the instance went with it — but say why it matters anyway: the same mistake against a live source leaves a slot pinning WAL, which is module 03's incident with nobody left to resume the pipe. This is the whole reason the teardown order is pipe first, instance second, and it is worth naming out loud rather than leaving it as a numbered step people follow without reading.
-
A slot remains on RDS after the pipe is deleted. Drop it explicitly with
pg_drop_replication_slot, then re-check. Do this beforeterraform destroy, while there is still an instance to check. -
A manual snapshot exists.
skip_final_snapshotis set so the destroy leaves none, but a participant who took one by hand keeps paying for it. Thedescribe-db-snapshotscheck is what catches it. -
The ClickHouse organization has other services. Check names before confirming a deletion. A participant deleting a colleague's service is the one irreversible mistake available in this module.
-
Somebody leaves early. Get the teardown done before the closing discussion, not after. If a participant must leave, walk them through it individually first — a promise to do it later is the most expensive artifact anyone takes home.
End checkpoint
- Every row of the learner page's "nothing left running" table is confirmed rather than assumed.
terraform state listprints nothing, and no AWS query names a workshop resource.- Both replication slots are gone, on both instances, and they went before the instances did.
bench/results/holds both runs with their config blocks..env.localis deleted.
Closing the session
Spend the last five minutes on what leaves the room, and be specific about it, because it is not the infrastructure:
runbook/cutover.md— written before there was a target to rush at, executed under a window, and amended with what it got wrong. That is the artifact with a shelf life, and it is the reason the workshop puts module 02 before module 03 rather than handing out a runbook.- The two
EXPLAINplans that show one role reaching ClickHouse and another staying local through the same connection string. - The two benchmark files with their config blocks, and the sentence that narrows what they prove.
Then ask the closing question: which of the day's failures would their current monitoring have caught? For most rooms the honest answer is that an inactive replication slot would not have been noticed until the volume filled, and a panel that stopped pushing down would not have been noticed at all.