Databricks Migration with MigrationRoom
Let an AI-assisted migration move a Databricks workload, then prove what should be copied, queried in place, or served from ClickHouse.
This four-hour workshop, including a 15-minute break, is a UI-led architecture exercise built around MigrationRoom. You will use its dashboard to migrate a Databricks TPC-H workload to ClickHouse, validate the copy, query two external catalogs, build a native hot path, and defend where each workload belongs.
What is MigrationRoom?
MigrationRoom is an AI-powered migration workspace for moving analytical workloads to ClickHouse Cloud. It combines a guided dashboard with a source-specific AI agent. The dashboard supplies the workload context and six migration actions; the agent inspects the source through governed tools, reasons about the ClickHouse design in chat, and executes approved work while the dashboard reports live state.
It is more than a SQL translator:
| Advantage | What it changes in this workshop |
|---|---|
| One guided workflow | Discovery, schema design, copy, validation, rewrite, benchmark, and optimization stay in one conversation and dashboard |
| Source-aware tooling | The agent discovers Databricks through its read-only MCP instead of asking you to paste schemas or run ad-hoc scripts |
| Human approval gates | You challenge physical design and approve DDL before data moves |
| Observable long-running work | Migration progress, table state, validation, and benchmarks remain visible even when the chat is quiet |
| ClickHouse-native redesign | The agent can move beyond syntax translation to ORDER BY, projections, dictionaries, materialized views, and catalog federation |
| Resumable context | Conversations and migration runs can be resumed, inspected, and re-fired without rebuilding the exercise from memory |
The AI accelerates the work; it does not own the decision. Learners remain responsible for approving schema choices, stopping on validation failures, interpreting benchmarks, and deciding what should be copied, federated, or retained in Databricks.
The site has three tracks. The Learner track is the hands-on playbook. The Instructor track mirrors every module with timing, delivery notes, recovery steps, and publication gates. Reference pages explain the operator workflow, type mapping, catalog and governance boundaries, and benchmark method.
Three paths, one decision
| Path | What you exercise | Where governance applies |
|---|---|---|
| Copy | MigrationRoom reads Databricks SQL and writes ClickHouse native tables | ClickHouse RBAC governs the copied rows after they cross the boundary |
| Zero-copy | DataLakeCatalog exposes Unity Catalog and an instructor REST catalog to ClickHouse | Catalog and object-store authorization govern the external reads; confirm the deployed identity flow |
| Hybrid | One query joins native hot data to Unity and REST catalog tables | Each external read and the native copy retain separate controls |
The goal is not an all-or-nothing replatform. Keep open, governed lake data where that fits; materialize only the workloads whose measured serving needs justify a native copy.
Before the session
Allow 25-45 minutes of prework for Module 00.
You need Git, Python 3.11 or newer, Terraform 1.5 or newer, Docker with Compose v2,
make, account-admin access to the Databricks account used for the workshop, a
ClickHouse Cloud service, and an LLM provider configured for MigrationRoom. Module 00
creates a new Databricks workspace and shows how to create its Terraform service
principal, client ID, and OAuth secret. Use workshop-scoped credentials and never paste
credentials into chat or screenshots.
The instructor pre-attaches read-only Unity and REST catalog databases, supplies the REST fixture's expected invariant, and publishes the tested ClickHouse version. Live cloud results are completed in the room; this guide does not pre-populate successful output.
Modules
| # | Module | Time | Outcome |
|---|---|---|---|
| 00 | Enter the room | 25-45 min prework | The new Databricks workspace and MigrationRoom dashboard are ready |
| 01 | Establish the source baseline | 20 min | Step 1 discovers live source counts, types, features, and workload inputs |
| 02 | Review and approve the design | 40 min | The proposed DDL is challenged and approved in chat before data moves |
| 03 | Migrate, validate, and rewrite | 45 min | UI-driven copy, parity, and query rewriting complete |
| — | Break | 15 min | A background migration may continue while the room pauses |
| 04 | Benchmark in the dashboard | 25 min | Correctness and timing context precede latency interpretation |
| 05 | Query multiple catalogs | 35 min | The agent queries Unity, REST, and native tables together through ClickHouse |
| 06 | Build the native hot path | 40 min | Step 6 implements an approved native optimization and re-benchmarks it |
| 07 | Defend and tear down | 20 min | Placement decisions are defended and participant resources are removed |