Benchmark fairness
Interpret MigrationRoom's UI comparison without overstating a single run.
The workshop passes on method and reasoning, not a fixed speedup. The MigrationRoom Benchmark view is interpretable only after query correctness and environment context are clear.
Correctness first
For every source/target pair, confirm both statements represent the same business question, return the same columns, and produce the same logical result. Pay special attention to NULL ordering, JSON extraction, array expansion, divide-by-zero behavior, and the selected dataset snapshot.
A missing, failed, or semantically different query blocks its performance conclusion. Do not read the aggregate speedup first and explain mismatches afterward.
Timing context
MigrationRoom prefers Databricks server time when query-history access is available; it falls back to wall time otherwise. Those measurements answer different questions. Label the timing type shown by the agent and dashboard.
Record cold and warm observations separately. A Databricks warehouse cold start is a real operational result, but it should not be compared silently with a warm ClickHouse query. Keep these fields beside any number you quote:
| Field | Why it matters |
|---|---|
| Query pair and returned result | Establishes the compared workload |
| Engine/version and region | Gives optimizer and topology context |
| Warehouse/service type and size | Describes available compute and scaling |
| CPU and memory per worker/replica | Makes exposed physical capacity explicit; record not exposed rather than infer it |
| Cluster/replica count and scaling range | Separates vertical capacity from horizontal concurrency capacity |
| Autoscaling or idle-resume event | Prevents a scale transition from masquerading as engine execution time |
| Cold/warm label and warm-up | Separates startup from steady execution |
| Server time or wall time | Prevents incompatible timing labels |
| Failure or omission | Prevents a fast partial run from looking complete |
Hardware and service equivalence
There is no reliable one-line conversion between a Databricks SQL warehouse size and a ClickHouse Cloud service size. Databricks serverless may abstract the underlying worker instance, while ClickHouse Cloud expresses capacity through replica count and vertically scalable memory per replica. Record what each control plane exposes and mark hidden physical details as not exposed.
For this workshop, record at least:
- Databricks warehouse mode and size, cloud/region, scaling configuration, channel or runtime, and whether the run incurred a start or scale event;
- ClickHouse Cloud tier, cloud/region, replica count, current and min/max memory per replica, vCPU per replica when exposed, version, and whether the service resumed or scaled; and
- benchmark concurrency, other active workloads, cache policy, and timing type.
Use the learner module's system.asynchronous_metrics query to capture the active
ClickHouse replicas and their current CGroupMaxCPU and CGroupMemoryTotal limits.
Those are effective runtime limits, not the service's configured autoscaling range.
Tier, region, and min/max scaling bounds remain control-plane metadata and must be
recorded separately. CGroupMemoryUsed is useful for detecting obvious pre-existing
pressure, but it must not be reported as hardware capacity.
The Terraform default is a serverless 2X-Small Databricks SQL warehouse. The workshop
does not provision the ClickHouse Cloud target, so no fixed ClickHouse hardware claim is
valid across participants. The official
Databricks sizing documentation
and ClickHouse Cloud scaling documentation
describe the different scaling controls.
Before and after optimization
Re-fire Benchmark after Step 6 using the same logical query, services, timing type, and cache-state policy. The physical SQL may change because the optimized version reads a projection, dictionary, or materialized serving table, but the returned business result must not.
One UI observation is not a concurrency, throughput, p95, or cost study. If the delivery needs those claims, run a separately designed load test outside the participant path and publish its request count, concurrency, timeout, failures, and distribution.
Comparison limits
Databricks and ClickHouse have different storage, caching, autoscaling, scheduling, and pricing semantics. Do not infer a universal product result from one TPC-H-derived workload, region, service size, or selectively optimized query.
The workshop demonstrates how ClickHouse can combine native MergeTree layout, dictionaries, insert-triggered materialized views, projections, and catalog-backed data for this serving workload. The defensible conclusion is scoped to the query and configuration shown in the MigrationRoom dashboard.
Review order
- Confirm semantic equivalence and successful results.
- Confirm engine, region, service size, cold/warm state, and timing type.
- Inspect every failed or omitted query.
- Compare the per-query source and target timings.
- State the scoped conclusion, operational cost, and retest condition.