Databricks MigrationRoomClickHouse Workshops

Benchmark fairness

Interpret MigrationRoom's UI comparison without overstating a single run.

The workshop passes on method and reasoning, not a fixed speedup. The MigrationRoom Benchmark view is interpretable only after query correctness and environment context are clear.

Correctness first

For every source/target pair, confirm both statements represent the same business question, return the same columns, and produce the same logical result. Pay special attention to NULL ordering, JSON extraction, array expansion, divide-by-zero behavior, and the selected dataset snapshot.

A missing, failed, or semantically different query blocks its performance conclusion. Do not read the aggregate speedup first and explain mismatches afterward.

Timing context

MigrationRoom prefers Databricks server time when query-history access is available; it falls back to wall time otherwise. Those measurements answer different questions. Label the timing type shown by the agent and dashboard.

Record cold and warm observations separately. A Databricks warehouse cold start is a real operational result, but it should not be compared silently with a warm ClickHouse query. Keep these fields beside any number you quote:

FieldWhy it matters
Query pair and returned resultEstablishes the compared workload
Engine/version and regionGives optimizer and topology context
Warehouse/service type and sizeDescribes available compute and scaling
CPU and memory per worker/replicaMakes exposed physical capacity explicit; record not exposed rather than infer it
Cluster/replica count and scaling rangeSeparates vertical capacity from horizontal concurrency capacity
Autoscaling or idle-resume eventPrevents a scale transition from masquerading as engine execution time
Cold/warm label and warm-upSeparates startup from steady execution
Server time or wall timePrevents incompatible timing labels
Failure or omissionPrevents a fast partial run from looking complete

Hardware and service equivalence

There is no reliable one-line conversion between a Databricks SQL warehouse size and a ClickHouse Cloud service size. Databricks serverless may abstract the underlying worker instance, while ClickHouse Cloud expresses capacity through replica count and vertically scalable memory per replica. Record what each control plane exposes and mark hidden physical details as not exposed.

For this workshop, record at least:

  • Databricks warehouse mode and size, cloud/region, scaling configuration, channel or runtime, and whether the run incurred a start or scale event;
  • ClickHouse Cloud tier, cloud/region, replica count, current and min/max memory per replica, vCPU per replica when exposed, version, and whether the service resumed or scaled; and
  • benchmark concurrency, other active workloads, cache policy, and timing type.

Use the learner module's system.asynchronous_metrics query to capture the active ClickHouse replicas and their current CGroupMaxCPU and CGroupMemoryTotal limits. Those are effective runtime limits, not the service's configured autoscaling range. Tier, region, and min/max scaling bounds remain control-plane metadata and must be recorded separately. CGroupMemoryUsed is useful for detecting obvious pre-existing pressure, but it must not be reported as hardware capacity.

The Terraform default is a serverless 2X-Small Databricks SQL warehouse. The workshop does not provision the ClickHouse Cloud target, so no fixed ClickHouse hardware claim is valid across participants. The official Databricks sizing documentation and ClickHouse Cloud scaling documentation describe the different scaling controls.

Before and after optimization

Re-fire Benchmark after Step 6 using the same logical query, services, timing type, and cache-state policy. The physical SQL may change because the optimized version reads a projection, dictionary, or materialized serving table, but the returned business result must not.

One UI observation is not a concurrency, throughput, p95, or cost study. If the delivery needs those claims, run a separately designed load test outside the participant path and publish its request count, concurrency, timeout, failures, and distribution.

Comparison limits

Databricks and ClickHouse have different storage, caching, autoscaling, scheduling, and pricing semantics. Do not infer a universal product result from one TPC-H-derived workload, region, service size, or selectively optimized query.

The workshop demonstrates how ClickHouse can combine native MergeTree layout, dictionaries, insert-triggered materialized views, projections, and catalog-backed data for this serving workload. The defensible conclusion is scoped to the query and configuration shown in the MigrationRoom dashboard.

Review order

  1. Confirm semantic equivalence and successful results.
  2. Confirm engine, region, service size, cold/warm state, and timing type.
  3. Inspect every failed or omitted query.
  4. Compare the per-query source and target timings.
  5. State the scoped conclusion, operational cost, and retest condition.

このページの内容

JA