Databricks MigrationRoomClickHouse Workshops

04 Benchmark in the dashboard — instructor notes

Enforce correctness and timing context before interpreting the UI comparison.

Timing

Budget 25 minutes. Fire Benchmark only after validation and accepted rewrites.

Talk track

The dashboard benchmark is a guided per-query comparison, not the previous standalone concurrency harness. Before latency, require matching logical results, explicit source and target query pairing, service configuration, cold/warm state, and Databricks server-time versus wall-time labeling.

Require the compute record before learners click Benchmark. The default source is a serverless 2X-Small Databricks SQL warehouse, but its physical host details may be platform-managed. For ClickHouse Cloud, record the service tier, replicas, and the effective vCPU and memory per replica returned by the system.asynchronous_metrics query. Use the Cloud settings only for control-plane facts SQL does not expose: tier, region, and configured scaling bounds. Do not let learners equate a Databricks size label with a ClickHouse replica size or call the result hardware-normalized.

Ask each pair to name the weakest ClickHouse query and the observed reason it may be slow. That becomes the Step 6 hypothesis. Attractive numbers without a successful query on both sides or without timing context are not results.

Common failures

  • A cold Databricks start is compared with a warm ClickHouse query without disclosure.
  • Wall time is labeled server time.
  • A query failed or was omitted, but the aggregate is still quoted.
  • Learners copy a service label but do not run the effective-capacity SQL probe.
  • Service size, replicas, region, or scaling state is missing from the evidence.
  • Learners describe a single dashboard run as a concurrency or throughput benchmark.

Reset steps

Preserve the first UI result, correct query pairing or service state, and re-fire the complete Benchmark step under declared conditions. Do not retain only the best rerun.

ในหน้านี้

TH