Sherlock is a forensics framework I architected and deployed across Slack's Vitess fleet of more than 11,000 database servers. Every tablet records a 5-second snapshot of system, MySQL and vttablet state and ships it to S3. When something breaks, you can rewind to the exact moment and see which query plan on which table was burning the CPU.
Why Prometheus wasn't enough
Grafana told us that something happened. It couldn't tell us why. Take a CPU spike to 100% for 30 seconds inside a 5-minute window that otherwise sits at 20%. Under rate()[5m] it averages out to about 28%, which looks like a mild bump. A 100% burst for 30 seconds and a 40% plateau for 90 seconds can produce the same average and are very different problems. And even when you spot the bump, Prometheus can't say which queries were running at that moment.
- 30 s scrapes; severity diluted by
rate()averaging - Short-and-intense looks the same as long-and-mild
- Correlating CPU to table or plan is manual
- Short retention at high resolution
- 5 s snapshots: a 30 s spike shows up as 6 consecutive readings
- Query plan breakdown captured in the same snapshot as the CPU reading
- Heap, file descriptors, memory, disk I/O and MySQL state all together
- Replay files on S3, analyzed with a single CLI command
What every snapshot captures
System
- CPU %, load average (1/5/15m)
- Memory: total, used, cached
- Disk utilization and I/O rates
- Network throughput and connections, open FDs
VTTablet
- Query plans by type and table, with request rates
- Query-time concurrency per
plan.table - Connection and transaction pool state, consolidation waits
- Rows returned and rows affected per plan
Runtime
- Go heap: alloc, in-use, idle, object count
- GC metrics
MySQL
- Threads running, connection states
- InnoDB buffer pool hit rate
- Slow query counter
Architecture
replay list / range / view, enable, historysherlock-analyze across all replicas, driven by Claude- Controller in DynamoDB. A switch hierarchy (global > keyspace > shard > tablet) turns collection on or off at any scope without a deploy. Every change is recorded in an audit history.
- Safety-gated core dumps. When a tablet stops responding, Sherlock can capture a core dump automatically, guarded by a CPU check and a shard lock so it never makes an incident worse.
- No data loss on restart. The active file is flushed and uploaded when vtagent stops. A 14-day local cleanup and an S3 lifecycle policy keep storage bounded.
- Stack: Go (embedded in the vtops agent), AWS SDK v2, S3, DynamoDB, Chef for fleet-wide configuration.
The CLI
Before the OPS Framework existed, Sherlock already had an operator interface built into vtops-go:
sherlock replay range — snapshot summary and query-plan analysis for a window (output abridged, table names anonymized)
$ sudo vtops-go sherlock replay range -t tablet-…-1k93 -s 08:11 -e 08:21 Timestamp CPU Heap Mem Disk Consol Rows 2026-03-19T08:11:02-07:00 41.6% 58MiB 77.1% 39.4% 1.0/s 473/s 2026-03-19T08:11:07-07:00 84.2% 65MiB 77.0% 39.4% 0.6/s 494/s 2026-03-19T08:11:12-07:00 ** 96.8% 82MiB 76.9% 39.4% 5.8/s 596/s --- Query Plan Analysis --- Top tables by avg query time concurrency: Insert.table_a: avg 0.144x, peak 0.374x (120 snapshots) Insert.table_b: avg 0.138x, peak 0.385x (120 snapshots)
replay view is an interactive TUI. You step through snapshots with h/j/k/l, filter the list by CPU threshold, and press p to copy a permalink to the exact snapshot for the Slack thread.
From recording to answers
Per-tablet replay was a big step, but the first real incidents exposed its limit. Comparing each replica at its own peak is misleading. That gap drove sherlock-analyze, the cross-replica Go analysis engine, and led to the OPS Framework that has Claude orchestrate it. Each incident fed the next round of tooling:
- Mar 2026 · consolidator deadlockReplay timed what a core dump couldn't
Sherlock pinpointed the exact onset of a deadlock to the second: a single global mutex with allocation inside the lock, causing a GC-assist circular wait. That led to better replay ranges and a consolidated
diag-tablet.sh. - Mar 2026 · pool exhaustionThree-dump comparison
Core dumps taken at 1 h, 7 h and near failure confirmed a progression theory and identified connection-pool exhaustion. That led to the
coredumpCLI and the shard-level analyzer. - Apr 2026 · comparative shardsDegraded capacity, amplified
Consolidator backlog from a 2-replica shard plus slow background queries drove 620–1065× concurrency spikes. That led to multi-tablet correlation in
sherlock-analyze. - May 2026 · transaction-pool stallDetection added the same day
An InnoDB row-lock convoy drained the primary's transaction pool to 0/250 for 35 seconds. Stall look-back detection for tx-pool exhaustion shipped within hours.
Roadmap I proposed
- Forensic collector and event bundler triggered on threshold breach
- Gzip bundles on S3 that survive a dead tablet, so no SSH is needed
- Baseline-driven anomaly detection: per-shard P5/P50/P95 envelopes
- Automatic "what's different?" overlays and recurring-signature detection
- An automated first responder through the OPS Framework