Case study · Slack Datastores · Jul 2025 – May 2026 · Architect

Sherlock: security-camera footage for the database fleet

Sherlock is a forensics framework I architected and deployed across Slack's Vitess fleet of more than 11,000 database servers. Every tablet records a 5-second snapshot of system, MySQL and vttablet state and ships it to S3. When something breaks, you can rewind to the exact moment and see which query plan on which table was burning the CPU.

5 s
snapshot interval on every enabled tablet
14 d
of replay history in S3, after the tablet itself is gone
25
active weeks from proof of concept to live production use
15
production incidents investigated with Sherlock data in 2026

Why Prometheus wasn't enough

Grafana told us that something happened. It couldn't tell us why. Take a CPU spike to 100% for 30 seconds inside a 5-minute window that otherwise sits at 20%. Under rate()[5m] it averages out to about 28%, which looks like a mild bump. A 100% burst for 30 seconds and a 40% plateau for 90 seconds can produce the same average and are very different problems. And even when you spot the bump, Prometheus can't say which queries were running at that moment.

Prometheus / Grafana
  • 30 s scrapes; severity diluted by rate() averaging
  • Short-and-intense looks the same as long-and-mild
  • Correlating CPU to table or plan is manual
  • Short retention at high resolution
Sherlock
  • 5 s snapshots: a 30 s spike shows up as 6 consecutive readings
  • Query plan breakdown captured in the same snapshot as the CPU reading
  • Heap, file descriptors, memory, disk I/O and MySQL state all together
  • Replay files on S3, analyzed with a single CLI command

What every snapshot captures

System

  • CPU %, load average (1/5/15m)
  • Memory: total, used, cached
  • Disk utilization and I/O rates
  • Network throughput and connections, open FDs

VTTablet

  • Query plans by type and table, with request rates
  • Query-time concurrency per plan.table
  • Connection and transaction pool state, consolidation waits
  • Rows returned and rows affected per plan

Runtime

  • Go heap: alloc, in-use, idle, object count
  • GC metrics

MySQL

  • Threads running, connection states
  • InnoDB buffer pool hit rate
  • Slow query counter

Architecture

vtagent · every tabletMetrics monitorCollects every 5 s, with threshold checks for CPU, memory and disk
vtagentReplay filesJSON snapshots, rotated at 10 MB (about 5 minutes of data)
vtagentS3 uploaderBatched upload, local files deleted on success, flushed on shutdown
vtctld · vtops-goCLIreplay list / range / view, enable, history
OPS FrameworkAnalysissherlock-analyze across all replicas, driven by Claude

The CLI

Before the OPS Framework existed, Sherlock already had an operator interface built into vtops-go:

sherlock replay range — snapshot summary and query-plan analysis for a window (output abridged, table names anonymized)

$ sudo vtops-go sherlock replay range -t tablet-…-1k93 -s 08:11 -e 08:21

Timestamp                       CPU    Heap    Mem     Disk    Consol   Rows
2026-03-19T08:11:02-07:00     41.6%   58MiB   77.1%   39.4%    1.0/s   473/s
2026-03-19T08:11:07-07:00     84.2%   65MiB   77.0%   39.4%    0.6/s   494/s
2026-03-19T08:11:12-07:00  ** 96.8%   82MiB   76.9%   39.4%    5.8/s   596/s

--- Query Plan Analysis ---
Top tables by avg query time concurrency:
  Insert.table_a:              avg 0.144x, peak 0.374x (120 snapshots)
  Insert.table_b:              avg 0.138x, peak 0.385x (120 snapshots)

replay view is an interactive TUI. You step through snapshots with h/j/k/l, filter the list by CPU threshold, and press p to copy a permalink to the exact snapshot for the Slack thread.

From recording to answers

Per-tablet replay was a big step, but the first real incidents exposed its limit. Comparing each replica at its own peak is misleading. That gap drove sherlock-analyze, the cross-replica Go analysis engine, and led to the OPS Framework that has Claude orchestrate it. Each incident fed the next round of tooling:

  1. Mar 2026 · consolidator deadlockReplay timed what a core dump couldn't

    Sherlock pinpointed the exact onset of a deadlock to the second: a single global mutex with allocation inside the lock, causing a GC-assist circular wait. That led to better replay ranges and a consolidated diag-tablet.sh.

  2. Mar 2026 · pool exhaustionThree-dump comparison

    Core dumps taken at 1 h, 7 h and near failure confirmed a progression theory and identified connection-pool exhaustion. That led to the coredump CLI and the shard-level analyzer.

  3. Apr 2026 · comparative shardsDegraded capacity, amplified

    Consolidator backlog from a 2-replica shard plus slow background queries drove 620–1065× concurrency spikes. That led to multi-tablet correlation in sherlock-analyze.

  4. May 2026 · transaction-pool stallDetection added the same day

    An InnoDB row-lock convoy drained the primary's transaction pool to 0/250 for 35 seconds. Stall look-back detection for tx-pool exhaustion shipped within hours.

Roadmap I proposed

Collection
  • Forensic collector and event bundler triggered on threshold breach
  • Gzip bundles on S3 that survive a dead tablet, so no SSH is needed
Analysis
  • Baseline-driven anomaly detection: per-shard P5/P50/P95 envelopes
  • Automatic "what's different?" overlays and recurring-signature detection
  • An automated first responder through the OPS Framework