Case study · Slack Datastores · 2026 · Tech lead (DRI)

OPS Framework: Claude Code as the execution engine for on-call

The OPS Framework is a collaborative investigation toolkit for on-call and incident response on Slack's Vitess fleet. The engineer forms a hypothesis. Claude helps validate it by running commands, parsing data and correlating across signals. The framework supplies the structure: runbooks, guardrails, memory and the analysis engines.

25
runbooks, from lag triage to AZ drains
~1.8k
lines of Go in the Sherlock analysis engine
~1.4k
lines of Python and shell parsers for other data sources
98
commits in 11 weeks, Mar–Jun 2026

The problem

Database incident response at Slack's scale meant juggling Grafana, SSH, MySQL, vtctld, S3 and Slack threads, under time pressure, with tribal knowledge about which commands to run and in what order. Handing an LLM a shell helps, but it brings two failure modes I wasn't willing to accept on production databases:

The design: split judgment from computation

Claude does

Judgment and orchestration

  • Match a short prompt to a runbook
  • Pull env, keyspace and shard out of hostnames
  • Decide the next step from what the data shows
  • Draft the write-up, data first and hypotheses second
Code does

Everything that must repeat

  • Parse replay JSON into typed snapshots
  • Align replicas to a common 5-second grid
  • Detect spikes and classify bottlenecks
  • Separate the root-cause table from amplified ones
Engineer does

Intent and accountability

  • Form the hypothesis
  • Approve confirm: true steps
  • Edit the interpretation
  • Post to the incident thread

How a run works

Promptlag {tablet}or dig, e-trx, schema-change {thread}…
LoadRunbook + pitfallsYAML steps, knowledge/common.md and per-capability notes
ExecuteAdaptive stepsCommand, reasoning, conditional and confirm steps
AnalyzeParser binariessherlock-analyze, dig-now-analyze.py, vtorc-hist-analyze.py
Recordops logEvery command stored in JSONL with hierarchical context tags

A runbook is guidance, not a script

Runbooks are YAML. Steps can be shell commands, reasoning checkpoints where Claude interprets output and picks a branch, condition gates, or confirm: true steps that stop for a human before anything changes state. Claude may skip irrelevant steps or add ad-hoc ones. The structure keeps it on rails without turning it into a brittle script.

runbooks/lag.yaml (abridged, internal hostnames removed)

name: lag
syntax: "lag {tablet}"
context_template: ["{env}", vitess, "{keyspace}", replication, lag]
steps:
  - name: check_replication
    command: ssh {tablet} 'sudo msql -- -e "SHOW REPLICA STATUS\G"' | grep -E "Seconds_Behind|..."
  - name: check_backup_lock_waiters
    condition: "seconds_behind > 0"
    command: ... PROCESSLIST WHERE State="Waiting for backup lock" ...
  - name: check_backup_process
    condition: "backup_lock_waiters > 0"
  - name: check_schema_change_progress
    condition: "ongoing schema change detected in vtops status"
    grafana: https://grafana/.../vitess-online-ddl?var-keyspace={keyspace}&var-shard={shard}
  - name: cancel_online_ddl
    condition: "schema change is causing significant lag and user confirms cancellation"
    confirm: true          # a human approves before anything changes state
  - name: analyze
    type: reasoning
    description: Determine the root cause. Common patterns: backup holding
      BACKUP LOCK and blocking DDL in replication; heavy binlog traffic;
      consolidator contention (check Sherlock replay); IO saturation.
      Report findings with specific numbers.

Pitfalls: lessons that stick

When a session goes wrong (the wrong binary name, a script piped over SSH that should run locally, UTC used where the tool expects local time, a "complete" report covering a failed job), the fix goes into the runbook's pitfalls: field or into a knowledge/ note. CLAUDE.md tells the agent to treat pitfalls as hard constraints, not suggestions. Each mistake costs one session at most.

runbooks/dig.yaml (excerpt)

pitfalls:
  - "Run diag-tablet.sh LOCALLY — it handles SSH internally. Do NOT pipe it over SSH to the tablet."
  - "Times (--start/--end) are LOCAL timezone (PDT), not UTC."
  - "Tablet naming for unsharded keyspaces: shard 0 is '00' not '00-00'."

Command logging and recall

Every command Claude runs is written to a daily JSONL audit log. The entry records the capability, session, step, command, exit code, an output summary and context tags such as prod, vitess, {keyspace}, {keyspace}/{shard}, replication. The ops CLI searches that history with AND logic:

# everything ever run against one keyspace, today
$ ops byteam -d today
# every schema-change command touching one table
$ ops schema-change native_rules
# replay a whole session
$ ops sessions -d today && ops session <id>

That gives the team an auditable record of what the agent actually did. It also lets Claude look up how a similar problem was solved last time.

The deterministic core: sherlock-analyze

The flagship parser is a standalone Go tool that processes Sherlock replay data from every tablet in a shard at once. It exists because comparing each replica at its own peak is misleading. You have to compare all replicas at the same timestamp.

StageFileWhat it does
Parseparse.goS3 replay JSON → typed snapshots, with data-gap detection
Alignalign.goPuts all replicas on a common 5-second grid, including tablets since removed from topology
Detectanalyze.goAdaptive spike threshold (6× baseline or 5 s/s absolute QTime), merging gaps under 20 s
Classifyanalyze.goLabels each spike as consolidation, pool exhaustion, tx-pool exhaustion or MySQL pressure, with a stall look-back
Attributeanalyze.goRoot-cause table (sublinear spike, high baseline budget) vs. amplified tables (superlinear), plus onset-share attribution
Reportreport.goMarkdown: topology, baselines, spike timeline, cross-replica comparison, root-cause analysis

Excerpt from a generated report (identifiers anonymized)

## Spike #1 Root Cause Analysis (primary at 12:39:30)

Root cause (dominant steady-state workload, sublinear spike):
Table                 Baseline QTime  Budget  Onset share  Peak QTime     Table spike
Update.inbox_items    1.91 s/s        151%    73%          3225.29 s/s    1685.2x
Select.inbox_items    1.18 s/s         93%     0%            22.00 s/s      18.6x

Primary tx pool: 250/250 → 0/250 for 35s  ·  replicas: no spikes, healthy pools

Run the same window twice and you get the same report. Claude's job is to collect the inputs, call the binary, and write the narrative around numbers it did not compute.

The other parsers

Runbook catalog

Investigate

  • sherlock-analyze cross-replica spike forensics
  • dig single-tablet deep dive at a point in time
  • dig-now live query analysis across a shard
  • lag replication lag triage
  • e-trx errant transaction GTID diff and binlog decode
  • coredump vttablet core dump analysis
  • phist / vtorc-hist reparent and recovery history

Operate

  • schema-change from an escalation thread to verified rollout
  • upscale / upscale-ks clone-based shard and keyspace upscales
  • clone / delete AZ-aware replica lifecycle
  • az-drain batched primary drains during AWS AZ incidents
  • backup, chef-env, vtadmin-fix, vtgate-cdc-hot

Triage & route

  • triage-placement reads the alert, checks history and topology, drafts a reply
  • escal-for owning team and on-call via the service catalog

Team memory

  • qa saves and answers operational Q&A from threads
  • tool catalogs observability links and tests MCP access
  • daily-save / daily-update standup queue

Integrations go through MCP: Slack (threads, drafts), Grafana/Prometheus, Honeycomb, logging and the internal service catalog. Messages bound for a channel are always posted as drafts for a human to send.

Guardrails for writing up incidents

The framework's instructions encode how findings get communicated, not just how they get computed:

Roadmap: toward an automated first responder

When I left, analyses were triggered by a person and took about 20 minutes each. Fifteen ran in 2026, and every one was validated by a human reviewer. The next step I designed was a phased path to an agent that posts data-backed triage to the alert thread within about 60 seconds:

  1. Phase 0–1Auto-trigger

    A cron poller, then a Slack socket-mode listener, detects unanalyzed alerts and starts the runbook headlessly.

  2. Phase 2Headless data access

    Swap SSH hops for direct S3 reads (IAM role), the vtadmin API and PromQL. This is the critical path.

  3. Phase 3–4Agent service + structured output

    A Go service running the Messages API tool-use loop. --json output and P0–P3 severity scoring built on the existing spike structs.

  4. Phase 5–7Graduated trust

    Shadow channel → drafts → 🤖-prefixed posts → full auto. Then a feedback loop from ✅/❌ reactions and proactive detection of baseline drift.

The trust ladder is deliberate. Each stage advances only on a measured record, for example 20 consecutive analyses with less than 5% of them needing correction.