The OPS Framework is a collaborative investigation toolkit for on-call and incident response on Slack's Vitess fleet. The engineer forms a hypothesis. Claude helps validate it by running commands, parsing data and correlating across signals. The framework supplies the structure: runbooks, guardrails, memory and the analysis engines.
The problem
Database incident response at Slack's scale meant juggling Grafana, SSH, MySQL, vtctld, S3 and Slack threads, under time pressure, with tribal knowledge about which commands to run and in what order. Handing an LLM a shell helps, but it brings two failure modes I wasn't willing to accept on production databases:
- Non-repeatable numbers. Ask a model to read 240 JSON snapshots per tablet and summarize the spike, and you get a different answer on each run.
- Lost lessons. A mistake corrected in one session gets repeated in the next.
The design: split judgment from computation
Judgment and orchestration
- Match a short prompt to a runbook
- Pull env, keyspace and shard out of hostnames
- Decide the next step from what the data shows
- Draft the write-up, data first and hypotheses second
Everything that must repeat
- Parse replay JSON into typed snapshots
- Align replicas to a common 5-second grid
- Detect spikes and classify bottlenecks
- Separate the root-cause table from amplified ones
Intent and accountability
- Form the hypothesis
- Approve
confirm: truesteps - Edit the interpretation
- Post to the incident thread
How a run works
lag {tablet}or dig, e-trx, schema-change {thread}…knowledge/common.md and per-capability notessherlock-analyze, dig-now-analyze.py, vtorc-hist-analyze.pyops logEvery command stored in JSONL with hierarchical context tagsA runbook is guidance, not a script
Runbooks are YAML. Steps can be shell commands, reasoning checkpoints where Claude interprets output and picks a branch, condition gates, or confirm: true steps that stop for a human before anything changes state. Claude may skip irrelevant steps or add ad-hoc ones. The structure keeps it on rails without turning it into a brittle script.
runbooks/lag.yaml (abridged, internal hostnames removed)
name: lag syntax: "lag {tablet}" context_template: ["{env}", vitess, "{keyspace}", replication, lag] steps: - name: check_replication command: ssh {tablet} 'sudo msql -- -e "SHOW REPLICA STATUS\G"' | grep -E "Seconds_Behind|..." - name: check_backup_lock_waiters condition: "seconds_behind > 0" command: ... PROCESSLIST WHERE State="Waiting for backup lock" ... - name: check_backup_process condition: "backup_lock_waiters > 0" - name: check_schema_change_progress condition: "ongoing schema change detected in vtops status" grafana: https://grafana/.../vitess-online-ddl?var-keyspace={keyspace}&var-shard={shard} - name: cancel_online_ddl condition: "schema change is causing significant lag and user confirms cancellation" confirm: true # a human approves before anything changes state - name: analyze type: reasoning description: Determine the root cause. Common patterns: backup holding BACKUP LOCK and blocking DDL in replication; heavy binlog traffic; consolidator contention (check Sherlock replay); IO saturation. Report findings with specific numbers.
Pitfalls: lessons that stick
When a session goes wrong (the wrong binary name, a script piped over SSH that should run locally, UTC used where the tool expects local time, a "complete" report covering a failed job), the fix goes into the runbook's pitfalls: field or into a knowledge/ note. CLAUDE.md tells the agent to treat pitfalls as hard constraints, not suggestions. Each mistake costs one session at most.
runbooks/dig.yaml (excerpt)
pitfalls: - "Run diag-tablet.sh LOCALLY — it handles SSH internally. Do NOT pipe it over SSH to the tablet." - "Times (--start/--end) are LOCAL timezone (PDT), not UTC." - "Tablet naming for unsharded keyspaces: shard 0 is '00' not '00-00'."
Command logging and recall
Every command Claude runs is written to a daily JSONL audit log. The entry records the capability, session, step, command, exit code, an output summary and context tags such as prod, vitess, {keyspace}, {keyspace}/{shard}, replication. The ops CLI searches that history with AND logic:
# everything ever run against one keyspace, today $ ops byteam -d today # every schema-change command touching one table $ ops schema-change native_rules # replay a whole session $ ops sessions -d today && ops session <id>
That gives the team an auditable record of what the agent actually did. It also lets Claude look up how a similar problem was solved last time.
The deterministic core: sherlock-analyze
The flagship parser is a standalone Go tool that processes Sherlock replay data from every tablet in a shard at once. It exists because comparing each replica at its own peak is misleading. You have to compare all replicas at the same timestamp.
| Stage | File | What it does |
|---|---|---|
| Parse | parse.go | S3 replay JSON → typed snapshots, with data-gap detection |
| Align | align.go | Puts all replicas on a common 5-second grid, including tablets since removed from topology |
| Detect | analyze.go | Adaptive spike threshold (6× baseline or 5 s/s absolute QTime), merging gaps under 20 s |
| Classify | analyze.go | Labels each spike as consolidation, pool exhaustion, tx-pool exhaustion or MySQL pressure, with a stall look-back |
| Attribute | analyze.go | Root-cause table (sublinear spike, high baseline budget) vs. amplified tables (superlinear), plus onset-share attribution |
| Report | report.go | Markdown: topology, baselines, spike timeline, cross-replica comparison, root-cause analysis |
Excerpt from a generated report (identifiers anonymized)
## Spike #1 Root Cause Analysis (primary at 12:39:30) Root cause (dominant steady-state workload, sublinear spike): Table Baseline QTime Budget Onset share Peak QTime Table spike Update.inbox_items 1.91 s/s 151% 73% 3225.29 s/s 1685.2x Select.inbox_items 1.18 s/s 93% 0% 22.00 s/s 18.6x Primary tx pool: 250/250 → 0/250 for 35s · replicas: no spikes, healthy pools
Run the same window twice and you get the same report. Claude's job is to collect the inputs, call the binary, and write the narrative around numbers it did not compute.
The other parsers
dig-now-analyze.pyreads livequerylogzHTML from every tablet in a shard and groups it by table, query pattern, owning service and team, and consolidator use.vtorc-hist-analyze.pyturns VTOrc Slack alerts andvtops-go statusoutput into a de-duplicated recovery timeline with links.coredump-analyze.shextracts Go runtime state from a vttablet core dump (goroutines, memory, scheduler), classifies goroutines and lines them up with Sherlock data.
Runbook catalog
Investigate
sherlock-analyzecross-replica spike forensicsdigsingle-tablet deep dive at a point in timedig-nowlive query analysis across a shardlagreplication lag triagee-trxerrant transaction GTID diff and binlog decodecoredumpvttablet core dump analysisphist/vtorc-histreparent and recovery history
Operate
schema-changefrom an escalation thread to verified rolloutupscale/upscale-ksclone-based shard and keyspace upscalesclone/deleteAZ-aware replica lifecycleaz-drainbatched primary drains during AWS AZ incidentsbackup,chef-env,vtadmin-fix,vtgate-cdc-hot
Triage & route
triage-placementreads the alert, checks history and topology, drafts a replyescal-forowning team and on-call via the service catalog
Team memory
qasaves and answers operational Q&A from threadstoolcatalogs observability links and tests MCP accessdaily-save/daily-updatestandup queue
Integrations go through MCP: Slack (threads, drafts), Grafana/Prometheus, Honeycomb, logging and the internal service catalog. Messages bound for a channel are always posted as drafts for a human to send.
Guardrails for writing up incidents
The framework's instructions encode how findings get communicated, not just how they get computed:
- Data first, interpretation second. Write "data suggests", not "this caused". Flag single-snapshot observations as such.
- Don't dismiss alternatives without evidence. Show the metric that points to row-level contention rather than declaring that an upscale won't help.
- Report outcomes honestly. "1086/1087 jobs succeeded, 1 timed out" is a different fact from "schema verified via replication".
Roadmap: toward an automated first responder
When I left, analyses were triggered by a person and took about 20 minutes each. Fifteen ran in 2026, and every one was validated by a human reviewer. The next step I designed was a phased path to an agent that posts data-backed triage to the alert thread within about 60 seconds:
- Phase 0–1Auto-trigger
A cron poller, then a Slack socket-mode listener, detects unanalyzed alerts and starts the runbook headlessly.
- Phase 2Headless data access
Swap SSH hops for direct S3 reads (IAM role), the vtadmin API and PromQL. This is the critical path.
- Phase 3–4Agent service + structured output
A Go service running the Messages API tool-use loop.
--jsonoutput and P0–P3 severity scoring built on the existing spike structs. - Phase 5–7Graduated trust
Shadow channel → drafts → 🤖-prefixed posts → full auto. Then a feedback loop from ✅/❌ reactions and proactive detection of baseline drift.
The trust ladder is deliberate. Each stage advances only on a measured record, for example 20 consecutive analyses with less than 5% of them needing correction.