I build operations automation for large database fleets. Most recently at Slack, I was the tech lead on the OPS Framework and Sherlock. Sherlock records 5-second forensic snapshots across the Vitess fleet (11,000+ database servers). The OPS Framework turns Claude Code into an on-call partner that runs codified runbooks against that data.
Claude handles the judgment. Code handles anything that has to come out the same way twice.
Claude reads the alert, picks the runbook, runs the commands and adapts when the data surprises it. Parsing, time alignment, spike detection and classification are done by tested Go and Python code. Any number that ends up in an incident thread comes from code, not from a model reading raw output.
Figures are from the project repositories and incident log as of mid-2026.
The pattern
Each OPS runbook splits the work the same way. The model never parses raw replay data itself, and the parsers never decide what to look at next.
sherlock-analyze pool/80-88 at 12:39pitfalls that later runs treat as hard constraints. Gaps in the data become new parser features, sometimes shipped the same day.Selected work
OPS Framework: Claude Code as an on-call force multiplier
An AI-native operations framework. It covers YAML runbooks with reasoning steps and confirmation gates, an audit log of every command with recall, a store of operational knowledge, and deterministic analysis engines that Claude calls instead of improvising.
Read the case study →Sherlock: database forensics
5-second replay snapshots from every Vitess tablet, uploaded to S3 and switched on or off per scope through DynamoDB. It works like security-camera footage for the database.
Read the case study →Ten years of platform automation
A Vitess v12–v14 upgrade control plane, Tier-0 disaster recovery, a patented RDS provisioner, and database platforms on Kubernetes.
See experience →What I can help with
Make Claude useful on call
- Turn tribal knowledge into runbooks Claude can follow
- Audit logs, recall, confirmation gates and guardrails
- A path from human-triggered runs to an automated first responder
Parsers and analyzers
- Go CLIs and analysis engines with tests
- Time-series alignment and anomaly classification
- Reports a reviewer can check line by line
Database infrastructure
- Vitess/MySQL, RDS and Aurora fleets
- Upgrade, provisioning and DR control planes
- Observability with incident-grade resolution