An Open Benchmark for AI in Cybersecurity Operations.
What you can't measure, you can't improve. What you can't measure, you also can't validate. AI in cybersecurity operations is outpacing both. Vendors ship new agents faster than anyone can compare them on telemetry that resembles what defenders actually see.
SOCBench is being built to be that benchmark: exhaustive and open, putting AI systems through the actual work a SOC does (detection, triage, investigation, hunting, detection engineering, threat intelligence) and scoring them on the dimensions that decide whether anyone can ship them.
Today, one capability is published: Detection. 3 frontier LLMs × 4 analyst personas × 1,205 network-flow units, scored against ground truth. The rest of the SOC follows.
SOCBench scores AI systems on four dimensions: efficacy, cost, latency, and reliability. All four count.
Live = published results · Roadmap = scoped, not yet measured · Exploring = open question whether a clean benchmark is possible