NEW LogLM and open-weights model results now in Detection
NEW On PatchEval (ByteDance), PatchLoop clears OpenHands and SWE-agent by 20+ points. See how

SOCBench

An Open Benchmark for AI in Cybersecurity Operations.

What you can't measure, you can't improve. What you can't measure, you also can't validate. AI in cybersecurity operations is outpacing both. Vendors ship new agents faster than anyone can compare them on telemetry that resembles what defenders actually see.

SOCBench is being built to be that benchmark: exhaustive and open, putting AI systems through the actual work a SOC does (detection, triage, investigation, hunting, detection engineering, threat intelligence) and scoring them on the dimensions that decide whether anyone can ship them.

Today, one capability is published: Detection. 3 frontier LLMs × 4 analyst personas × 1,205 network-flow units, scored against ground truth. The rest of the SOC follows.

SOCBench grew out of work that DeepTempo does to validate and improve their LogLM and the open source AI-SOC Vigil. Everyone is welcome to contribute and to help lead the SOCBench project.

The Benchmark

SOCBench scores AI systems on four dimensions: efficacy, cost, latency, and reliability. All four count.

Efficacy
Does it get the right answer?
F1 precision recall verdict accuracy
Cost
What does it actually take to run?
USD / alert token economics
Latency
How fast does it respond?
MTTD P50 P95
Reliability
Does it return well-formed output, every time, within budget?
completion rate defect rate

Current Capabilities and Road Map

Detection
Live
Classify flows, logs, or alerts as malicious or benign (the core work of the monitoring tier).
See results →
Triage & escalation
Roadmap
Rank alerts, dedupe, decide what gets paged and what gets closed.
Investigation & DFIR
Roadmap
Multi-step reasoning over evidence; timeline reconstruction and root-cause attribution.
Threat hunting
Roadmap
Proactive search across telemetry for adversary TTPs without a triggering alert.
Detection engineering
Roadmap
Synthesize rules and queries, explain false-positives, propose detection coverage.
Threat intelligence
Roadmap
IOC enrichment, threat-actor attribution, MITRE ATT&CK mapping.
Vulnerability detection and patching
Live
Identify exploitable vulnerabilities across assets, prioritize by risk, and propose or apply patches.
See PatchLoop harness →
Response & remediation
Exploring
Execute playbooks, recommend containment, draft customer-facing comms.

Live = published results  ·  Roadmap = scoped, not yet measured  ·  Exploring = open question whether a clean benchmark is possible