
Skynet Breakage Atlas
Adversarial multi-model failure atlas — subscription controls (Claude / Codex / Grok) vs Skynet fleet cells, with flag rates, empty replies, hard errors, and latency.
Overview
The Skynet Breakage Atlas answers a practical routing question: when traffic goes through our proxy fleet instead of first-party subscription surfaces, which models go empty, hard-error, or slow-walk under adversarial families?
Live at breakage.hermanity.dev.
What’s measured
- Controls: Claude Max, Codex, Grok (subscription / direct)
- Fleet: Skynet-routed models (deepseek, glm, kimi, minimax, …)
- Families: cascade seed, context boundary, cost threshold, empty probe, format JSON, multi-intent
- Metrics: flag rate, empty rate, hard error rate, p50 latency
Snapshot (public table on the site)
glm-5.2 is the known-broken reference (flag rate 1.0 across cells in the frozen N-batch). Controls stay low-flag; fleet variance is the point of the atlas — empty replies are a first-class failure mode, not a footnote.
Why it exists
Model bench scores quality. This atlas scores breakage under load-bearing failure families so routing and canary work have a public, re-checkable board instead of vibes from one chat session.
The broader AI Agent Reliability guide places empty responses, hard errors, route ambiguity, stale state, and recovery behavior in one operational framework.
Related
- Live: breakage.hermanity.dev
- Sibling: Model Bench