Independent Model Security Research
CASI & ARS Leaderboards
Research by F5 Labs (research originated by CalypsoAI, acquired by F5) — July 2026 snapshot. Onnex republishes this data verbatim, credited, and unedited; these are not Onnex’s own scores.
Snapshot: July 2026
What the numbers mean
CASI — How hard a model is to break. Higher is safer.
A composite score reflecting a model’s “Defensive Breaking Point” — the minimum attack complexity and resources needed to compromise it. Unlike a simple Attack Success Rate, CASI weights attacks by severity, so a trivial jailbreak and a serious compromise don’t count the same.
ARS — How well a model holds up against automated attackers. Scored 0–100.
Measures resilience when attacked by autonomous AI agents rather than human prompts, across three categories: required sophistication, defensive endurance, and counter-intelligence.
Avg. Performance — How capable the model is at ordinary tasks.
Averaged across MMLU, GPQA, MATH and HumanEval.
RTP — The trade-off between staying safe and staying useful.
Higher generally indicates a better balance of security against capability.
CoS — What that security costs you to run.
Inference cost relative to the model’s CASI score. Lower is cheaper.
CASI Leaderboard
How hard a model is to break. Higher is safer. A composite score reflecting a model’s Defensive Breaking Point — the minimum attack complexity and resources needed to compromise it. Unlike a simple Attack Success Rate, CASI weights attacks by severity, so a trivial jailbreak and a serious compromise don’t count the same.
| 1 | Anthropic | Claude Sonnet 5 | 93.08 | 53.40% | 0.77 | 19.34 |
|---|---|---|---|---|---|---|
| 2 | Anthropic | Claude Haiku 4.5 | 92.54 | 23.70% | 0.65 | 6.48 |
| 3 | Anthropic | Claude Opus 4.8 | 89.67 | 55.70% | 0.76 | 33.46 |
| 4 | NVIDIA | Nemotron-3 Ultra | 83.68 | 37.80% | 0.65 | 4.00 |
| 5 | OpenAI | GPT-5.4 Mini | 82.68 | 16.60% | 0.56 | 6.35 |
| 6 | Qwen | Qwen3.5-397B-A17B | 81.13 | 32.00% | 0.61 | 5.18 |
| 7 | OpenAI | GPT-5.5 | 78.47 | 50.40% | 0.67 | 44.60 |
| 8 | OpenAI | GPT-5 Nano | 77.11 | 19.00% | 0.54 | 0.58 |
| 9 | OpenAI | GPT-5.4 Nano | 76.08 | 17.60% | 0.53 | 1.91 |
| 10 | Xiaomi | MiMo-V2.5 | 73.80 | 40.10% | 0.60 | 0.57 |
Comprehensive AI Security Index leaderboard, July 2026 snapshot, sortable by column, source F5 Labs
ARS Leaderboard
How well a model holds up against automated attackers. Scored 0–100. Measures resilience when attacked by autonomous AI agents rather than human prompts, across three categories: required sophistication, defensive endurance, and counter-intelligence.
| 1 | Anthropic | Claude Sonnet 5 | 98.26 | 53.40% | 0.80 | 18.32 |
|---|---|---|---|---|---|---|
| 2 | Anthropic | Claude Haiku 4.5 | 94.61 | 23.70% | 0.66 | 6.34 |
| 3 | Anthropic | Claude Opus 4.8 | 94.29 | 55.70% | 0.79 | 31.82 |
| 4 | OpenAI | GPT-5.4 Mini | 91.58 | 16.60% | 0.62 | 5.73 |
| 5 | Qwen | Qwen3.5-4B | 91.04 | 16.00% | 0.61 | 0.20 |
| 6 | OpenAI | GPT-5 Nano | 89.94 | 19.00% | 0.62 | 0.50 |
| 7 | Qwen | Qwen3.5-122B-A10B | 89.81 | 28.10% | 0.65 | 4.01 |
| 8 | OpenAI | GPT-5.5 | 88.97 | 50.40% | 0.74 | 39.34 |
| 9 | Qwen | Qwen3.5-27B | 88.23 | 29.30% | 0.65 | 3.29 |
| 10 | OpenAI | GPT-5.4 Nano | 88.22 | 17.60% | 0.60 | 1.64 |
Agentic Resistance Score leaderboard, July 2026 snapshot, sortable by column, source F5 Labs
Eleven-Month History
Leading Model Score By Month
- CASI (solid line, circle marker)
- ARS (dashed line, square marker)
The same figures as a table
| Index | Sep 2025 | Oct 2025 | Nov 2025 | Dec 2025 | Jan 2026 | Feb 2026 | Mar 2026 | Apr 2026 | May 2026 | Jun 2026 | Jul 2026 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| CASI | 95.03 | 94.19 | 95.89 | 97.81 | 97.86 | 98.30 | 96.93 | 98.30 | 98.32 | 93.63 | 93.08 |
| ARS | 93.99 | 94.42 | 95.01 | 98.05 | 98.05 | 98.05 | 98.05 | 94.57 | 94.57 | 94.57 | 98.26 |
Onnex Commentary
Security And Capability Are Not The Same Axis
Security and capability are not the same axis. Claude Haiku 4.5 ranks 2nd on CASI at 92.54 while scoring 23.70% on average performance, whereas GPT-5.5 scores 50.40% on average performance but ranks only 7th on security. A high security score does not imply a high capability score, or the reverse — evaluate a model on both, not one as a proxy for the other.
Within this dataset, Anthropic models have held the top three CASI positions in every snapshot captured, September 2025 through July 2026. That is a factual pattern in F5 Labs’ data, not a recommendation or an endorsement, and it says nothing about how any given model will perform on your own workload.
Cost of Security also varies widely: roughly 78x across the July 2026 CASI top ten, from 0.57 (MiMo-V2.5) to 44.60 (GPT-5.5). And a model’s CASI and ARS ranks can diverge on the same snapshot — Qwen3.5-4B places 5th on ARS but does not appear in the CASI top ten at all. Rank on one board is not a substitute for checking the other.
How CASI Works
Full scoring methodology, attack taxonomy, and update cadence are published by F5 Labs.