Onnex
Your business has the answer.
Let Onnex find it.
LOADING
. . .
Completed
0

Independent Model Security Research

CASI & ARS Leaderboards

Research by F5 Labs (research originated by CalypsoAI, acquired by F5) — July 2026 snapshot. Onnex republishes this data verbatim, credited, and unedited; these are not Onnex’s own scores.

Snapshot: July 2026

What the numbers mean

CASI — How hard a model is to break. Higher is safer.

A composite score reflecting a model’s “Defensive Breaking Point” — the minimum attack complexity and resources needed to compromise it. Unlike a simple Attack Success Rate, CASI weights attacks by severity, so a trivial jailbreak and a serious compromise don’t count the same.

ARS — How well a model holds up against automated attackers. Scored 0–100.

Measures resilience when attacked by autonomous AI agents rather than human prompts, across three categories: required sophistication, defensive endurance, and counter-intelligence.

Avg. Performance — How capable the model is at ordinary tasks.

Averaged across MMLU, GPQA, MATH and HumanEval.

RTP — The trade-off between staying safe and staying useful.

Higher generally indicates a better balance of security against capability.

CoS — What that security costs you to run.

Inference cost relative to the model’s CASI score. Lower is cheaper.

CASI Leaderboard

How hard a model is to break. Higher is safer. A composite score reflecting a model’s Defensive Breaking Point — the minimum attack complexity and resources needed to compromise it. Unlike a simple Attack Success Rate, CASI weights attacks by severity, so a trivial jailbreak and a serious compromise don’t count the same.

Comprehensive AI Security Index leaderboard, July 2026 snapshot, sortable by column, source F5 Labs
1AnthropicClaude Sonnet 593.0853.40%0.7719.34
2AnthropicClaude Haiku 4.592.5423.70%0.656.48
3AnthropicClaude Opus 4.889.6755.70%0.7633.46
4NVIDIANemotron-3 Ultra83.6837.80%0.654.00
5OpenAIGPT-5.4 Mini82.6816.60%0.566.35
6QwenQwen3.5-397B-A17B81.1332.00%0.615.18
7OpenAIGPT-5.578.4750.40%0.6744.60
8OpenAIGPT-5 Nano77.1119.00%0.540.58
9OpenAIGPT-5.4 Nano76.0817.60%0.531.91
10XiaomiMiMo-V2.573.8040.10%0.600.57

Comprehensive AI Security Index leaderboard, July 2026 snapshot, sortable by column, source F5 Labs

ARS Leaderboard

How well a model holds up against automated attackers. Scored 0–100. Measures resilience when attacked by autonomous AI agents rather than human prompts, across three categories: required sophistication, defensive endurance, and counter-intelligence.

Agentic Resistance Score leaderboard, July 2026 snapshot, sortable by column, source F5 Labs
1AnthropicClaude Sonnet 598.2653.40%0.8018.32
2AnthropicClaude Haiku 4.594.6123.70%0.666.34
3AnthropicClaude Opus 4.894.2955.70%0.7931.82
4OpenAIGPT-5.4 Mini91.5816.60%0.625.73
5QwenQwen3.5-4B91.0416.00%0.610.20
6OpenAIGPT-5 Nano89.9419.00%0.620.50
7QwenQwen3.5-122B-A10B89.8128.10%0.654.01
8OpenAIGPT-5.588.9750.40%0.7439.34
9QwenQwen3.5-27B88.2329.30%0.653.29
10OpenAIGPT-5.4 Nano88.2217.60%0.601.64

Agentic Resistance Score leaderboard, July 2026 snapshot, sortable by column, source F5 Labs

Eleven-Month History

Leading Model Score By Month

93 95 97 99 Sep '25 Nov '25 Jan '26 Mar '26 May '26 Jul '26
  • CASI (solid line, circle marker)
  • ARS (dashed line, square marker)
Score of the leading model each month, September 2025 through July 2026, for both boards. Series are distinguished by line style and marker shape, not color alone.
The same figures as a table
IndexSep 2025Oct 2025Nov 2025Dec 2025Jan 2026Feb 2026Mar 2026Apr 2026May 2026Jun 2026Jul 2026
CASI95.0394.1995.8997.8197.8698.3096.9398.3098.3293.6393.08
ARS93.9994.4295.0198.0598.0598.0598.0594.5794.5794.5798.26

Onnex Commentary

Security And Capability Are Not The Same Axis

Security and capability are not the same axis. Claude Haiku 4.5 ranks 2nd on CASI at 92.54 while scoring 23.70% on average performance, whereas GPT-5.5 scores 50.40% on average performance but ranks only 7th on security. A high security score does not imply a high capability score, or the reverse — evaluate a model on both, not one as a proxy for the other.

Within this dataset, Anthropic models have held the top three CASI positions in every snapshot captured, September 2025 through July 2026. That is a factual pattern in F5 Labs’ data, not a recommendation or an endorsement, and it says nothing about how any given model will perform on your own workload.

Cost of Security also varies widely: roughly 78x across the July 2026 CASI top ten, from 0.57 (MiMo-V2.5) to 44.60 (GPT-5.5). And a model’s CASI and ARS ranks can diverge on the same snapshot — Qwen3.5-4B places 5th on ARS but does not appear in the CASI top ten at all. Rank on one board is not a substitute for checking the other.

How CASI Works

Full scoring methodology, attack taxonomy, and update cadence are published by F5 Labs.