Skip to content

SCW AI Trust Index Ranks LLMs by Security Risk

The index lets you compare the security performance of 16 popular LLMs within various frameworks.

Blue AI letters and twisted black strands on purple hexagon background
Image by Steve A Johnson on Unsplash

AI-generated code carries "an average of 15 confirmed vulnerabilities per codebase, 4.3 of which are considered severe,” according to Secure Code Warrior, which has released the new SCW AI Trust Index to help organizations better understand and predict security risks introduced by LLMs.

The research behind the index — done in collaboration with RMIT University — evaluated 1,760 complete codebases generated by 16 LLMs from OpenAI, Anthropic, Google, and others to determine how often they produce vulnerabilities, what types of vulnerabilities are most common, how security outcome differs by framework, and more. The index scores models from 0 to 100, with a higher score indicating fewer vulnerabilities. 

A quick look shows the following models in the top five:

  • Claude Sonnet 5 — 80.4
  • Claude Fable 5 — 76.4
  • GPT 5.3 Codex — 75.3
  • Claude Opus 4.8 — 74.5
  • GPT 5.1 — 71.6

with these models trailing far behind:

  • Gemini 2.5 Flash — 39.4
  • Qwen3 Coder — 32.7
  • GPT 5 Mini — 21.6

The scores don’t necessarily add up to a clear winner, however, as framework context affects model performance. In other words, “AI-generated coding risk is not random," the announcement states. "It's predictable by model and framework, giving security leaders the data to safely scale AI-assisted development.”

Additionally, “no single AI model consistently produces the most secure code." For example:

  • GPT-5.1 leads in Java Enterprise API. 
  • Claude Sonnet 4.5 leads in Java Spring. 
  • Claude Opus 4.8 leads in Python Django. 
  • GPT-5.5 leads in C# (.NET).
  • Claude Fable 5 leads in C.

In terms of specific security risks, most of the vulnerabilities reflect things that models fail to do rather than actively generate, with the most common involving logging failures, injection vulnerabilities, insecure design, and broken access control. 

You can learn more and compare model performance against various frameworks through the SCW AI Index.

Add ADMIN IT Infrastructure & Operations on Google