---
title: "Capability Without Security: Measuring the Functionality-Security Gap in AI-Generated Code"
date: "2026-07-08T18:25:32+00:00"
url: "https://staging.checkmarx.com/capability-without-security-measuring-functionality-security-gap-ai-generated-code/"
description: "New research reveals the functionality-security gap in AI-generated code, introducing a secure code generation benchmark to evaluate how leading AI coding assistants balance functionality with software security."
---

# Capability Without Security: Measuring the Functionality-Security Gap in AI-Generated Code

## Capability Isn’t Enough. Security Matters.

### Thank you!

 ![TY Form Visuals](https://staging.checkmarx.com/wp-content/uploads/2025/12/TY-Form-Visuals.svg)

Research Report

# The Weather Report: Capability Without Security

 ![Capability Without Security Chart](https://staging.checkmarx.com/wp-content/uploads/2026/07/Capability-Without-Security-Chart.webp)

## Measuring the functionality-security gap in AI-generated code

Google says roughly 75% of its new code is now AI-generated and engineer-reviewed. As AI-generated code becomes standard, Checkmarx commissioned The Weather Report to independently examine how secure that code is in practice.

The report tested this by rerunning two published benchmarks – CyberSecEval and SusVibes – on Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, and Gemini 3.5 Flash. The result: models write working code 83% to 95% of the time, but only 24% to 36% of that code is secure.

## Key Highlights:

 Frontier models have improved, but their code still needs security review.

 Better coding does not mean safer code.

 Generic security prompts and max-effort settings add little.

 Threat modeling plus review boosts secure output, but costs tokens.

 The industry needs a maintained benchmark, not another one-off test.

Download the report for the complete benchmark results.

*This study was initiated and funded by Checkmarx. Per the report’s Acknowledgements section, The Weather Report Inc. — an independent research organization — retained full control over benchmark and model selection, experimental design, evaluation, analysis, and conclusions.*

 ## Market Technology Leadership

40%

of Fortune 100

1800+

Customers in 70 countries

75+

Languages &amp; 100+ frameworks

7X

Leader at Gartner® Magic Quadrant™ for Application Security Testing

## Industry Recognition

 ![SAST Forrester Wave Leader 2025 Award logo](https://staging.checkmarx.com/wp-content/uploads/2025/09/FORRESTER-2025-Checkmarx-Badge.png)

 ![gartner_checkmarx](https://staging.checkmarx.com/wp-content/uploads/2025/10/gartner_checkmarx.webp)

 ![Latio Application Security Testing Leader 2026 badge. The circular badge features a blue center with black text 'APPLICATION SECURITY TESTING LEADER' and 'Latio' in script at the top. A light blue ribbon at the bottom displays '2026'.](https://staging.checkmarx.com/wp-content/uploads/2026/02/Testing-Leader-1.png)

 ![Shortlist Badge](https://staging.checkmarx.com/wp-content/uploads/2026/02/Shortlist-Badge.webp)
