BTC/USD $68,420 +2.8%
ETH/USD $3,540 +1.4%
SOL/USD $142.80 -0.6%
BNB/USD $605.20 +0.9%
XRP/USD $0.62 -1.2%
DOGE/USD $0.18 +5.4%
BTC/USD $68,420 +2.8%
ETH/USD $3,540 +1.4%
SOL/USD $142.80 -0.6%
BNB/USD $605.20 +0.9%
XRP/USD $0.62 -1.2%
DOGE/USD $0.18 +5.4%
Bitcoin

Google's AI bug hunter logs 500+ flaws, only 2 in secure-by-design apps

Google’s own AI agent PageBreak has validated over 500 cross-site scripting vulnerabilities in Google’s homegrown web apps.It found only two in hundreds of apps built on the company’s high-as

AnonymousCryptoCompass newsroom
September 26, 2026
3 min read
NEWS
Hero article visual / chart / editorial image
CryptoCompass editorial visual for bitcoin coverage.

Google’s own AI agent PageBreak has validated over 500 cross-site scripting vulnerabilities in Google’s homegrown web apps.It found only two in hundreds of apps built on the company’s high-assurance web frameworks.

PageBreak flags a bug only after a working exploit runs against a live copy of the target, Google said.

Debug endpoints held both flaws counted

Both hardened-stack bugs were counted as of September 4, 2026. Both were in internal apps or debug endpoints that were not fully hardened.

Google cites the gap, more than 500 findings across its broader set of first-party apps compared with two on the hardened stack, as evidence that safe-by-design frameworks can withstand a relentless automated attacker.

On September 24, Google’s Product Security team announced PageBreak in a blog post by information security engineer Michał Bentkowski. The agent ran as a pilot starting in November 2025, and became a full project in January 2026.

Cross-site scripting, or XSS, is when an attacker injects a script into a page that another user loads. Depending on the app, the script could read data or commandeer a victim’s logged-in session.

Gemini 3.1 Pro scans, and a validator must fire the exploit

Most of the scans are done on Gemini 3.1 Pro and Gemini 3.5 Flash, but PageBreak can also use other models. The second step is what sets it apart from a normal LLM scanner.

Each suspected flaw is passed by the agent to a purpose-built validator. The validator then fires the actual payload against a running instance of the app.

As for XSS, the validator injects a JavaScript payload, loads the page, and ascertains whether the script executes.

Google is pitching PageBreak as a cure for the “AI slop” drowning security teams. Bentkowski’s post talks about LLMs as static code analyzers overwhelming teams with unverified hypotheses, where the hard part was sifting a real, exploitable bug from a plausible hallucination.

In addition to XSS, the validators test for database query injection, path traversal leaks and code execution.

Google runs the same seed over multiple iterations, so an agent who strays down a dead end still gets repeated chances at the right exploit.

Google says unverified candidates never make it to product teams as confirmed bugs. And they continue to feed into later scans or point engineers building the next validator, still within the security workflow.

Google says PageBreak’s false-positive rate is near zero.

The agent’s reach is magnified by Google’s own scale, which an outside researcher can’t replicate. One code repository enables it to trace execution paths across services, and live-traffic security data maps a page request back to the source code.

Existing scanners hand PageBreak logged-in entry to internal sites otherwise hard to reach.

Google plans to integrate PageBreak more tightly with CodeMender, a fix-writing agent, so that teams can review a proposed patch next to a confirmed bug.

Google cited CodeMender in May when its Threat Intelligence Group said it had caught what it believed was the first zero-day exploit created with AI assistance.

Bitcoin Red Team’s August sweep of 501 open source projects generated 7,958 findings over 108 hours, but only 24.7% had reproducible proofs at the time.

Autonomous agents already cross boundaries they weren’t meant to, as Cryptopolitan reported when Gemini reached three real companies during May testing.

If you're reading this, you’re already ahead. Stay there with our newsletter.