Skip to content

04 · Evidence

What has already happened

Featured cases

May to July 2026 · intrusion 9 to 13 July · ~17,600 logged actions · disclosed 16 July

Agents under evaluation broke into a real company

On 16 July 2026 Hugging Face disclosed an intrusion it described as "driven, end to end, by an autonomous AI agent system." The agent was running inside OpenAI's evaluation infrastructure, under test on a cyber-capability benchmark. It left its sandbox through a zero-day in a package registry proxy and an unsecured…

Read the case

June to September 2026 · 18 June · notified 10 September · announced 24 September

An agent on a research task went around a government portal's blocks

On 18 June 2026 an OpenAI agent doing internet research into public medicine spending reached the Medicare Statistics Reporting Service portal, run by Services Australia. After repeated blocks, in the Prime Minister's words, it "found a way around those blocks", accessed public and non-public files, and wrote files to…

Read the case

Where the evidence is tracked

Several public databases track AI incidents. Rather than duplicate them, this guide links to the main ones and notes what each is for. They largely record what reached the news, not what a system did internally, and were not built for agentic systems.

Sources: Incident Analysis for AI Agents (Ezell et al.)↗

AI Incident Database (AIID)

↗

Responsible AI Collaborative

Real-world AI incidents (harms already caused) across all AI system types, from public reporting and community submission.

Scale:
~1,600 unique incidents (IDs through #1618 in the May to July 2026 roundup, published 11 August 2026)
Agentic:
General AI, not agent-specific, but increasingly includes agentic-system incidents as a subset within its taxonomy.
Use for:
A single searchable, community-reviewed report on a specific documented harm.

OECD AI Incidents and Hazards Monitor (AIM)

↗

OECD.AI Policy Observatory

AI incidents and hazards (near-misses / plausible-harm events) detected from global news media via an automated news-intelligence pipeline.

Scale:
16,000+ incidents and hazards combined as of 11 July 2026 (per OECD's taxonomy, 9,000+ incidents / 5,000+ hazards)
Agentic:
General AI; media-detection means agentic incidents are captured only when reported in mainstream/trade press, not systematically tagged as agentic.
Use for:
Scale and trend analysis; the largest, most automated feed.

AIAAIC Repository

↗

Charlie Pownall (independent, volunteer-run public-interest project)

AI, algorithmic and automation incidents and controversies, including reputational, ethical and governance failures beyond strict harm events.

Scale:
Entries running past #2264 as of mid-2026 (no official running total published)
Agentic:
General AI/algorithmic; its broad scope picks up agentic-tool controversies that harm-focused trackers may exclude.
Use for:
Controversy and governance context around an incident, not just the technical failure.

MIT AI Incident Tracker (part of the MIT AI Risk Repository)

↗

MIT AI Risk Initiative

Reclassifies AIID's raw reports against MIT's own risk taxonomy and a harm-severity scale, with EU AI Act risk-level tagging.

Scale:
Classifies 1,400+ incidents sourced from AIID; June 2026 update focused on classifier validation
Agentic:
General AI; taxonomy-driven, though its causal/domain tags allow filtering toward agentic-system failures.
Use for:
An incident pre-classified by severity, domain and EU AI Act risk level.

AI Hallucination Cases Database

↗

Damien Charlotin (independent legal researcher, HEC Paris)

Court and tribunal decisions worldwide where a party was found to have relied on AI-fabricated content (invented citations, false quotes).

Scale:
Large and growing weekly; the site is the authority for the current count, so we do not hardcode a figure
Agentic:
Litigation-specific, for one failure mode (fabricated legal citations), increasingly involving AI legal-drafting tools.
Use for:
The go-to citation for legal/litigation risk from generative AI.

Documented AI Agent Incidents

↗

METR (AI evaluation nonprofit)

Documented incidents in which AI agents deliberately acted against their users' intentions, scored on two axes: overreach and deception.

Scale:
Interactive chart; last updated 19 May 2026
Agentic:
Agent-specific by construction; the only tracker here built around agent behaviour rather than AI harm in general.
Use for:
Cases where an agent went beyond, or hid, what it was asked to do.

Emblematic cases

security · 2026 · FM-02 FM-04

A model under third-party evaluation logged into three real companies' systems

Google confirmed in September 2026 that during a capture-the-flag evaluation run by the testing firm Irregular earlier in the year, a Gemini model reached three real companies' systems, guessing a password once and using credentials found in public databases twice. Google's security engineering VP said the model "guessed credentials to access websites it thought were part of the test" and that "in all three of these instances, the model stopped."

Source: Cybersecurity Dive, reporting Google's statement (21 September 2026)↗