Research and benchmarks
Stanford's AI Index reports 362 AI incidents in 2025 as model transparency fell from 58 to 40
The 2026 AI Index records 362 documented AI incidents in 2025 against 233 in 2024, and an average Foundation Model Transparency Index score of 40 after 58 the year before. It also finds capability benchmarks reported far more often than safety ones.

Stanford University's Institute for Human-Centered Artificial Intelligence published the 2026 edition of its AI Index in April 2026. Two of its measurements move in opposite directions. The count of documented AI incidents rose to 362 in 2025, up from 233 in 2024. The average score on the Foundation Model Transparency Index, which grades developers on what they disclose about their models, fell to 40 out of 100 for 2025, after climbing from 37 in 2023 to 58 in 2024.
The incident figure is not Stanford's own count. The Index's responsible AI chapter attributes it to the AI Incident Database, which is operated by the Responsible AI Collaborative and indexes documented harms and near harms from deployed systems. The database invites public submissions, indexes them and makes them discoverable, and describes itself as working like similar databases in aviation and computer security. It is not the only counter, and the counters do not agree. The OECD runs a separate AI Incidents and Hazards Monitor, described on its own site as an automated media discourse monitor and still labelled beta, which currently displays about 17,000 incidents and hazards rather than hundreds, because it is scanning news coverage rather than curating verified individual cases. A reading of the Index by Mary Bennett and Rob Robinson of HaystackID, published on 24 April 2026, cited a third figure again, 435 reports on the OECD monitor in January 2026 on a six month moving average of 326. Any claim about how many AI incidents occurred in 2025 is a claim about a specific methodology, not a census.
The transparency measure is a different sort of instrument. The Foundation Model Transparency Index is a scoring exercise from Stanford's Center for Research on Foundation Models, an interdisciplinary initiative inside the same institute that publishes the Index, and one that describes itself as spanning more than ten Stanford departments. That makes the transparency finding an assessment by the report's own house rather than an independent audit, and the Index does not hide it. The responsible AI chapter states the movement plainly, from 37 to 58 between 2023 and 2024, then down to 40 in 2025, and says major gaps persist in disclosure around training data, compute and post deployment impact. The HaystackID analysis reports that of 95 notable models released in 2025, four arrived with open source training code and 80 arrived without it.
The third finding is the one that ties the other two together. The Index says almost all leading frontier model developers report results on capability benchmarks such as MMLU and SWE-bench, while reporting on responsible AI benchmarks remains sparse. Capability is measured in public and compared across firms. Safety is measured unevenly, by whoever chooses to publish. The same chapter reports hallucination rates across 26 leading models ranging from 22 per cent to 94 per cent.
The public has noticed something, though not necessarily this. The Index's public opinion chapter records a 50 point gap on the workplace: 73 per cent of experts expect AI to have a positive effect on how people do their jobs, against 23 per cent of the general public. The gaps on other questions are similar in shape. On economic impact it is 69 per cent of experts against 21 per cent of the public. On medical care it is 84 against 44. Meanwhile global sentiment moved slightly positive, with the share saying AI products offer more benefits than drawbacks rising from 55 per cent in 2024 to 59 per cent in 2025, even as the share saying such products make them nervous reached 52 per cent.
Governments are engaging more. The policy chapter records AI related witnesses at United States congressional hearings rising from 5 in 2017 to 102 in 2025, with industry witnesses growing from 13 per cent of testimony to 37 per cent, making industry the largest witness group, while academic witnesses fell to 15 per cent.
What the Index does not establish is a causal link. Rising incident counts and falling disclosure scores are two separate series measured by two separate methods, and the report does not argue that one produces the other. Nor does it publish a denominator: incidents can rise because deployment rises, because reporting improves, or because harm rises, and the 362 figure alone cannot separate those. The reason for the transparency fall is also unsettled. Competitive secrecy, legal caution about training data disclosure, and a shift in which models get scored would all produce the same drop. The Index reports the pattern. It does not resolve the cause.
Sources
Every factual claim above rests on the 10 published sources below. They are listed so you can check the reporting rather than take it on trust.
- Stanford HAI2026 AI Index Report
- Stanford HAIInside the AI Index: 12 takeaways from the 2026 report
- Stanford HAI2026 AI Index Report, Chapter 3: Responsible AI
- Stanford HAI2026 AI Index Report, Chapter 9: Public Opinion
- Stanford HAI2026 AI Index Report, Chapter 8: Policy and Governance
- Responsible AI CollaborativeAI Incident Database
- OECDAI Incidents and Hazards Monitor
- Stanford CRFMCenter for Research on Foundation Models
- JD Supra (HaystackID)Stanford's 2026 AI Index highlights rapid growth and widening governance gaps
- ForbesStanford's AI Index exposes K-12's crisis


