AI safety is a complex, evolving field with no universal standard — terms like alignment, red-teaming and risk thresholds mean different things to different developers. Grading frontier labs on governance without accounting for the technical nuance behind safeguards, evaluations and deployment contexts oversimplifies a deeply layered challenge. A glossary of competing definitions isn't a failure of the industry — it reflects how genuinely hard this problem is.
Nine of the biggest AI labs just got graded on safety and governance — and the results are embarrassing. Anthropic led the pack with a C+, three labs outright failed, and nobody earned anything close to a passing mark. Frontier AI buyers can't just take model cards at face value anymore; they need hard answers on risk thresholds, external testing access and incident reporting before deploying these systems.
© 2026 Improve the News Foundation.
All rights reserved.
Version 7.4.1