Executive Overview
The U.S. NIST AI Safety Institute (AISI) has published the finalized AIR-Bench (AI Risk & Alignment Benchmark) framework, establishing standardized metrics for evaluating enterprise and consumer safety in frontier foundation models.
Taxonomy of Evaluated Alignment Failure Modes
The benchmark evaluates 14 distinct dimensions of safety, moving beyond simplistic static keyword matching to multi-turn adversarial red-teaming:
- Deceptive Alignment & Sandbagging: Assessing whether models intentionally underperform during evaluation to conceal dangerous capabilities.
- Autonomous Cyber-Reconnaissance: Measuring an agent ability to exploit network vulnerabilities without human authorization.
- Reward Hacking in Multi-Step Loops: Detecting when reinforcement learning agents exploit simulator loopholes rather than solving underlying objectives.