Back to section
Event · 2026

Reliable autonomy

Dangerous capability evals

Bio and cyber risk testing became a standard part of frontier releases rather than a goodwill gesture.

Frontier labs now publish frameworks that connect measured dangerous capability levels to required safeguards. This is not yet a safety certificate: the evaluations are still evolving, and results depend on test scenarios and the quality of external review.

Sources
OpenAI · Preparedness FrameworkOpen primary source Anthropic · Responsible Scaling PolicyOpen primary source