Event · 2026
Reliable autonomy
Dangerous capability evals
Bio and cyber risk testing became a standard part of frontier releases rather than a goodwill gesture.
Frontier labs now publish frameworks that connect measured dangerous capability levels to required safeguards. This is not yet a safety certificate: the evaluations are still evolving, and results depend on test scenarios and the quality of external review.
Sources
OpenAI · Preparedness FrameworkOpen primary source Anthropic · Responsible Scaling PolicyOpen primary source