Silent failures
The core problem of long tasks: failing without signalling failure — confidently returning a wrong result.
Before trusting AI with an operating theatre, a power plant or a company's books, you must be able to prove it will not fail silently. No such method exists today.
Autonomy is limited by guarantees, not intelligence. Without formal behavioural verification, AI will not enter critical loops.
The core problem of long tasks: failing without signalling failure — confidently returning a wrong result.
Bio and cyber risk testing became a standard part of frontier releases rather than a goodwill gesture.
The attempt to see inside the model rather than judge it by its output. So far it works at toy scale.
Regulation turns verifiability from a virtue into a shipping requirement.
A formal guarantee that a system stays inside drawn boundaries — the equivalent of aircraft certification.