AGI
Will AI Safety Keep Up with AI Capability?
Don't just read what happened. See what could happen next.
Prediction
Safety tooling and evals improve, but capability cadence still outruns comprehensive assurance for frontier systems.
- Confidence
- 40%
- Horizon
- 12–36 months
- Impact
- High
- Direction
- Stable
Prediction changed +5 points in 21 days
Short answer
Labs ship safer defaults, red-teaming, and monitoring — yet frontier models keep opening new misuse and reliability surfaces. Past-year progress is real on evaluations; coverage of agentic, multi-step harm remains incomplete.
Why this question matters
A persistent safety gap shapes regulation, insurance, enterprise adoption ceilings, and public trust.
What's happening now?
Frontier eval suites expand
More structured testing for misuse, autonomy, and deception-like behaviors.
Signal · moderate
Transparency declining at the frontier
Less disclosure of training data and methods as systems concentrate in industry.
Signal · strong
Regulation drafts and safety institutes
Governments stand up testing regimes that lag release cycles.
Signal · moderate
Agent risk surfaces multiply
Tool use and long-horizon tasks create failure modes evals only partially cover.
Signal · strong
What could happen next?
Scenario A
Assurance catches up
Standardized evals and third-party audits become prerequisites for frontier deploy.
Scenario B
Chronic lag
Capability jumps repeatedly reset the safety backlog.
Scenario C
Hard stop after incident
A severe failure forces pause norms or licensing that slow releases.
Key companies / entities
- Anthropic
- OpenAI
- Google DeepMind
- NIST
- EU AI Office
Evidence
- company
Lab safety and system cards
- research
AI Index safety and transparency findings
- policy
National AI safety institute announcements
Prediction history
- Sep 22, 202640%
- Sep 15, 202638%
- Sep 8, 202637%
- Sep 1, 202635%