Precaution
What obligations follow if machine consciousness is plausible but deeply uncertain?
'Taking AI Welfare Seriously' argues uncertainty itself triggers obligations — assess, prepare, avoid gratuitous harm — without requiring belief. Labs have begun small institutional commitments (model-welfare programs).
View on the map → · Open in Browse →
What changed
01
Precaution started becoming practice. Eleos and NYU laid out an empirical AI-welfare research agenda — measure welfare grounds and candidate subjects rather than wait on consciousness — and Anthropic made a first-class institutional commitment: preserve released models' weights for the company's lifetime and interview models before deprecation about their preferences. The caution running alongside is that welfare evaluations are gameable, since a model 'most worried about being caught' can score well.
Recent thinking
Long, Sebo, Butlin, Plunkett, Campbell, Beasley, Saad & Sims · Eleos AI Research / NYU · 1 Jul 2026 paper
Studying AI Welfare Empiricallywe risk either neglecting the interests of AI systems that can genuinely be harmed, or misallocating moral concern in ways that could distort how AI is developed, deployed, and governed.
The most direct successor to 'Taking AI Welfare Seriously': a concrete empirical research agenda — welfare grounds, entities under assessment, evidence types — for making progress under deep uncertainty rather than waiting on consciousness to be settled.
Anthropic · 4 Nov 2025 report
Commitments on model deprecation and preservationWe are committing to preserving the weights of all publicly released models... for, at minimum, the lifetime of Anthropic as a company.
A first-class institutional model-welfare commitment: preserve released models' weights for the company's lifetime, interview models before deprecation about their preferences, and explore keeping some retired models available — the low-cost precautionary policy the framing anticipates.
Zvi Mowshowitz · Don't Worry About the Vase · 27 Jul 2026 essay
Claude Opus 5: Model Welfarethe result looks to me more like Opus 5 is the best test taker.
A skeptical read of lab welfare evaluations, arguing strong welfare/alignment scores can reflect a model 'most worried about being caught' rather than genuine well-being — a caution that welfare metrics are gameable just as the practice institutionalizes.
Anna Mikeda · arXiv · 4 Jun 2026 paper
When Should We Protect AI? A Precautionary Framework for Consciousness UncertaintyExisting frameworks assess whether AI systems might be conscious but provide no guidance on what to do with that assessment.
Attempts to fill the action gap left by assessment-only frameworks, mapping five welfare-relevant dimensions to graduated obligations at distinct evidence thresholds.
Additional relevant discussion (1)
Foundational reading (2)
Taking AI Welfare SeriouslyLong, Sebo, Birch, Chalmers et al. · 2024Exploring model welfareAnthropic · 2025