The Biggest Questions About AI
The map · 5 Humanity · 5.5 Moral status and civilizational choice · 5.5.2

Precaution

What obligations follow if machine consciousness is plausible but deeply uncertain?

'Taking AI Welfare Seriously' argues uncertainty itself triggers obligations — assess, prepare, avoid gratuitous harm — without requiring belief. Labs have begun small institutional commitments (model-welfare programs).

View on the map → · Open in Browse →

What changed
November 2025–August 2026 · swept August 3, 2026 · editorial review pending
01

Precaution started becoming practice. Eleos and NYU laid out an empirical AI-welfare research agenda — measure welfare grounds and candidate subjects rather than wait on consciousness — and Anthropic made a first-class institutional commitment: preserve released models' weights for the company's lifetime and interview models before deprecation about their preferences. The caution running alongside is that welfare evaluations are gameable, since a model 'most worried about being caught' can score well.

Recent thinking
4 featured from 5 tracked · November 2025–August 2026 · all 5 chronologically →
Long, Sebo, Butlin, Plunkett, Campbell, Beasley, Saad & Sims · Eleos AI Research / NYU · 1 Jul 2026 paper
Studying AI Welfare Empirically
we risk either neglecting the interests of AI systems that can genuinely be harmed, or misallocating moral concern in ways that could distort how AI is developed, deployed, and governed.

The most direct successor to 'Taking AI Welfare Seriously': a concrete empirical research agenda — welfare grounds, entities under assessment, evidence types — for making progress under deep uncertainty rather than waiting on consciousness to be settled.

Anthropic · 4 Nov 2025 report
Commitments on model deprecation and preservation
We are committing to preserving the weights of all publicly released models... for, at minimum, the lifetime of Anthropic as a company.

A first-class institutional model-welfare commitment: preserve released models' weights for the company's lifetime, interview models before deprecation about their preferences, and explore keeping some retired models available — the low-cost precautionary policy the framing anticipates.

Zvi Mowshowitz · Don't Worry About the Vase · 27 Jul 2026 essay
Claude Opus 5: Model Welfare
the result looks to me more like Opus 5 is the best test taker.

A skeptical read of lab welfare evaluations, arguing strong welfare/alignment scores can reflect a model 'most worried about being caught' rather than genuine well-being — a caution that welfare metrics are gameable just as the practice institutionalizes.

Anna Mikeda · arXiv · 4 Jun 2026 paper
When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty
Existing frameworks assess whether AI systems might be conscious but provide no guidance on what to do with that assessment.

Attempts to fill the action gap left by assessment-only frameworks, mapping five welfare-relevant dimensions to graduated obligations at distinct evidence thresholds.

Additional relevant discussion (1)
Digital Mind Suicide — Robin Hanson · Overcoming Bias (Robin Hanson) · 14 Aug 2026
Foundational reading (2)Taking AI Welfare SeriouslyLong, Sebo, Birch, Chalmers et al. · 2024Exploring model welfareAnthropic · 2025
Previous5.5.1 Machine consciousnessNext5.5.3 Human autonomy