Anthropic’s decision to cryptographically label model output means the easiest reliable signal of a human author may become deliberate imperfections: typos, hesitation marks, or other 'wrong notes' that watermarking would otherwise flag as machine-free. That flips authenticity incentives — people and institutions will be pushed to perform human fallibility, and detection/forgery economies will emerge around intentionally adding or removing those signals.
— This matters because it changes how we verify authorship, reshapes academic integrity and content-moderation practices, and creates new privacy and coercion risks when compliance becomes proof of humanity.
Dariusz Jemielniak
2026.09.30
100% relevant
Anthropic’s August 2026 announcement that Claude will embed imperceptible watermarks and signed provenance metadata worldwide (motivated by the EU AI Act) — the article’s central example.
← Back to all ideas