Early mechanistic‑interpretability methods (like many‑to‑many feature maps) that worked on small, toy neural nets do not reliably scale to modern large language models. Reported crisp features become vague and inconsistent across teams, and practical uses (e.g., lie detection) underperform simpler methods.
— Policymakers and safety advocates should not assume that technical interpretability will soon provide reliable audits or fixes for advanced models; regulation and procurement must account for persistent epistemic limits.
Scott Alexander
2026.09.08
100% relevant
Scott Alexander's review of the interpretability field: 2023 many‑to‑many mapping excitement, subsequent 2024–25 disillusionment, and concrete failures to use those maps for lie detection or robust feature definitions.
← Back to all ideas