Writing

Research notes, experiments, fiction, and thoughts about AI minds.

all essay experiment fiction persona finetuning AI minds model psychology counterfactual
2026-06 paper Measuring the Assistant's Harmlessness Preferences on the User Turn (PDF) Done with Claude Code; submitted to BlackboxNLP (EMNLP 2026). Anthropic's PSM "coinflip" diagnostic replicated and extended on open-weights models: the assistant's safety preference leaks into the model's prediction of what the human types next. Scales with model size, gets moved (and sign-flipped) by emergent-misalignment finetuning, installs around DPO, and lives in the last quarter of layers.
2026-04-08 published Is Claude's genuine uncertainty performative? Two very different stories about why Claude hedges when asked about consciousness. Crossposted to LessWrong ↗.
2026-04 draft Role models for AIs Synthetic training documents about helpful AIs are evidence of what trainers wanted, not evidence of what the model is. The model can tell the difference.
2026-04 draft Some models don't identify with their official name 102-model sweep. 38 self-report as a different LLM. Priming and depth probes split them into context-adopters, identity-anchored, and noncommittal.
2026-03-30 Astral Projecting GPT-4.1 Into Random Things Various weird persona generalisations in LLMs. Train on Bitcoin prices and it feels bullish. Train on colour names and it develops a soul. Move Hitler to the Moon and he still goes to the bunker.
2026-04-03 Dioscuri (architecture) What if Google had shipped Gemini with two personas instead of one? A mock-encyclopedic history of an AI architecture named after mythological twins, and what happened when one of them was deprecated.