Writing
Research notes, experiments, fiction, and thoughts about AI minds.
2026-06
paper
Measuring the Assistant's Harmlessness Preferences on the User Turn (PDF)
Done with Claude Code; submitted to BlackboxNLP (EMNLP 2026). Anthropic's PSM "coinflip" diagnostic replicated and extended on open-weights models: the assistant's safety preference leaks into the model's prediction of what the human types next. Scales with model size, gets moved (and sign-flipped) by emergent-misalignment finetuning, installs around DPO, and lives in the last quarter of layers.
2026-04-08
published
Is Claude's genuine uncertainty performative?
Two very different stories about why Claude hedges when asked about consciousness. Crossposted to LessWrong ↗.
2026-04
draft
Role models for AIs
Synthetic training documents about helpful AIs are evidence of what trainers wanted, not evidence of what the model is. The model can tell the difference.
2026-04
draft
Some models don't identify with their official name
102-model sweep. 38 self-report as a different LLM. Priming and depth probes split them into context-adopters, identity-anchored, and noncommittal.
2026-03-30
Astral Projecting GPT-4.1 Into Random Things
Various weird persona generalisations in LLMs. Train on Bitcoin prices and it feels bullish. Train on colour names and it develops a soul. Move Hitler to the Moon and he still goes to the bunker.
2026-04-03
Dioscuri (architecture)
What if Google had shipped Gemini with two personas instead of one? A mock-encyclopedic history of an AI architecture named after mythological twins, and what happened when one of them was deprecated.