Machine learning researcher. I write about LLMs, interpretability, and running experiments.
Can Chain of Thought Replace a Model's Internal Workspace?
I tested whether Anthropic’s finding—that written reasoning protects a model from J-space interventions—also holds for Qwen3-4B. I did not reproduce the math result, and follow-up experiments showed why that null result remains inconclusive.
From a Compromised Mac to an Auditable Malware Analysis
A real compromise became a case study in supervising capable agents: bounded autonomy, auditable evidence, and claims designed to fail closed.