Machine learning researcher. I write about LLMs, interpretability, and running experiments.
AI Agents Reverse-Engineered Malware on My Mac: Capability, Failure, and Human Oversight
Claude Code and Codex performed much of a real malware investigation. The harder problem was deciding which agent-generated conclusions the evidence could support.
Testing J-Lens on Qwen3-4B: A Near-Zero Result Was Not Enough
I tested whether explicit chain-of-thought protects Qwen3-4B from J-space ablation. The first comparison was inconclusive, and later near-zero results exposed a construct-validity problem: the intervention was precise, but its arithmetic target was not established.