I write about thinking, and end up doing interp, safety-adjacent work, and post-training evals. A lot of my source material is from other backgrounds and registers. I like to play with the mechanics of reasoning.
Currently, I'm focused on the gap between what models say they think and what they do.
Part of the question ends up being about the meaning behind the thinking, that is to say, epistemic principles of machine thinking.