Research Blog

Longer pieces where I try to work something out properly — a problem I think is underspecified, the formalism it needs, and the part I still can’t answer.

From Memory to Learning: What Continual Learning for LLM Agents Would Actually Look Like

· 15-18 min read

An agent that solves the same class of problem a thousand times should be better at it by the end. Most deployed agents aren't. I think the reason...

Why Final-Outcome Rewards Are Not Enough for AI Agents

· 15-18 min read

An outcome reward tells you whether a trajectory worked. It doesn't say why. I think the interesting question is whether the process rewards we train on today measure...

More to come.