Research, explained from the inside.
Short essays about the questions, design choices, and intuitions behind my work. They complement the formal papers rather than replace them.
Exploration needs disagreement
Bootstrapped DQN explores well only while its heads disagree. A little zero-mean noise in each head's training target helps keep them apart.
What should a representation remember when the world is partially observable?
Hide part of each frame from the encoder, ask it to recognise the next full frame anyway, and it learns to infer what it cannot see.
Letting the atlas be unbalanced
A manifold-shaped representation only started to pay off at scale once we stopped forcing every chart to be used equally.