I use independent research to explore ideas that feel important before their product shape is obvious. Lately, I have been asking how multimodal systems should represent space, time, and long context, and what architectures might move beyond today's dominant recipes.
I publish code, artifacts, experimental details, and the paths that did not work. The aim is to leave behind something another person can inspect, reproduce, or build upon.
One question I keep coming back to is whether the architectural choices inherited from language are right for vision. This project tests rotation as an inductive bias for contrastive vision-language encoders. The experiments study when 2D RoPE helps, when it does not, and whether the same idea can extend context without retraining.
More on the videos page.