Vinod Anbalagan

Research

When a model succeeds, has it learned the structure, or the surface cues?

One question runs through the work: what changes in a model's internal representations when you constrain how it learns. I answer it with probes and causal interventions rather than benchmark scores.

The programme · one question, three knobs, one toolkit

later

Learning rule

Backprop versus local and biologically inspired alternatives.

later

Architecture

Symmetry and equivariance built into the network.

now

Objective

Predictive world models: physical state or surface statistics?

Instruments held constant: linear and nonlinear probing · activation patching · dynamical-systems analysis · matched baselines and registered negative results

In progressPre-registeredWorld models · interpretability

Reverie

Does a world model represent what it cannot see?

Abstract

World models trained on video are often said to learn physics. Reverie tests one narrow version of that claim. A recurrent state-space model (RSSM) is trained on simulated video of balls bouncing and colliding in a box, with no physics in the loss. When a ball passes behind an occluder, we ask two questions: is its position still decodable from the model's internal state, and is that representation causally used to predict where the ball reappears?

Decodability is measured with probes of increasing capacity, reported as selectivity against a randomly initialised encoder. Causal use is tested by patching the recurrent state with a counterfactual trajectory, against a matched-norm random-direction control. The central control is hidden collisions: collisions behind the occluder that no extrapolation can predict, which separate learned dynamics from straight-line continuation. Results will be validated on CLEVRER.

RSSM · linear and MLP probes · activation patching · CLEVRER · JAX

The animation on the home page illustrates the hidden-collision control. It is a sketch of the question, not model output.

MRL 2026First authorMultilingual evaluation

CulturalRiddles

A multicultural benchmark of riddles across 51 languages.

To be presented at MRL 2026 · being submitted to TACL · arXiv preprint coming soon

Riddles are a setting where surface fluency and real understanding come apart sharply. A model can produce perfect Tamil and still miss what a Tamil riddle is actually about. CulturalRiddles collects around 6,100 riddles written by native speakers from 61+ language communities, and evaluates how well language models solve them in the cultural context they come from.

Paper, dataset and code will be linked here on publication.

The Lab is where the smaller experiments live: eight measured, interactive studies that build the toolkit Reverie depends on. Follow along on The Meta Gradient.