Vinod Anbalagan

Machine learning engineer and independent researcher working on world models and representation learning: what a model actually represents, and whether it has learned the structure of a problem or only the surface cues that correlate with it.

Research direction

My current work probes predictive world models, using probing and causal intervention to test whether they keep physical state they can no longer see, or only predictive surface statistics (Reverie). The broader programme asks what changes in a model’s representations when you constrain how it learns, through its update rule, its architecture or its objective.

Alongside this: first author on CulturalRiddles, a 51-language benchmark with Cohere Labs, and 2nd place in the Adaption Labs AutoScientist Challenge. Background in electronic and computer engineering, and a decade of forecasting in real operations.

Selected work

Reverie: Does a World Model Represent What It Cannot See?

2026 to present

Independent research · Pre-registered

Probing and activation patching of an RSSM world model trained on simulated physics video with no physics in the loss. Tests whether an occluded object’s position stays decodable from the latent state and whether that representation is causally used, with hidden collisions as the central control and CLEVRER for validation.

CulturalRiddles: Multicultural Riddles Benchmark

Apr 2026 to present

Cohere Labs Open Science Community · First author

Living cultural-reasoning benchmark: 51 languages, 61+ communities, about 6,100 riddles, 100+ contributors. To be presented at MRL 2026, being submitted to TACL, arXiv preprint coming soon.

Adaption Labs AutoScientist Challenge

2026

2nd place, Data & Visualization track · $1,000 prize

Adaption Charts P2: synthetic multimodal dataset for chart QA and chart-to-table extraction, built so difficulty comes from reading the chart (truncated axes, missing labels, varied visual styles) rather than from arithmetic. Paired Llama-3.2-3B LoRA adapter for table-grounded QA.

DocuNative: Offline Multilingual Document QA

Mar 2026

Cohere Expedition Hackathon · Pipeline lead, 7-person team, one week

Fully offline document QA for newcomers and migrants; nothing leaves the user’s machine. PDF extraction, BGE-M3 embeddings, ChromaDB retrieval, Aya generation and mDeBERTa hallucination checks. 9,000+ automated evaluations across Chinese, Hindi and Polish showed cross-lingual embedding quality, not the LLM, was the main retrieval bottleneck.

Tamil Agricultural Advisory Dataset

2026

Adaption Labs Uncharted Data Challenge · Grade A (9.4/10)

187 curated Q&A records across 20 categories and 48 crops, built from public-sector sources over ten iterative submissions. Honorary recognition from Sara Hooker. Follow-up: an 11-language Indian agricultural dataset.

Plus six public datasets and a LoRA adapter on Hugging Face (580+ downloads).

Experience

First Author, Multicultural Riddles Benchmark

Apr 2026 to present

Cohere Labs Open Science Community · Remote

  • Lead data creation, validation, evaluation and analysis for a benchmark spanning 51 languages and about 6,100 riddles.
  • Align 100+ contributors from 61+ communities across data, evaluation and analysis tracks on schema design, validation rules and release requirements.
  • Built Python tooling that turns validator output into contributor-actionable guidance, improving annotation consistency and cutting review cycles.
  • Designed the canonical schema, benchmarking pipeline and structured dataset releases on GitHub and Hugging Face.

Independent ML Research & Engineering

Oct 2025 to present

The Meta Gradient · Toronto (Remote)

  • Study representations, loss landscapes and learning dynamics; currently probing world models for hidden physical state (Reverie).
  • Build and release public ML artifacts: reproducible code, model adapters, datasets and implementation-focused technical writing.

Machine Learning Engineer Intern

Feb 2025 to Oct 2025

M2M Tech · Vancouver (Remote)

  • Built end-to-end Python ML pipelines: data cleaning, feature engineering, XGBoost training, hyperparameter tuning and probability calibration.
  • Evaluated models through structured experiments and deployed real-time inference with Docker, FastAPI and Hugging Face Spaces.

Buyer / Procurement Analyst

Nov 2021 to Apr 2023

Whole Foods Market · Toronto

  • Forecast demand and optimised inventory against a $50K monthly budget in a perishable category, where forecast errors carried direct cost.
  • Ran quarterly analysis of sales, waste, inventory and vendor performance to balance stock against demand and reduce shrink.

Selected writing · The Meta Gradient

Education

MASc, Electronic & Computer Science Engineering

2010 to 2012

University of Windsor

Thesis: FPGA implementation of face recognition using Eigenfaces

BE, Electronics & Communication Engineering

2004 to 2008

Anna University

Machine Learning Foundations

2025

University of Toronto

Certifications

  • Cohere Labs ML Summer School (2026)
  • Wolfram · ML Statistical Foundations
  • NVIDIA · AI Infrastructure & Operations
  • Stanford Online · ML Specialization
  • Microsoft Azure AI-900
  • OpenEDG · Python Professional