Reverie: Does a World Model Represent What It Cannot See?
2026 to presentIndependent research · Pre-registered
Probing and activation patching of an RSSM world model trained on simulated physics video with no physics in the loss. Tests whether an occluded object’s position stays decodable from the latent state and whether that representation is causally used, with hidden collisions as the central control and CLEVRER for validation.
Abstract ↗
CulturalRiddles: Multicultural Riddles Benchmark
Apr 2026 to presentCohere Labs Open Science Community · First author
Living cultural-reasoning benchmark: 51 languages, 61+ communities, about 6,100 riddles, 100+ contributors. To be presented at MRL 2026, being submitted to TACL, arXiv preprint coming soon.
Research page ↗
Adaption Labs AutoScientist Challenge
20262nd place, Data & Visualization track · $1,000 prize
Adaption Charts P2: synthetic multimodal dataset for chart QA and chart-to-table extraction, built so difficulty comes from reading the chart (truncated axes, missing labels, varied visual styles) rather than from arithmetic. Paired Llama-3.2-3B LoRA adapter for table-grounded QA.
Dataset ↗Adapter ↗
DocuNative: Offline Multilingual Document QA
Mar 2026Cohere Expedition Hackathon · Pipeline lead, 7-person team, one week
Fully offline document QA for newcomers and migrants; nothing leaves the user’s machine. PDF extraction, BGE-M3 embeddings, ChromaDB retrieval, Aya generation and mDeBERTa hallucination checks. 9,000+ automated evaluations across Chinese, Hindi and Polish showed cross-lingual embedding quality, not the LLM, was the main retrieval bottleneck.
Code ↗Write-up ↗
Tamil Agricultural Advisory Dataset
2026Adaption Labs Uncharted Data Challenge · Grade A (9.4/10)
187 curated Q&A records across 20 categories and 48 crops, built from public-sector sources over ten iterative submissions. Honorary recognition from Sara Hooker. Follow-up: an 11-language Indian agricultural dataset.
Dataset ↗11-language ↗