I'm an MSc student in Data Science at ETH Zurich, working on LLM evaluation and agents: how to measure what LLM systems can do, and when their outputs can be trusted.
- CroissantMiner (NeurIPS 2026, lead author): a benchmark and systems for extracting Croissant metadata from ML dataset papers. Across 24 systems, a single pass over the full paper beats four agentic designs. Code
- SkillEval (under review): profiles 3,811 LLMs across 100 interpretable skills; routing questions by skill matches the strongest model at 25% of its cost.
- Before that: 3D human pose and motion capture at the AIT Lab (ETH Zurich), and pose estimation and action recognition at STROMA.
