AI Research Scientist @ Apodex · Pretraining data, agents & evaluation
Based in Singapore. I work on what goes into AI models and how to tell if they're any good: training data pipelines, agent execution, and evaluation. My earlier research spans audio-language models, sentence representations, and dialogue summarization.
- Pretraining data infrastructure — Acquisition and quality assessment of web data for pretraining, built from scratch.
- Apodex 1.1 — Research agents that work with files, data, code, and tools. Open model: Apodex-1.1-mini.
- Apodex 1.0 — Verification-focused agents for deep research. Open model: Apodex-1.0-mini.
- AudioBench — I led this benchmark for audio-language models, covering speech, audio-scene, and voice understanding (first author, NAACL 2025). Paper.
- audio-ai-hub — I maintain a curated collection of audio AI papers, models, benchmarks, and datasets. Browse the hub.
- MiroThinker and MiroFlow — Earlier work on deep research agents and their execution frameworks.
- SBERT-WK — Sentence representations from BERT-based word models (IEEE/ACM TASLP 2020).
- Research Archive on Hugging Face — Historical models and datasets retained for reproducibility.
Website & writing · LinkedIn · X · Hugging Face · ORCID · Email
Currently focused on industry work and not taking on new research collaborations.





