Skip to content
View berkearda's full-sized avatar
  • ETH Zurich
  • Zurich, Switzerland

Highlights

  • Pro

Block or report berkearda

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
berkearda/README.md

Hi, I'm Berke 👋

I'm an MSc student in Data Science at ETH Zurich, working on LLM evaluation and agents: how to measure what LLM systems can do, and when their outputs can be trusted.

  • CroissantMiner (NeurIPS 2026, lead author): a benchmark and systems for extracting Croissant metadata from ML dataset papers. Across 24 systems, a single pass over the full paper beats four agentic designs. Code
  • SkillEval (under review): profiles 3,811 LLMs across 100 interpretable skills; routing questions by skill matches the strongest model at 25% of its cost.
  • Before that: 3D human pose and motion capture at the AIT Lab (ETH Zurich), and pose estimation and action recognition at STROMA.

LinkedIn

Pinned Loading

  1. croissantminer croissantminer Public

    Benchmark and systems for extracting Croissant metadata from ML dataset papers (NeurIPS 2026, Evaluations and Datasets Track)

    Python 2 1

  2. skilleval skilleval Public

    Skill-level evaluation of large language models with cognitive diagnosis models

    Python 1