An in-the-wild benchmark for AI agents in the OpenClaw Environment.
-
Updated
Aug 6, 2026 - Python
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
Gen AI Evaluation Toolkit on AWS is a flexible, cloud-native accelerator built on AWS serverless architecture that enables comprehensive evaluation of generative AI applications.
DataClawEval: A Benchmark for Engineering Data Agents in Real Industrial Harness
A SnitchBench-style benchmark inverted for the dark-forest problem: does a listener AI alert humans about an alien signal when alerting may doom humanity?
A small benchmark for agent skills, verification artifacts, and fresh-session resumability.
🚀 AI Evolution Factory - From evaluation tool to continuous AI self-improvement platform. Agentic evaluation, auto-finetuning, global P2P testing, and hardware telemetry for local LLMs
Two-tier honest-evaluation harness for agentic RTL design (research, WIP)
Runnable lab measuring invalidation/staleness in agent memory — paper: Are We Ready For An Agent-Native Memory System? (arXiv:2606.24775)
Deterministic synthetic fixtures and hard-gate scoring for agentic biosafety safeguard routing.
Pipeline to investigate structured reasoning and instruction adherence in multimodal LLMs
Measure AI agents’ performance with standardized tests across 314 tasks, 33 domains, and 4 difficulty levels for clear, reproducible comparison.
Add a description, image, and links to the agentic-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the agentic-evaluation topic, visit your repo's landing page and select "manage topics."