Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
-
Updated
Aug 5, 2026 - Python
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
Context & Guide For Reinforcement Learning with Verifiable Rewards with Large Language Models
Score the trustworthiness of outputs from any LLM in real-time
Write loops, not prompts. The loop is easy — the verifier is the whole game. VCN night one at Network School.
Open Arena: SLM verifiers across observability platforms, dataset and model catalogs, and value scenarios for LLM, agentic and harness evals.
An RL Enviorment for AES Inversion
A verifiers RLM environment for testing whether adaptive recursive search outperforms brittle manual RAG choreography on long synthetic corpora.
A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.
An open reinforcement-learning (RL) environment that trains LLM agents to use the current fact, not the stale one — verifiable reward for temporal fact-currency, built on verifiers / prime-rl (GRPO, LoRA).
Typed asset shapes + visual + headless views for AI agents. One asset definition. Three rendering targets (HTML / Markdown / Text).
A verifiers RL environment that trains models to propose novel, evidence-grounded, falsifiable hypotheses. Rewards novelty with accountability.
Research proposal for verifier-gated on-policy distillation with explicit evaluation and claim boundaries
Verifiers hello world repo
A verifiable RL environment for TRP ion-channel ligand pharmacology, built on Prime Intellect's Verifiers
Verifiable RL environments for corporate law & governance — deterministic reward, no LLM judge, fully synthetic worlds.
Prime-RL / verifiers TSP environment (10-city hard config, lenient parser, eval-ready)
Reproducible verifier audits, datasheets, agreement metrics, and release gates
Three mechanically-verifiable RFT environments for the checkable 'rails' of a crypto market-structure (CLARITY Act) legal Impact Mapper.
Surge AI is a human-data company that provides large-scale, expert-quality labeled data for training and evaluating frontier AI models. The product surface spans RL environments and agents (rich, complex environments that challenge agentic models), rubrics and verifiers (scoring systems for AI outputs), RLHF (preference and reward data), SFT…
Add a description, image, and links to the verifiers topic page so that developers can more easily learn about it.
To associate your repository with the verifiers topic, visit your repo's landing page and select "manage topics."