AI agent for solving tasks from the GAIA Benchmark.
The project uses an LLM through the Groq API together with smolagents and a set of external tools to solve multi-step research and reasoning tasks.
- AI agent based on
smolagents openai/gpt-oss-120bmodel via Groq- Tool calling for external operations
- GAIA Benchmark evaluation
- Automatic generation of
submission.jsonl - API keys loaded from
.env - Ability to test a single task before running the full benchmark
AlbertAgent/
│
├── evaluate_gaia.py # Main GAIA evaluation script
├── tools.py # Agent tools
├── .env # API keys (not committed)
├── .gitignore
├── submission.jsonl # Generated evaluation results
└── README.md
- Python 3.10+
- Groq API key
- Hugging Face token
- Internet connection
Clone the repository:
git clone https://github.com/YOUR_USERNAME/AlbertAgent.git
cd AlbertAgentCreate a virtual environment.
python -m venv venv
venv\Scripts\activatepython3 -m venv venv
source venv/bin/activateInstall dependencies:
pip install -r requirements.txtIf requirements.txt does not exist yet:
pip install smolagents openai datasets python-dotenv tqdmCreate a .env file in the project root:
GROQ_API_KEY=your_groq_api_key
HF_TOKEN=your_huggingface_tokenNever commit .env to GitHub.
Add it to .gitignore:
.env
venv/
__pycache__/
*.pyc
submission.jsonlThe project currently uses:
openai/gpt-oss-120b
through the Groq OpenAI-compatible API:
https://api.groq.com/openai/v1
The agent uses ToolCallingAgent from smolagents to allow the model to call available tools while solving tasks.
Run:
python evaluate_gaia.pyThe program will:
- Load the GAIA validation dataset.
- Select the configured tasks for evaluation.
- Send each task to the AI agent.
- Allow the model to use available tools.
- Receive the final answer.
- Save the result to
submission.jsonl.
Example output:
STEP 1: Starting program...
STEP 2: Loading GAIA dataset...
STEP 3: Dataset loaded. Tasks: 165
STEP 4: Starting evaluation...
Processing task 1/165
========== PROMPT SENT TO AGENT ==========
...
========== MODEL ANSWER ==========
...
Evaluation finished!
Results saved to submission.jsonl
During development, it is recommended to test only one task first.
In evaluate_gaia.py:
test_dataset = dataset.select(range(1))This runs only the first task.
After the agent works correctly, the project can be configured to evaluate the full dataset:
test_dataset = datasetResults are saved to:
submission.jsonl
Each line contains a task ID and the model's answer:
{
"task_id": "example-task-id",
"model_answer": "Example answer"
}The basic workflow is:
GAIA Dataset
|
v
Task
|
v
ToolCallingAgent
|
v
GPT-OSS-120B
|
+---------> Tool
| |
| v
| Tool Result
| |
<-------------+
|
v
Final Answer
|
v
submission.jsonl
Available tools are defined in:
tools.py
and collected through:
ALL_TOOLSThe agent can use these tools to perform operations required by individual GAIA tasks.
Make sure .env exists in the project root:
GROQ_API_KEY=your_keyMake sure the configured model is available through Groq.
Current model:
openai/gpt-oss-120b
The evaluator configures UTF-8 output:
os.environ["PYTHONUTF8"] = "1"
os.environ["PYTHONIOENCODING"] = "utf-8"and configures stdout and stderr to use UTF-8.
Make sure the project uses:
from smolagents import ToolCallingAgentand:
agent = ToolCallingAgent(
tools=ALL_TOOLS,
model=model,
verbosity_level=0,
)Do not publish API keys or tokens.
Never commit:
.env
or any file containing:
GROQ_API_KEY
HF_TOKEN
If a key is accidentally pushed to GitHub, revoke it immediately and generate a new one.
Recommended development workflow:
1. Modify the agent or tools
|
v
2. Run one GAIA task
|
v
3. Check the model answer
|
v
4. Fix errors
|
v
5. Run several tasks
|
v
6. Run the full benchmark
This makes debugging easier than immediately running all GAIA tasks.
this project is under MIT License