Skip to content

evaluate_eval_set silently passes when eval_results_by_eval_id is empty (zero eval cases or num_runs=0) #6951

Description

@gaurav-gandhi-2411

🔴 Required Information

Describe the Bug:
AgentEvaluator.evaluate_eval_set reports overall success on an eval set that
was never actually evaluated. evaluate_eval_set builds eval_results_by_eval_id
from the eval set, iterates it to populate a failures list, and finishes with
assert not failures. If eval_results_by_eval_id ends up empty for any
reason, the loop body never runs, failures stays [], and the assert passes
— the function returns normally with no signal that zero cases were evaluated.

Two independent, ordinary public-API calls reach this with no validation
bypass:

  1. An EvalSet with eval_cases=[]EvalSet.eval_cases has no
    min_length/validator, so this constructs cleanly.
  2. evaluate_eval_set(..., num_runs=0) on a non-empty eval set — num_runs
    has no bounds check anywhere in the file, and inference_requests = [InferenceRequest(...)] * num_runs becomes [].

Both silently produce the same "nothing was evaluated, but nothing failed
either" outcome.

Steps to Reproduce: minimal, self-contained repro (writes a real temp
agent module, no third-party packages):

"""Minimal, self-contained repro for AgentEvaluator.evaluate_eval_set's
vacuous pass on an empty eval set (agent_evaluator.py, evaluate_eval_set).
"""
from __future__ import annotations
import asyncio, tempfile, uuid
from pathlib import Path
from google.adk.evaluation.agent_evaluator import AgentEvaluator
from google.adk.evaluation.eval_config import EvalConfig
from google.adk.evaluation.eval_set import EvalSet


def _write_agent_module(tmp_path: Path) -> str:
    package_name = f"vacuous_pass_repro_{uuid.uuid4().hex}"
    package_dir = tmp_path / package_name
    package_dir.mkdir()
    (package_dir / "__init__.py").write_text("", encoding="utf-8")
    (package_dir / "agent.py").write_text(
        "from google.adk.agents.llm_agent import LlmAgent\n"
        "from google.adk.models.base_llm import BaseLlm\n"
        "from google.adk.models.llm_response import LlmResponse\n"
        "from google.genai import types as genai_types\n\n\n"
        "class _FakeLlm(BaseLlm):\n"
        "    model: str = 'fake-model'\n\n"
        "    @classmethod\n"
        "    def supported_models(cls):\n"
        "        return ['fake-model']\n\n"
        "    async def generate_content_async(self, llm_request, stream=False):\n"
        "        yield LlmResponse(content=genai_types.Content(\n"
        "            parts=[genai_types.Part(text='ok')], role='model'))\n\n\n"
        "root_agent = LlmAgent(name='vacuous_pass_repro_agent', model=_FakeLlm(),\n"
        "                       instruction='Answer briefly.')\n",
        encoding="utf-8",
    )
    return f"{package_name}.agent"


async def main() -> None:
    with tempfile.TemporaryDirectory() as tmp_dir_str:
        tmp_path = Path(tmp_dir_str)
        module_name = _write_agent_module(tmp_path)
        import sys
        sys.path.insert(0, str(tmp_path))
        try:
            empty_eval_set = EvalSet(eval_set_id="empty_set", eval_cases=[])
            result = await AgentEvaluator.evaluate_eval_set(
                agent_module=module_name,
                eval_set=empty_eval_set,
                eval_config=EvalConfig(criteria={"tool_trajectory_avg_score": 1.0}),
            )
            print(f"RESULT: completed WITHOUT raising. Returned: {result!r}")
            print("VERDICT: PASSED -- vacuous pass on zero eval cases.")
        finally:
            sys.path.remove(str(tmp_path))
            sys.modules.pop(module_name, None)
            sys.modules.pop(module_name.rsplit(".", 1)[0], None)


if __name__ == "__main__":
    asyncio.run(main())

Expected Behavior: evaluate_eval_set should raise or otherwise signal
that zero cases were evaluated — a "pass" should mean "every requested case
was checked and none failed," not "nothing was checked."

Observed Behavior:

RESULT: completed WITHOUT raising. Returned: None
VERDICT: PASSED -- vacuous pass on zero eval cases.

Process exit code 0. A second, independent repro (non-empty EvalSet,
num_runs=0) reproduces the identical outcome through the same mechanism.

Environment: google-adk main branch (this repo),
src/google/adk/evaluation/agent_evaluator.py:289 (the assert not failures
line), AgentEvaluator.evaluate_eval_set.

CI implication

Any test suite or CI gate that calls AgentEvaluator.evaluate_eval_set/.evaluate()
directly (rather than going through the adk eval CLI) will report a clean
pass for an eval set that was silently emptied by a config bug, a bad filter,
or an accidental num_runs=0, rather than failing loudly.

Note: this is a separate code path from the adk eval CLI command
(cli_tools_click.py::cli_eval, which builds its own eval_run_summary via
LocalEvalService directly and never calls AgentEvaluator) — confirmed by
reading both call chains. The CLI has an analogous "empty collection defaults
to a pass" shape in its own exit-code logic, which is unaffected by this
issue's scope either way.

I'm happy to send a PR for this if useful — the fix is a small, additive
check (e.g. treat an empty eval_results_by_eval_id as NOT_EVALUATED
rather than a silent pass), similar in shape to a couple of other
NOT_EVALUATED-masking fixes already proposed against this same subsystem.

Metadata

Metadata

Labels

eval[Component] This issue is related to evaluationrequest clarification[Status] The maintainer need clarification or more information from the author

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions