Skip to content

Milestone 4C: retrieval evaluation - #3

Merged
princerm06 merged 4 commits into
mainfrom
milestone-4c-retrieval-evaluation
Oct 5, 2026
Merged

princerm06 merged 4 commits into
mainfrom
milestone-4c-retrieval-evaluation

Conversation

@princerm06

Copy link
Copy Markdown
Owner

Summary

  • add deterministic retrieval evaluation primitives for Hit Rate@K, Recall@K, and Mean Reciprocal Rank
  • support multiple relevant source pages per natural-language query
  • add unit tests covering metric definitions, ranking behavior, aggregation, and invalid inputs
  • document how to build a representative 15–25 query benchmark from real saved Synapse data
  • extend backend CI to run the evaluation tests alongside 4A/4B regression tests

Design / safety

  • evaluation math is independent of FastAPI, Supabase, and the embedding model
  • CI receives no database credentials or production secrets
  • no database schema, migration, deployment, or user-facing behavior changes
  • a live benchmark against real saved pages remains a separate verification step so we measure actual retrieval quality rather than inventing a flattering score

Verification

GitHub Actions should run all backend unit tests automatically. Leave this PR unmerged until CI is green. After that, the metric harness is verified; collecting the first real retrieval-quality baseline requires a labeled set of representative saved pages/queries and should not be fabricated in CI.

@princerm06
princerm06 marked this pull request as ready for review October 5, 2026 09:09
@princerm06
princerm06 merged commit 83ead6a into main Oct 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant