Is your feature request related to a problem?
We need a system to automatically clean up query and response data in the LLM calls table after a specified retention period. Without this, the database may become cluttered and inefficient.
Describe the solution you'd like
- Implement a cron job to purge query/response data after a configurable retention window (default 1 week).
- Handle one-time backlog separately from regular cron operations.
- Write a cleanup strategy document detailing cron design, batching, and backlog handling.
- Conduct a staging load test by duplicating ~20k LLM call rows ~5x in the
copy_dev database to measure cleanup time. Do not use production data.
- Ensure S3 cleanup is managed separately and not included in the DB cron/migration code.
Original issue
Context
query + response data in the LLM calls table will be cleaned up via a cron job after a configurable retention window (starting with 1 week).
A step-by-step cleanup strategy doc is to be written (cron design, batching, one-time backlog handling).
Scope / Acceptance criteria
Notes
- Guardrails input/output data is retained separately (anonymized, tagged as product enhancement) and is NOT part of this cleanup.
Is your feature request related to a problem?
We need a system to automatically clean up query and response data in the LLM calls table after a specified retention period. Without this, the database may become cluttered and inefficient.
Describe the solution you'd like
copy_devdatabase to measure cleanup time. Do not use production data.Original issue
Context
query + response data in the LLM calls table will be cleaned up via a cron job after a configurable retention window (starting with 1 week).
A step-by-step cleanup strategy doc is to be written (cron design, batching, one-time backlog handling).
Scope / Acceptance criteria
copy_devdatabase, duplicate staging's ~20k LLM call rows ~5x to reach ~1 lakh, and measure cleanup time. Do NOT pull prod data.Notes