Skip to content

Assessment Pipeline: Add audio/video support #1184

Description

@kartpop

Is your feature request related to a problem?
The AI Assessments pipeline currently lacks support for audio and video modalities. This limits the assessment's effectiveness and prevents a comprehensive analysis of inputs.

Describe the solution you'd like

  • Extend multimodal support to include audio/video in AI Assessments.
  • Experiment with Gemini for video handling and explore methods with OpenAI/Anthropic.
  • Sample an audio/video dataset from partners, run the pipeline, conduct human evaluations, and iterate.
  • Support a language mix of ~50–60% English, ~10–15% Tamil, and a strong presence of South Indian languages like Telugu.
Original issue

Context

Extend multimodal support to include audio/video in AI Assessments.

Investigation

  • Video: Gemini does appear to support video directly. Experiments to be done to figure out reasonable way of handling video in Open AI / Anthropic (sampling frames + full audio transcript ?)

Approach / acceptance criteria

  • Sample audio/video dataset from partners, run pipeline, run human evals, iterate. Building the tech is easy; evaluation is the bottleneck.
  • Language mix to support: ~50–60%+ English, Tamil ~10–15%, strong South Indian language presence, Telugu

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions