Skip to content

Build replacement embeddings beside active vectors with Ptah #50

Description

@denisvmedia

Hi, I'm Denis, a maintainer of Ptah, an open-source tool for database schema and embedding migrations. We are testing it against real application workflows so model changes can be built and verified before they affect serving traffic. There are docs and a browser playground for trying Ptah's schema commands without installing it.

I noticed that api/scripts/reembed.js changes models in place. When the dimension changes, it drops HNSW and clears memories.vector before filling it with the new embeddings. Could an optional workflow that builds a replacement column beside the active one be useful for larger corpora?

We have published a detailed walkthrough of this proposal: Changing Embedding Models: Zengram Before and After Ptah. It compares the current in-place workflow with building a candidate beside the active vectors, with diagrams, tested commands, and the handover and rollback limitations.

flowchart LR
    A[Current API and encoder] --> B[Existing vectors and HNSW]
    C[Ptah and replacement encoder] --> D[Separate candidate column]
    D --> E[Catch up changes, index, verify]
    E --> F[Deploy matching encoder and column]
Loading

I prepared a small integration in a fork: PGVECTOR_COLUMN selects the prepared column for the existing search and write paths. Its default remains vector. Ptah handles backfill, outbox catch-up, index creation, verification, and explicit cutover approval. It is an operator tool, not a new API service or runtime dependency.

The test uses released Ptah 0.11.4 and real PostgreSQL/pgvector with Zengram's storage code. A local deterministic embedding endpoint lets it test a 1536-to-384 dimension change and a provider outage without credentials. It checks that old search keeps working while the candidate is built, retries preserve the original vectors and payloads, live edits/deletes are caught up, and the new column supports both write paths and tenant/collection filters. This tests migration behavior, not the replacement model's retrieval quality.

The handover is deliberately explicit: briefly pause writers, catch up and verify, then deploy the matching encoder and column together. Zengram does not automatically follow Ptah's active pointer. The one-time setup also adds a stored generated text column, so that table rewrite must be scheduled. Keeping the old vectors alone is not a safe rollback after new writes; the guide explains that boundary.

The fork branch passes CI on Node 20 and 22 plus the PostgreSQL migration test. The implementation is in #51.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions