From b50d2ad12adf45a438dd4cdcf385b0aab938b3c8 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Mon, 5 Oct 2026 18:30:15 +0530 Subject: [PATCH 01/11] docs(v2): isolate API reference and response fixes for PRO-2457 Signed-off-by: SohamRatnaparkhi --- .../v2/endpoint/configure-connector.mdx | 2 +- .../v2/endpoint/connectors-overview.mdx | 2 +- .../v2/endpoint/create-connector.mdx | 2 +- api-reference/v2/endpoint/create-tenant.mdx | 37 ++--- .../v2/endpoint/delete-collection.mdx | 10 +- .../v2/endpoint/delete-connector.mdx | 2 +- api-reference/v2/endpoint/delete-source.mdx | 34 ++--- api-reference/v2/endpoint/delete-tenant.mdx | 16 +-- api-reference/v2/endpoint/fetch-content.mdx | 61 ++++---- api-reference/v2/endpoint/ingest-context.mdx | 88 ++++++------ .../v2/endpoint/list-connector-providers.mdx | 2 +- api-reference/v2/endpoint/list-documents.mdx | 42 ++++-- .../v2/endpoint/list-sub-tenants.mdx | 6 +- api-reference/v2/endpoint/list-tenants.mdx | 6 +- .../v2/endpoint/list-webhook-deliveries.mdx | 2 + api-reference/v2/endpoint/query-overview.mdx | 30 ++-- api-reference/v2/endpoint/query.mdx | 95 ++++++------- .../v2/endpoint/register-webhook.mdx | 2 + .../v2/endpoint/retry-webhook-delivery.mdx | 2 +- .../v2/endpoint/source-relations.mdx | 54 ++++--- api-reference/v2/endpoint/source-status.mdx | 42 +++--- .../v2/endpoint/sources-overview.mdx | 20 +-- api-reference/v2/endpoint/subgraph.mdx | 50 ++++--- api-reference/v2/endpoint/submit-feedback.mdx | 50 ++++--- api-reference/v2/endpoint/tenant-stats.mdx | 6 +- api-reference/v2/endpoint/tenant-status.mdx | 18 +-- .../v2/endpoint/tenants-overview.mdx | 9 +- api-reference/v2/endpoint/test-webhook.mdx | 2 +- .../v2/endpoint/update-connector.mdx | 2 + .../v2/endpoint/update-metadata-schema.mdx | 71 ++++------ .../v2/endpoint/update-source-metadata.mdx | 44 ++++-- api-reference/v2/error-responses.mdx | 66 +++++---- api-reference/v2/index.mdx | 57 ++++++-- api-reference/v2/sdks.mdx | 134 ++++++++++++------ essentials/v2/api-results.mdx | 39 ++--- 35 files changed, 649 insertions(+), 456 deletions(-) diff --git a/api-reference/v2/endpoint/configure-connector.mdx b/api-reference/v2/endpoint/configure-connector.mdx index 4e0fa988..bde80070 100644 --- a/api-reference/v2/endpoint/configure-connector.mdx +++ b/api-reference/v2/endpoint/configure-connector.mdx @@ -84,4 +84,4 @@ Use [Get Connector Status](/api-reference/v2/endpoint/get-connector-status) to f - **Next:** [Get Connector Status](/api-reference/v2/endpoint/get-connector-status): follow the first sync - [Update Connector Resource](/api-reference/v2/endpoint/update-connector-resource): change one resource's instructions or access rule - [Discover Resources](/api-reference/v2/endpoint/discover-connector-resources): find resource ids before configuring -- [Connectors - Overview](/api-reference/v2/endpoint/connectors-overview): how synced metadata is merged +- [Connectors: Overview](/api-reference/v2/endpoint/connectors-overview): how synced metadata is merged diff --git a/api-reference/v2/endpoint/connectors-overview.mdx b/api-reference/v2/endpoint/connectors-overview.mdx index 5f4af542..8f86a3f6 100644 --- a/api-reference/v2/endpoint/connectors-overview.mdx +++ b/api-reference/v2/endpoint/connectors-overview.mdx @@ -1,5 +1,5 @@ --- -title: "Connectors - Overview" +title: "Connectors: Overview" description: "Every connector endpoint, the order to call them in, and which one to use to change what." --- diff --git a/api-reference/v2/endpoint/create-connector.mdx b/api-reference/v2/endpoint/create-connector.mdx index 420a09fb..18f63ba4 100644 --- a/api-reference/v2/endpoint/create-connector.mdx +++ b/api-reference/v2/endpoint/create-connector.mdx @@ -73,4 +73,4 @@ The response repeats `database` and `collection` under their deprecated names `t - **Next:** [Discover Resources](/api-reference/v2/endpoint/discover-connector-resources): see what the credentials can reach - **Next:** [Configure Connector](/api-reference/v2/endpoint/configure-connector): choose resources and start syncing - [List Connector Providers](/api-reference/v2/endpoint/list-connector-providers): the credential schema for a provider -- [Connectors - Overview](/api-reference/v2/endpoint/connectors-overview) +- [Connectors: Overview](/api-reference/v2/endpoint/connectors-overview) diff --git a/api-reference/v2/endpoint/create-tenant.mdx b/api-reference/v2/endpoint/create-tenant.mdx index 3b6c045a..279feae1 100644 --- a/api-reference/v2/endpoint/create-tenant.mdx +++ b/api-reference/v2/endpoint/create-tenant.mdx @@ -1,6 +1,6 @@ --- title: "Create Database" -description: "Creates a space for storing context. " +description: "Creates a space for storing context." openapi: "api-reference/v2/openapi.json POST /databases" --- @@ -16,7 +16,6 @@ response = client.databases.create( "name": "category", "data_type": "VARCHAR", "max_length": 256, - "enable_match": True, }, { "name": "product_description", @@ -37,7 +36,6 @@ const response = await client.databases.create({ name: "category", dataType: "VARCHAR", maxLength: 256, - enableMatch: true, }, { name: "product_description", @@ -61,8 +59,7 @@ curl -X POST 'https://api.hydradb.com/databases' \ { "name": "category", "data_type": "VARCHAR", - "max_length": 256, - "enable_match": true + "max_length": 256 }, { "name": "product_description", @@ -85,7 +82,7 @@ curl -X POST 'https://api.hydradb.com/databases' \ | Name | Description | | --- | --- | -| | Account-scoped database identifier. Use a stable, case-sensitive ID up to 25 characters; prefer lowercase letters, numbers, and underscores for portability. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | +| | Account-scoped database identifier. Use a stable ID up to 255 characters of lowercase letters, digits, `-`, and `_`; anything else, including uppercase or spaces, returns `400`. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | | | Defines database-level metadata fields. See the [Scoping using metadata](/essentials/v2/metadata#step-1a-declare-the-schema-at-database-creation) guide for detailed schema parameters. Formerly `tenant_metadata_schema`; the `tenant_metadata_schema` alias is still accepted (deprecated). (default=`null`) | ## Successful response @@ -131,21 +128,27 @@ Always check if a database is ready before using it. Use [Database Status](/api- 1. Create the database with `POST /databases` 2. **Default collection:** No collection exists until your first write. The first time you ingest without an explicit `collection`, HydraDB creates the database's default collection, which then stores all context written without a `collection`. Create additional collections at any time to scope data to users, teams, or projects. -3. **Retry failed databases:** If a database appears in `data.failed_databases`, re-create that database with `POST /databases` after addressing the reported issue. Poll status again before ingestion. +3. **Retry failed databases:** If a database appears in `data.failed_databases` from [List Databases](/api-reference/v2/endpoint/list-tenants), re-create that database with `POST /databases` after addressing the reported issue. Poll status again before ingestion. 4. Start [ingesting context](/api-reference/v2/endpoint/ingest-context) once databases are ready -5. Check status of [ingestion](/api-reference/v2/endpoint/source-status). Start querying the database once the recently ingested sources show `completed` +5. Check status of [ingestion](/api-reference/v2/endpoint/source-status). Start querying the database once the recently ingested sources show `graph_creation` (searchable) or `completed` --- ## Defining metadata schema - Schema field names are **immutable** after database creation. You can add per-document free-form metadata fields at ingestion time, and add new database-level fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but updates are additive only: no delete, rename, type change, or Milvus backfill for newly added dense/sparse metadata lanes. Plan your schema carefully before creating the database. + Schema field names are **immutable** after database creation. You can add per-document free-form metadata fields at ingestion time, and add new database-level fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but updates are additive only: no delete, rename, or type change. Dense and sparse metadata lanes (`enable_dense_embedding`, `enable_sparse_embedding`) can only be declared here, at creation. Plan your schema carefully before creating the database. -You can define a custom schema at database creation to enable exact-match metadata filtering (`enable_match`) or semantic/BM25 search over metadata text fields (`enable_dense_embedding` / `enable_sparse_embedding`). +You can define a custom schema at database creation to declare the `metadata` fields you filter on, and to enable semantic/BM25 search over metadata text fields (`enable_dense_embedding` / `enable_sparse_embedding`). Each dense or sparse flag adds one vector field, so a field with both uses two; a database can have at most 6. Going over returns `400`, as does declaring an `ARRAY` field. -For detailed parameters, valid data types, limits, shorthand flags, and comprehensive examples, see the [metadata](/essentials/v2/metadata) guide. +For detailed parameters, valid data types, limits, and comprehensive examples, see the [metadata](/essentials/v2/metadata) guide. + +--- + +## Errors + +Common codes: `400 INVALID_INPUT` (missing or invalid `database`, or an invalid schema), `403 FORBIDDEN` (your plan's database limit is reached), `409 DATABASE_ALREADY_EXISTS` (the `database` is already in use; the deprecated `POST /tenants` route returns `INVALID_INPUT` instead), and `500 INTERNAL_ERROR` (retry; if the message says the rollback also failed, delete the database first, then create it again). See [Error Responses](/api-reference/v2/error-responses) for the full list. --- @@ -153,9 +156,9 @@ For detailed parameters, valid data types, limits, shorthand flags, and comprehe ## **Related Resources** -- **Next:** [Database Status](/api-reference/v2/endpoint/tenant-status) - poll until provisioning completes -- **Next:** [Ingest Context](/api-reference/v2/endpoint/ingest-context) - start ingesting data once status is ready -- **Related:** [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema) - add metadata schema fields later -- **Related:** [Delete Database](/api-reference/v2/endpoint/delete-tenant) - teardown -- **Read more:** [Concepts → Multi-Tenant Support](/essentials/v2/multi-tenant) -- **Read more:** [Usage → Metadata](/essentials/v2/metadata) +- **Next:** [Database Status](/api-reference/v2/endpoint/tenant-status): poll until provisioning completes +- **Next:** [Ingest Context](/api-reference/v2/endpoint/ingest-context): start ingesting data once status is ready +- **Related:** [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema): add metadata schema fields later +- **Related:** [Delete Database](/api-reference/v2/endpoint/delete-tenant): teardown +- **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) +- **Read more:** [Usage: Metadata](/essentials/v2/metadata) diff --git a/api-reference/v2/endpoint/delete-collection.mdx b/api-reference/v2/endpoint/delete-collection.mdx index 9850a8fd..b7b5f5c9 100644 --- a/api-reference/v2/endpoint/delete-collection.mdx +++ b/api-reference/v2/endpoint/delete-collection.mdx @@ -95,15 +95,15 @@ You do not need one to reuse the name safely. Ingestion creates a missing collec ## Errors -Common codes: `400 VALIDATION_ERROR`, `404 DATABASE_NOT_FOUND`, `401 UNAUTHORIZED`. See [Error Responses](/api-reference/v2/error-responses) for the full list. +Common codes: `400 INVALID_INPUT` (`database` or `collection` is missing), `404 DATABASE_NOT_FOUND` (a database that is itself being deleted returns `404 NOT_FOUND`), `401 UNAUTHORIZED`, `503 SERVICE_UNAVAILABLE` (an earlier failed deletion of this collection is still releasing its lock; the collection stays fenced, so retry in a few seconds), and `500 INTERNAL_ERROR` (transient; retry the same `DELETE`). See [Error Responses](/api-reference/v2/error-responses) for the full list.
**Related Resources** -- **Before this:** [List Collections](/api-reference/v2/endpoint/list-sub-tenants) - find the collection ID -- **Alternative:** [Delete Context](/api-reference/v2/endpoint/delete-source) - remove specific knowledge or memories without deleting the collection -- **Larger scope:** [Delete Database](/api-reference/v2/endpoint/delete-tenant) - remove the entire database -- **Read more:** [Concepts → Multi tenancy](/essentials/v2/multi-tenant) +- **Before this:** [List Collections](/api-reference/v2/endpoint/list-sub-tenants): find the collection ID +- **Alternative:** [Delete Context](/api-reference/v2/endpoint/delete-source): remove specific knowledge or memories without deleting the collection +- **Larger scope:** [Delete Database](/api-reference/v2/endpoint/delete-tenant): remove the entire database +- **Read more:** [Concepts: Multi tenancy](/essentials/v2/multi-tenant) diff --git a/api-reference/v2/endpoint/delete-connector.mdx b/api-reference/v2/endpoint/delete-connector.mdx index f056b5f8..834e88ed 100644 --- a/api-reference/v2/endpoint/delete-connector.mdx +++ b/api-reference/v2/endpoint/delete-connector.mdx @@ -33,5 +33,5 @@ curl -X DELETE 'https://api.hydradb.com/connectors/{id}' \ ## Related Resources -- [Connectors - Overview](/api-reference/v2/endpoint/connectors-overview) +- [Connectors: Overview](/api-reference/v2/endpoint/connectors-overview) - [List Connectors](/api-reference/v2/endpoint/list-connectors): check the connector is gone diff --git a/api-reference/v2/endpoint/delete-source.mdx b/api-reference/v2/endpoint/delete-source.mdx index 9f7827a8..df29ca80 100644 --- a/api-reference/v2/endpoint/delete-source.mdx +++ b/api-reference/v2/endpoint/delete-source.mdx @@ -8,8 +8,8 @@ import { Field } from "/snippets/field.jsx"; Specify the resource category with the `type` parameter: -- `type=knowledge` _(default)_ - delete knowledge sources. -- `type=memory` - delete memories. +- `type=knowledge` _(default)_: delete knowledge sources. +- `type=memory`: delete memories. Pass one or more IDs in `ids`. Send `database`, `collection`, `ids`, and `type` as top-level fields in the request body. @@ -59,7 +59,7 @@ curl -X DELETE 'https://api.hydradb.com/context' \ "success": true, "data": { "success": true, - "message": "Delete completed", + "message": "Successfully deleted 2 source(s)", "results": [ { "id": "policy_main", "deleted": true }, { "id": "runbook_deploy", "deleted": true } @@ -137,7 +137,7 @@ curl -X DELETE 'https://api.hydradb.com/context' \ } ``` -```json 404 - nothing matched (strict mode) +```json 404: nothing matched (strict mode) { "success": false, "data": { @@ -168,11 +168,11 @@ curl -X DELETE 'https://api.hydradb.com/context' \ ## Status codes -**By default, every outcome returns `200`** - including a delete that removed +**By default, every outcome returns `200`**, including a delete that removed nothing. The real result is in the body, so check `data.deleted_count` and `data.results[]` rather than the status code. -```bash Default - always 200 +```bash Default: always 200 curl -X DELETE 'https://api.hydradb.com/context' \ -H "Authorization: Bearer " \ -H "API-Version: 2" \ @@ -184,10 +184,10 @@ curl -X DELETE 'https://api.hydradb.com/context' \ Send `X-HydraDB-Delete-Status: strict` and a delete that did not happen returns `404`, `409`, or `500` instead of `200`. **This is the recommended mode for new -integrations** - it is the only way to detect a failed delete from the status +integrations**: it is the only way to detect a failed delete from the status code alone. -```bash Strict - honest status codes +```bash Strict: honest status codes curl -X DELETE 'https://api.hydradb.com/context' \ -H "Authorization: Bearer " \ -H "API-Version: 2" \ @@ -202,13 +202,13 @@ In strict mode: | --- | --- | --- | | `200` | n/a | At least one source was deleted. Check `results[]` for per-ID outcomes. | | `404` | `NOT_FOUND` | None of the given `ids` matched anything to delete. | -| `409` | `SOURCE_PROCESSING` | A source is still indexing. Retry after ingestion completes - see the `Retry-After` header. | +| `409` | `SOURCE_PROCESSING` | A source is still indexing. Retry after ingestion completes (see the `Retry-After` header). | | `500` | `INTERNAL_ERROR` | A store failed to remove the source. The delete is retryable. | Deleting a source that is still indexing is the case worth handling, and the main reason to turn strict mode on. Ingestion is asynchronous, so an - ingest-then-delete sequence - what most teardown and test scripts do - can + ingest-then-delete sequence (what most teardown and test scripts do) can reach the source before it finishes indexing. The source is **not** deleted. In the default mode that comes back as a `200` with `deleted_count: 0`, which @@ -221,7 +221,7 @@ On `404`, `409`, and `500` the response `data` still carries the same `results` / `deleted_count` payload a `200` carries, so you can read per-ID outcomes on a failure exactly as you would on success: -```json 409 - still indexing (strict mode) +```json 409: still indexing (strict mode) { "success": false, "data": { @@ -253,7 +253,7 @@ The header always wins. Without it, the server default applies. | Request | Behaviour | | --- | --- | -| No header | The server default - currently `legacy`, so `200` for every outcome. | +| No header | The server default, currently `legacy`, so `200` for every outcome. | | `X-HydraDB-Delete-Status: strict` | Honest `404` / `409` / `500`. | | `X-HydraDB-Delete-Status: legacy` | `200` for every outcome, whatever the server default. | @@ -263,7 +263,7 @@ The header always wins. Without it, the server default applies. When it happens, `X-HydraDB-Delete-Status: legacy` keeps the unconditional `200` for any integration that is not ready. Both header values are supported - and neither has a removal date - if that ever changes, we will announce it. + and neither has a removal date. If that ever changes, we will announce it. If your integration checks `response.ok` or `status == 200` today, it is treating blocked deletes as successful. That is the failure strict mode @@ -272,9 +272,9 @@ The header always wins. Without it, the server default applies. ## Some additional notes -- **Partial-success semantics:** For `type=knowledge`, each ID is reported independently in `results[]`. A failure on one ID does not stop the rest. For `type=memory`, the response reports an aggregate `user_memory_deleted` reflecting _all_ listed IDs. One exception: if any source in the request is still indexing, the whole request is refused and nothing is deleted - reported as `409` in strict mode, and as a `200` with `deleted_count: 0` by default. +- **Partial-success semantics:** For `type=knowledge`, each ID is reported independently in `results[]`. A failure on one ID does not stop the rest. For `type=memory`, the response reports an aggregate `user_memory_deleted` reflecting _all_ listed IDs. One exception: if any source in the request is still indexing, the whole request is refused and nothing is deleted (reported as `409` in strict mode, and as a `200` with `deleted_count: 0` by default). - **Retrieval drops the source immediately:** Even before background cleanup finishes, deleted IDs disappear from `/query` and `/context/list` responses. -- **Mixed deletes need two calls:** To delete both knowledge and memory items, send two requests - one with `type=knowledge`, one with `type=memory`. +- **Mixed deletes need two calls:** To delete both knowledge and memory items, send two requests: one with `type=knowledge`, one with `type=memory`.
@@ -283,6 +283,6 @@ The header always wins. Without it, the server default applies. - **Find IDs:** [List Documents](/api-reference/v2/endpoint/list-documents) - **Perform a query:** [Query](/api-reference/v2/endpoint/query) - **Re-add content:** [Ingest Context](/api-reference/v2/endpoint/ingest-context) - - **Bigger hammer:** [Delete Database](/api-reference/v2/endpoint/delete-tenant) - removes the entire database - - **Read more:** [Context Management - Overview](/api-reference/v2/endpoint/sources-overview) + - **Bigger hammer:** [Delete Database](/api-reference/v2/endpoint/delete-tenant), which removes the entire database + - **Read more:** [Context Management: Overview](/api-reference/v2/endpoint/sources-overview) diff --git a/api-reference/v2/endpoint/delete-tenant.mdx b/api-reference/v2/endpoint/delete-tenant.mdx index 4fa82d43..2d5405b6 100644 --- a/api-reference/v2/endpoint/delete-tenant.mdx +++ b/api-reference/v2/endpoint/delete-tenant.mdx @@ -9,7 +9,7 @@ import { Field } from "/snippets/field.jsx"; This action is irreversible. Deleting a database removes all of its associated data, including ingested content, memories, metadata schema, vector indices, and graphs. There is no soft-delete and no recovery window. -The examples below use a placeholder name, `database_to_delete`. Replace it with the database you actually mean to destroy before running them - and check the name twice on a shared or team account, where you may not be the only one using it. +The examples below use a placeholder name, `database_to_delete`. Replace it with the database you actually mean to destroy before running them, and check the name twice on a shared or team account, where you may not be the only one using it. @@ -77,24 +77,24 @@ After deletion completes, the same `database` can be used in a new `POST /databa ## Behavior notes -**Irreversible action.** Ingested documents, memories, embeddings, graph nodes, metadata schema, and storage objects are permanently removed. There is no recovery window - ensure you have a backup if the content matters. +**Irreversible action.** Ingested documents, memories, embeddings, graph nodes, metadata schema, and storage objects are permanently removed. There is no recovery window, so ensure you have a backup if the content matters. - **Stop in-flight work first:** Stop all ingestion, polling, query, and background jobs targeting this database before deleting. Calls made after deregistration can fail with `DATABASE_NOT_FOUND` even while infrastructure cleanup is still running. - **Async cleanup:** The endpoint returns immediately after deregistering the database. Infrastructure cleanup of vector stores, graphs, and storage objects runs in the background and may take a few minutes to complete. -- **Repeat calls:** Deleting an already-deleted database returns `404 DATABASE_NOT_FOUND`. Deleting a database that is still provisioning or deleting is treated as a request to tear down that database. +- **Repeat calls:** Deleting an already-deleted database returns `404 DATABASE_NOT_FOUND`. Deleting a database that is still deleting returns `200` again. Deleting a database that is still provisioning, or whose provisioning or earlier deletion failed, is treated as a request to tear down that database. ## Errors -Common codes: `404 DATABASE_NOT_FOUND`, `401 UNAUTHORIZED`, `422 VALIDATION_ERROR`. See [Error Responses](/api-reference/v2/error-responses) for the full list. +Common codes: `400 INVALID_INPUT` (missing `database`), `401 UNAUTHORIZED`, `404 DATABASE_NOT_FOUND`, and `500 INTERNAL_ERROR` (transient; retry the same `DELETE`). See [Error Responses](/api-reference/v2/error-responses) for the full list.
**Related Resources** -- **Before this:** [List Databases](/api-reference/v2/endpoint/list-tenants) - find the database ID -- **Alternative:** [Delete Collection](/api-reference/v2/endpoint/delete-collection) - remove one collection without deleting the whole database -- **Alternative:** [Delete Context](/api-reference/v2/endpoint/delete-source) - remove specific knowledge or memories without deleting the whole database -- **Read more:** [Concepts → Multi-Tenant Support](/essentials/v2/multi-tenant) +- **Before this:** [List Databases](/api-reference/v2/endpoint/list-tenants): find the database ID +- **Alternative:** [Delete Collection](/api-reference/v2/endpoint/delete-collection): remove one collection without deleting the whole database +- **Alternative:** [Delete Context](/api-reference/v2/endpoint/delete-source): remove specific knowledge or memories without deleting the whole database +- **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) diff --git a/api-reference/v2/endpoint/fetch-content.mdx b/api-reference/v2/endpoint/fetch-content.mdx index 8dcfa551..4531f9c7 100644 --- a/api-reference/v2/endpoint/fetch-content.mdx +++ b/api-reference/v2/endpoint/fetch-content.mdx @@ -49,20 +49,21 @@ curl -G 'https://api.hydradb.com/context/inspect' \ | | Owning database. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | | | Collection scope. If omitted, the default collection is used. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`null`) | | | What to return. See [Fetch modes](#fetch-modes). (default=`"both"`) | -| | TTL of the presigned URL (when `mode` includes `url`). Range `60 ≤ x ≤ 604800` (7 days). (default=`3600`) | +| | TTL of the presigned URL (when `mode` includes `url`). Range `60` to `604800` (7 days). (default=`3600`) | +| | Principals to answer as (document ACLs). A source they cannot see returns the same `404` as a missing one. Repeated (`acl=a&acl=b`) or comma-separated. Omit for no ACL scoping. See [Access Control](/essentials/v2/access-control). | ## Fetch modes | Mode | Returns | Use when | | --- | --- | --- | -| `content` | The payload in `content_base64` (base64-encoded); `content` is `null` in this mode. Decode `content_base64` to recover the parsed text or the raw binary bytes. | You want to render the content in-app or feed it to another model. | -| `url` | Presigned URL (`presigned_url`) valid for `expiry_seconds`. | You want a client or a service to download the original file directly without proxying through your backend. | -| `both`_(default)_ | Text **and** presigned URL. | UI flows that show parsed text inline plus a "Download original" link. | +| `content` | The stored bytes. Valid UTF-8 (plain text, Markdown, CSV, memory text, app sources) comes back as text in `content`; anything else (PDF, DOCX, images) comes back base64-encoded in `content_base64`. The other field is `null`. | You want to render the content in-app or feed it to another model. | +| `url` | Presigned URL (`presigned_url`) valid for `expiry_seconds`. `content`, `content_base64`, and `inferred_content` are `null`. | You want a client or a service to download the original file directly without proxying through your backend. | +| `both`_(default)_ | Everything `content` returns, **and** presigned URL. | UI flows that show text inline plus a "Download original" link. | -Regardless of `mode`, the response also includes `inferred_content` when available - the model-derived text for the source (for memories, the inferred memory statement; for knowledge, derived/normalized content). It is `null` when the source has no inferred content. +In `content` and `both` modes, the response also includes `inferred_content` when available: the model-derived text for the source (for memories, the inferred memory statement; for knowledge, derived/normalized content). It is `null` when the source has no inferred content. -Use `mode=url` when you need the original file. Use `mode=content` when you only need extracted text for display, summarization, or prompting. +Use `mode=url` when you need the original file. Use `mode=content` when you only need the content inline for display, summarization, or prompting. It is the stored original, so a PDF comes back as base64 bytes, not extracted text. ### Mode examples @@ -76,12 +77,13 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "success": true, "id": "policy_main", "content": null, - "content_base64": "U2VjdGlvbiAxLiBBdXRoZW50aWNhdGlvbiBwb2xpY2llcy4uLg==", + "content_base64": "JVBERi0xLjcKJeLjz9MKMSAwIG9iago8PC9UeXBlL0NhdGFsb2c...", "inferred_content": null, "presigned_url": null, "content_type": "application/pdf", "size_bytes": 1842233, - "message": "File fetched successfully" + "message": "File fetched successfully", + "error": null }, "error": null, "meta": { @@ -104,7 +106,8 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "presigned_url": "https://storage.hydradb.com/.../policy_main.pdf?X-Amz-...", "content_type": "application/pdf", "size_bytes": 1842233, - "message": "File fetched successfully" + "message": "File fetched successfully", + "error": null }, "error": null, "meta": { @@ -127,7 +130,8 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "presigned_url": "https://storage.hydradb.com/.../diagram.png?X-Amz-...", "content_type": "image/png", "size_bytes": 92844, - "message": "File fetched successfully" + "message": "File fetched successfully", + "error": null }, "error": null, "meta": { @@ -137,7 +141,7 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only } ``` - + ```json { "success": true, @@ -147,10 +151,11 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "content": "Prefers concise answers and dark mode.", "content_base64": null, "inferred_content": "User prefers concise answers and dark mode.", - "presigned_url": null, - "content_type": "text/plain", - "size_bytes": null, - "message": "Memory fetched successfully" + "presigned_url": "https://storage.hydradb.com/.../mem_user_alex_tone?X-Amz-...", + "content_type": "text/plain; charset=utf-8", + "size_bytes": 38, + "message": "File fetched successfully", + "error": null }, "error": null, "meta": { @@ -170,13 +175,14 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "data": { "success": true, "id": "policy_main", - "content": "Section 1. Authentication policies...", - "content_base64": null, + "content": null, + "content_base64": "JVBERi0xLjcKJeLjz9MKMSAwIG9iago8PC9UeXBlL0NhdGFsb2c...", "inferred_content": null, "presigned_url": "https://storage.hydradb.com/.../policy_main.pdf?X-Amz-...", "content_type": "application/pdf", "size_bytes": 1842233, - "message": "File fetched successfully" + "message": "File fetched successfully", + "error": null }, "error": null, "meta": { @@ -191,8 +197,8 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "success": false, "data": null, "error": { - "code": "SOURCE_NOT_FOUND", - "message": "Source not found" + "code": "NOT_FOUND", + "message": "Source 'policy_main' not found. Verify the id is correct and the source has been ingested." }, "meta": { "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", @@ -215,7 +221,8 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only "presigned_url": "https://storage.hydradb.com/.../diagram.png?X-Amz-...", "content_type": "image/png", "size_bytes": 92844, - "message": "File fetched successfully" + "message": "File fetched successfully", + "error": null }, "error": null, "meta": { @@ -228,13 +235,17 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only ## Behavior notes - **Text vs binary handling.** In `mode=both`, text-parseable sources (PDF, DOCX, MD, TXT) populate `content` with parsed text while non-text/binary files come back in `content_base64`; check both fields when handling unknown content types. In `mode=content`, the payload is always returned base64-encoded in `content_base64` and `content` is `null` - decode `content_base64` to recover the text or bytes. + **Text vs binary handling:** In `mode=content` and `mode=both`, you get the stored original, not parsed text. Valid UTF-8 sources (TXT, MD, CSV, memory text, app sources) populate `content` with text, while binary files (PDF, DOCX, images) come back in `content_base64`; check both fields when handling unknown content types. -- **`inferred_content`:** Alongside the raw `content`, the response carries `inferred_content` - the model-derived text for the source. For **memories** this is the inferred memory statement (e.g. `"User prefers concise answers and dark mode."`); for **knowledge** sources it is typically `null` unless derived content exists. It is returned for every `mode`. +- **`inferred_content`:** Alongside the raw `content`, the response carries `inferred_content`: the model-derived text for the source. For **memories** this is the inferred memory statement (e.g. `"User prefers concise answers and dark mode."`); for **knowledge** sources it is typically `null` unless derived content exists. It is returned in `content` and `both` modes; `mode=url` returns it as `null`. -- **Recently ingested sources:** Fetching immediately after ingestion may return a record before the parsed text is ready. For reliable reads, use [Ingestion Status](/api-reference/v2/endpoint/source-status) first. +- **Recently ingested sources:** Fetching immediately after ingestion returns the stored original, but `inferred_content` stays `null` until processing produces it. For reliable reads, use [Ingestion Status](/api-reference/v2/endpoint/source-status) first. - **Presigned URL TTL:** The URL is valid only for `expiry_seconds`. Anyone with the URL can download the file during that window, so treat it as a short-lived secret. -- **Memory items:** Fetching a memory's `id` returns its raw text content. There are no presigned URLs for memory items - `mode=url` and `mode=both` return `presigned_url: null`. +- **Memory items:** Fetching a memory's `id` returns its raw text content (a conversation memory returns its stored JSON, with `user_name` and `pairs`). Memories are stored like files, so `mode=url` and `mode=both` also return a `presigned_url`. + +## Errors + +Common codes: `400 INVALID_INPUT` (missing `id` or `database`, an unknown `mode`, or `expiry_seconds` outside `60` to `604800`), `404 DATABASE_NOT_FOUND`, `404 NOT_FOUND` (the source does not exist, or `acl` is set and its principals may not see it), `500 INTERNAL_ERROR` (a transient storage error; retry). See [Error Responses](/api-reference/v2/error-responses) for the full list.
diff --git a/api-reference/v2/endpoint/ingest-context.mdx b/api-reference/v2/endpoint/ingest-context.mdx index 552f4af7..58c970d1 100644 --- a/api-reference/v2/endpoint/ingest-context.mdx +++ b/api-reference/v2/endpoint/ingest-context.mdx @@ -8,10 +8,10 @@ import { Field } from "/snippets/field.jsx"; When context is of `type=knowledge`: -1. **Documents** - Use the `documents` field for binary uploads that HydraDB should parse: PDFs, Office and iWork files, spreadsheets, images and plain text. See [Supported file formats](#supported-file-formats) for the full list. -2. **App Sources** - Use `app_knowledge` for pre-extracted JSON content (Slack threads, Notion pages, emails, tickets). Read more about [ingesting knowledge from your apps](/essentials/v2/app-sources). +1. **Documents:** Use the `documents` field for binary uploads that HydraDB should parse: PDFs, Office and iWork files, spreadsheets, images and plain text. See [Supported file formats](#supported-file-formats) for the full list. +2. **App Sources:** Use `app_knowledge` for pre-extracted JSON content (Slack threads, Notion pages, emails, tickets). Read more about [ingesting knowledge from your apps](/essentials/v2/app-sources). -When context is of `type=memory` +When context is of `type=memory`: Use `memories` for per-user content, scoped with `collection`. Set `infer: true` to let HydraDB extract preferences from raw signals, or `infer: false` to store the text verbatim. Read more about [ingesting memories](/essentials/v2/memories). @@ -94,7 +94,7 @@ const knowledgeResult = await client.context.ingest({ type: "knowledge", database: "acme_corp", collection: "team_docs", - upsert: true, + upsert: "true", // Use documents when HydraDB should parse PDFs, DOCX, CSV, Markdown, or text files. documents: [ @@ -313,18 +313,18 @@ response.raise_for_status() | --- | --- | | | Use singular `"memory"` when writing memories. (default=`"knowledge"`) | | | Target database. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | -| | Logical partition inside the database. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`""` - default collection) | -| | Replace existing sources with the same ID. Set to `false` to error on conflict. (default=`true`) | +| | Logical partition inside the database. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`""`, the default collection) | +| | Replace existing sources with the same ID. Set to `false` to error on conflict: a document or app source whose ID exists comes back `failed` with `error_code: "E9003"`. (default=`true`) | | | Binary uploads, **knowledge only**. Required when `type=knowledge` and you want HydraDB to parse documents. Omit when ingesting memories. See [Supported file formats](#supported-file-formats). (default=`[]`) | -| | One entry per file in `documents`, in the same order. If omitted, documents index with inferred defaults such as filename/title. See the item shape below. | +| | One entry per file in `documents`, in the same order. The counts must match, or the request returns `400`. If omitted, documents index with inferred defaults such as filename/title. See the item shape below. | | | Pre-extracted source objects (Slack, Notion, web pages, etc.), **knowledge only**. See the `app_knowledge` item shape below. | -| | Map of source id → your own entities + relations - replaces LLM graph extraction for each keyed source. Works for `type=knowledge` (key = a `document_metadata` id or `app_knowledge` item id) and `type=memory` (key = a memory `id`). See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) and the shape below. | +| | Map of source id to your own entities + relations. Replaces LLM graph extraction for each keyed source. Works for `type=knowledge` (key = a `document_metadata` id or `app_knowledge` item id) and `type=memory` (key = a memory `id`). See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) and the shape below. | | | Memory items, **memory only**. Required and non-empty when `type=memory`. Use plural `memories` for the form field, even though `type` is singular `memory`. See the `memories` item shape below. | -1. **`id` must not contain a comma (`,`).** The comma is reserved as the id separator on [Ingestion Status](/api-reference/v2/endpoint/source-status) (`GET /context/status?ids=a,b`), so an `id` containing a comma cannot be looked up unambiguously. This applies to every `id` you supply - `document_metadata`, `app_knowledge`, and `memories` items. Ingesting an item whose `id` contains a comma is rejected with a `400`. +1. **`id` must not contain a comma (`,`).** The comma is reserved as the id separator on [Ingestion Status](/api-reference/v2/endpoint/source-status) (`GET /context/status?ids=a,b`), so an `id` containing a comma cannot be looked up unambiguously. This applies to every `id` you supply: `document_metadata`, `app_knowledge`, and `memories` items. Ingesting an item whose `id` contains a comma is rejected with a `400`. -2. **`202 Accepted` means queued, not indexed.** Ingestion is asynchronous. A successful response only confirms your sources were accepted, not that they are ready to query. Before querying, poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with the returned IDs until each source reaches `completed` or `errored`. Alternatively, register a webhook for `indexing.status_changed` events (see [Webhooks](/essentials/v2/webhooks)). +2. **`202 Accepted` means queued, not indexed.** Ingestion is asynchronous. A successful response only confirms your sources were accepted, not that they are ready to query. Before querying, poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with the returned IDs until each source reaches `graph_creation` (searchable), `completed`, or `errored`. Alternatively, register a webhook for `indexing.status_changed` events (see [Webhooks](/essentials/v2/webhooks)). @@ -421,9 +421,9 @@ The request returns `202` as usual. The rejected file comes back with `status: " ### Document metadata -Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can be passed alongside each uploaded document to control indexing, filtering, and display. The key list is closed - an item carrying any other key is rejected with a `400` naming it, rather than being silently dropped. See the field reference below. +Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `infer`, `evidence_kind`, `evidence_subject`) can be passed alongside each uploaded document to control indexing, filtering, and display. The key list is closed: an item carrying any other key is rejected with a `400` naming it, rather than being silently dropped. See the field reference below. - + ```python Python SDK @@ -477,19 +477,24 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can | Field | Description | | --- | --- | - | | Optional context ID. If set, becomes the `id` for this document (use your app's document ID for parity). Must not contain a comma (`,`) - it is reserved as the id separator on `/context/status?ids=`. | - | | Database-schema fields for filtering and search. Keys must be declared in `database_metadata_schema`. (default=`{}`) | - | | Free-form per-document fields for display or bookkeeping. To filter on these at search time, nest under `metadata_filters.additional_metadata`. (default=`{}`) | + | | Optional context ID. If set, becomes the `id` for this document (use your app's document ID for parity). Must not contain a comma (`,`). It is reserved as the id separator on `/context/status?ids=`. | + | | Database-schema fields for filtering and search. Keys must be declared in `database_metadata_schema`. At most 16 KiB. (default=`{}`) | + | | Free-form per-document fields for display or bookkeeping. To filter on these at search time, nest under `metadata_filters.additional_metadata`. At most 1 KiB. (default=`{}`) | | | Declare forceful relations to other sources. Shape: `{ "ids": ["...", "..."] }`. Surfaced via `additional_context` in `mode: "thinking"` search. | + | | Run model enrichment on this file. (default=`false`) | + | | How this content came to be known: `assertion`, `record`, `said`, `done`, `third_party`, or `inferred`. Any other value returns `400`. | + | | A stable handle for who or what the content is about. | + + Size limits are measured on the compact JSON of the whole object, in UTF-8 bytes, and apply the same way to `app_knowledge` and `memories` items. - Those four are the only keys accepted. Anything else - including `title`, `type`, `url` and `timestamp` - is rejected with a `400` naming the unsupported key, rather than being silently dropped. + Those seven, plus the legacy aliases `source_id` and `file_id` (for `id`) and `document_metadata` (for `additional_metadata`), are the only keys accepted. Anything else (including `title`, `type`, `url` and `timestamp`) is rejected with a `400` naming the unsupported key, rather than being silently dropped. In particular, a document's **title is derived, not settable**. It defaults to the uploaded filename and is returned as `source_title` on query results and `title` on `/context/list`. It cannot be overridden at ingest, and [`PATCH /context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata) only merges `additional_metadata` and `database_metadata`. Use `additional_metadata` for your own display fields, or ingest through `app_knowledge`, whose items carry an explicit `title`. - + ```python Python SDK @@ -556,7 +561,7 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can | Field | Description | | --- | --- | - | | Context ID. Treated as the upsert key. Send an empty string to have one generated upstream. Must not contain a comma (`,`) - it is reserved as the id separator on `/context/status?ids=`. | + | | Context ID. Treated as the upsert key. Send an empty string to have one generated upstream. Must not contain a comma (`,`). It is reserved as the id separator on `/context/status?ids=`. | | | Target database. Must match the form-level `database`. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | | | Logical partition inside the database. Must match the form-level `collection`. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). | | | Short title or subject shown in search results. | @@ -565,13 +570,14 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can | | Canonical URL or reference link. | | | ISO-8601 timestamp (creation or last-updated). | | | Content payload. Use `{ "text": "..." }` for plain text. Required for app sources. | - | | Database-schema fields. (default=`{}`) | - | | Free-form per-document fields. (default=`{}`) | + | | Database-schema fields. At most 16 KiB. (default=`{}`) | + | | Free-form per-document fields. At most 1 KiB. (default=`{}`) | | | Optional related attachments. (default=`[]`) | - | | Forceful relations, same shape as on `metadata`. | + | | Forceful relations, same shape as on `document_metadata` items. | + | | Principals allowed to retrieve this source. Omit to leave it unrestricted. A malformed principal rejects the whole request with `400`. See [Access Control](/essentials/v2/access-control). | - + ```python Python SDK @@ -626,21 +632,19 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can | Field | Description | | --- | --- | - | | Optional unique ID. Acts as the upsert key. Must not contain a comma (`,`) - it is reserved as the id separator on `/context/status?ids=`. (default=auto-generated) | - | | Short label for display in `/context/list` and search hits as `source.title`. (default=truncated `text`) | + | | Optional unique ID. Acts as the upsert key. Must not contain a comma (`,`). It is reserved as the id separator on `/context/status?ids=`. (default=auto-generated) | + | | Short label for display in `/context/list` and search hits as `source.title`. (default=the memory's `id`) | | | Raw text or markdown content. Required unless `user_assistant_pairs` is provided. | | | Conversation pairs `{ user, assistant }`. Required unless `text` is provided. | | | Treat `text` as markdown for chunking. (default=`false`) | | | When `true`, HydraDB extracts the underlying preference from raw signal. (default=`false`) | | | Guides extraction when `infer: true`. Ignored when `infer: false`. | | | The user's name. Feeds inference. (default=`"User"`) | - | | TTL in seconds. Memory stops surfacing after expiry. | - | | Database-schema fields as a **JSON-stringified** object (e.g. `"{\"department\":\"legal\"}"`). Unlike `metadata` and `app_knowledge`, memory items take this as a string, not an object. (default=`""`) | - | | Free-form per-document fields as a **JSON-stringified** object. Same string-vs-object difference as `metadata` above. (default=`""`) | - | | Forceful relations within the Memories store. Shape: `{ "ids": ["...", "..."] }`. | + | | Database-schema fields as an object (e.g. `{ "department": "legal" }`), the same shape as on `document_metadata` and `app_knowledge` items. At most 16 KiB. (default=`{}`) | + | | Free-form per-memory fields as an object, like `metadata` above. At most 1 KiB. (default=`{}`) | - + ```python Python SDK @@ -700,8 +704,8 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can - - `graph_payload` is a **map of source id → graph** that **replaces LLM graph extraction** for each keyed source. For `type=knowledge`, the key is a `document_metadata` id or an `app_knowledge` item id; for `type=memory`, the key is a memory `id`. Keyed sources are still chunked and embedded, so they stay searchable. See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) for the full guide. + + `graph_payload` is a **map of source id to graph** that **replaces LLM graph extraction** for each keyed source. For `type=knowledge`, the key is a `document_metadata` id or an `app_knowledge` item id; for `type=memory`, the key is a memory `id`. Keyed sources are still chunked and embedded, so they stay searchable. See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) for the full guide. @@ -777,30 +781,30 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`) can | Field | Description | | --- | --- | | | Top-level key: a `document_metadata` id or `app_knowledge` item id for `type=knowledge`, or a memory `id` for `type=memory`. Value is that source's graph. A key matching no source returns `400`. | - | | Map keyed by a caller-local id; each value is an entity. The key is only a handle for `relations` to reference - it is not stored. | - | | Entity name. Normalized (lowercased) server-side so it matches at query time. ≤ 256 chars. | + | | Map keyed by a caller-local id; each value is an entity. The key is only a handle for `relations` to reference. It is not stored. | + | | Entity name. Normalized (lowercased) server-side so it matches at query time. At most 256 chars. | | | Entity type (e.g. `PERSON`, `POLICY`). Stored as supplied. | | | Logical grouping for the entity. Stored as supplied. | - | | Optional external id (email, URL, etc.) - display only. | + | | Optional external id (email, URL, etc.), display only. | | | Edges referencing entity-map keys. | | | Entity-map key of the source entity. | | | Entity-map key of the target entity. | - | | Relationship label, any plain string. ≤ 256 chars. | - | | Optional sentence supporting the edge. ≤ 2,000 chars. | + | | Relationship label, any plain string. At most 256 chars. | + | | Optional sentence supporting the edge. At most 2,000 chars. | | | Optional timing info (e.g. "since 2021", "in Q3"). | - **Per-source replace mode.** Each top-level key must match a `document_metadata` id or `app_knowledge` item id for `type=knowledge`, or a memory `id` for `type=memory`, in the same request; attach graphs to multiple sources at once. Extraction is skipped for keyed sources. Caps per graph: ≤ 5,000 entities, ≤ 10,000 relations, ≤ 500 relations per entity; over-cap returns `400`. Graphs survive re-ingest (re-upload or connector re-sync re-applies the stored graph). + **Per-source replace mode:** Each top-level key must match a `document_metadata` id or `app_knowledge` item id for `type=knowledge`, or a memory `id` for `type=memory`, in the same request; attach graphs to multiple sources at once. Extraction is skipped for keyed sources. Caps per graph: at most 5,000 entities, 10,000 relations, and 500 relations per entity; over-cap returns `400`. Graphs survive re-ingest (re-upload or connector re-sync re-applies the stored graph). ## Some important notes -- **Async indexing.** `202 Accepted` means HydraDB queued the work, not that content is searchable. Poll [Ingestion Status](/api-reference/v2/endpoint/source-status) until `indexing_status` reaches `graph_creation` (searchable) or `completed` (graph-ready). -- **Multipart, not JSON.** This endpoint uses `multipart/form-data`. Stringify all JSON arrays (`metadata`, `app_knowledge`, `memories`) before placing them in the form field. -- **Declare hot schema fields upfront.** Put frequently filtered fields in `metadata`, define them in `database_metadata_schema` with `enable_match: true`, and use `additional_metadata` for free-form display/bookkeeping fields. Define filterable fields when creating the database via [Create Database](/api-reference/v2/endpoint/create-tenant). Additive schema updates exist, but newly added dense/sparse metadata lanes are not backfilled into existing Milvus collections. -- **Memory vs knowledge.** Use `type: "memory"` for memory ingestion, listing, and deletion. Use `type: "all"` on `POST /query` when results should combine both. The multipart field name for memories is always `memories`. -- **Collection defaulting.** Omitting `collection` writes to the default collection. List available collections with [List Collections](/api-reference/v2/endpoint/list-sub-tenants). +- **Async indexing:** `202 Accepted` means HydraDB queued the work, not that content is searchable. Poll [Ingestion Status](/api-reference/v2/endpoint/source-status) until `indexing_status` reaches `graph_creation` (searchable) or `completed` (graph-ready). +- **Multipart, not JSON:** This endpoint uses `multipart/form-data`. Stringify all JSON arrays (`document_metadata`, `app_knowledge`, `memories`) before placing them in the form field. A body that is neither a form nor JSON returns `415`. +- **Declare hot schema fields upfront:** Put frequently filtered fields in `metadata`, define them in `database_metadata_schema`, and use `additional_metadata` for free-form display/bookkeeping fields. Define filterable fields when creating the database via [Create Database](/api-reference/v2/endpoint/create-tenant). Additive schema updates exist ([Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema)), but dense/sparse metadata lanes can only be declared at database creation. +- **Memory vs knowledge:** Use `type: "memory"` for memory ingestion, listing, and deletion. Use `type: "all"` on `POST /query` when results should combine both. The multipart field name for memories is always `memories`. +- **Collection defaulting:** Omitting `collection` writes to the default collection, created on first write. List available collections with [List Collections](/api-reference/v2/endpoint/list-sub-tenants).
diff --git a/api-reference/v2/endpoint/list-connector-providers.mdx b/api-reference/v2/endpoint/list-connector-providers.mdx index 2b0a7373..bf5ee85c 100644 --- a/api-reference/v2/endpoint/list-connector-providers.mdx +++ b/api-reference/v2/endpoint/list-connector-providers.mdx @@ -110,4 +110,4 @@ For Gmail, filter on `account_email` (`additional_metadata.account_email`) to sc ## Related Resources - [Create Connector](/api-reference/v2/endpoint/create-connector): connect a provider -- [Connectors - Overview](/api-reference/v2/endpoint/connectors-overview) +- [Connectors: Overview](/api-reference/v2/endpoint/connectors-overview) diff --git a/api-reference/v2/endpoint/list-documents.mdx b/api-reference/v2/endpoint/list-documents.mdx index ffc2e4f4..e6db571e 100644 --- a/api-reference/v2/endpoint/list-documents.mdx +++ b/api-reference/v2/endpoint/list-documents.mdx @@ -8,8 +8,8 @@ import { Field } from "/snippets/field.jsx"; Specify the category using the `type` parameter to filter and view ingested knowledge or user memories within a database or collection: -- `type=knowledge` _(default)_ - knowledge sources (documents, app sources). -- `type=memory` - user memories. +- `type=knowledge` _(default)_: knowledge sources (documents, app sources). +- `type=memory`: user memories. Supports pagination, metadata filters, and field projection. For metadata design and query-time behavior, see [Scoping using metadata](/essentials/v2/metadata). @@ -37,7 +37,7 @@ const sources = await client.context.list({ pageSize: 50, filters: { metadata: { department: "legal" }, - source_fields: { type: "slack" }, + sourceFields: { type: "slack" }, }, includeFields: ["title", "type", "timestamp", "additional_metadata"], }); @@ -70,17 +70,18 @@ curl -X POST 'https://api.hydradb.com/context/list' \ | | Owning database. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | | | Collection scope. If omitted, the default collection is used. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`null`) | | | Bucket to list. (default=`"knowledge"`) | -| | When provided and non-empty, only items with these IDs are returned (pagination \+ filters still apply). (default=`null`) | +| | When provided and non-empty, only items with these IDs are returned (pagination \+ filters still apply). At most `100` IDs. (default=`null`) | | | Page number (1-indexed). (default=`1`) | | | Items per page, from `1` to `100`. (default=`50`) | | | Structured exact-match filters. See [Filters](#1-filters). (default=`null`) | -| | Field projection. Only the listed fields plus `id`, `database`, `collection` are populated. Only applies to `type=knowledge`. (default=`null` - all fields) | +| | Field projection. Only the listed fields plus `id`, `database`, `collection` are populated. Only applies to `type=knowledge`. (default=`null`, all fields) | +| | Principals to list as (document ACLs). Only knowledge sources those principals may see come back. Omitted, `[]`, or `["*"]` lists everything. Ignored for `type=memory`. See [Access Control](/essentials/v2/access-control). | ### 1. Filters -- `filters` is a structured object with three optional categories. Filters are exact-match constraints i.e. filtered values are matched against stored values as exact values. There are no range, contains, or OR operators on this endpoint; run multiple calls and merge client-side for OR behavior. +- `filters` is a structured object with three optional categories. Filters are exact-match constraints i.e. filtered values are matched against stored values as exact values (except `source_fields.title`, a case-insensitive prefix match). There are no range, contains, or OR operators on this endpoint; run multiple calls and merge client-side for OR behavior. A `null` filter value returns `400`. - **AND/OR:** All filter pairs combine with a logical AND. To express OR semantics, run multiple calls and union them client-side. -- `ids `**\+ filters:** When `ids` is non-empty, only those IDs are considered, but other `filters` still apply on top - useful for "show me items 1, 2, 3 that also belong to department=legal". +- **`ids` \+ filters:** When `ids` is non-empty, only those IDs are considered, but other `filters` still apply on top: useful for "show me items 1, 2, 3 that also belong to department=legal". ```json { @@ -94,18 +95,18 @@ curl -X POST 'https://api.hydradb.com/context/list' \ | Category | Matched against | Notes | | --- | --- | --- | -| | Context item's schema-aligned `metadata` payload | Use for database metadata fields. `tenant_metadata` is accepted as a legacy alias. Keys must be declared in the database's `database_metadata_schema` with `enable_match: true`; undeclared keys are silently ignored. | +| | Context item's schema-aligned `metadata` payload | Use for database metadata fields. `tenant_metadata` is accepted as a legacy alias. Each key is an equality check on the stored `metadata`; none is ignored. In a database with a schema, ingest rejects undeclared keys, so filtering on one returns nothing (except `connector_id` and `provider`, which connectors always write). | | | Context item's `additional_metadata` payload | Free-form per-document JSON. No schema declaration required. `document_metadata` is accepted as a legacy alias. | -| | Built-in source fields (`type`, `title`, `description`, `url`, `timestamp`) | Use for app-source categories or quick title lookups. | +| | Built-in source fields (`type`, `title`, `description`, `url`, `timestamp`) and connector fields (`app_provider`, `app_kind`, `app_external_id`, `app_parent_id`) | Use for app-source categories or quick title lookups. `title` must be a string and matches by case-insensitive prefix: `"standup"` finds "Standup notes 2026-05-12". External IDs are unique only per provider, so pair `app_external_id` or `app_parent_id` with `app_provider`. Any other key returns `400`. | ### 2. Including Fields for convenient data objects When you don't need every field on every row, pass `include_fields` to keep response payloads small. Only the listed fields are populated; omitted fields should be treated as unavailable in that response. `id`, `database`, and `collection` are always returned. -Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadata`, `additional_metadata`, and `relations`. Omit or pass `null` to return everything. +Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadata`, `additional_metadata`, and `relations`, plus the legacy names `tenant_metadata` and `document_metadata`. Omit or pass `null` to return everything. - **Projectable vs. fetchable fields.** `content`, `url`, and `attachments` are **not** valid `include_fields` values - they are stripped from list responses, and requesting one returns `400`. Fetch them per-source via [Inspect Context](/api-reference/v2/endpoint/fetch-content). + **Projectable vs. fetchable fields:** `content`, `url`, and `attachments` are **not** valid `include_fields` values: they are stripped from list responses, and requesting one returns `400`. Fetch them per-source via [Inspect Context](/api-reference/v2/endpoint/fetch-content). `include_fields` only applies to `type=knowledge`. It is ignored for `type=memory`. @@ -117,12 +118,14 @@ Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadat "success": true, "data": { "success": true, - "message": "Sources retrieved successfully", + "message": "Successfully fetched sources", "sources": [ { "id": "policy_main", "database": "acme_corp", "collection": "team_docs", + "tenant_id": "acme_corp", + "sub_tenant_id": "team_docs", "title": "Compliance Policy", "type": "pdf", "timestamp": "2026-05-12T08:14:00Z", @@ -152,11 +155,17 @@ Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadat "success": true, "data": { "success": true, + "message": "Successfully fetched sources", "user_memories": [ { "memory_id": "mem_user_alex_tone", - "memory_content": "Prefers concise answers and dark mode.", - "inferred_content": "User prefers concise answers and dark mode." + "database": "acme_corp", + "collection": "team_docs", + "tenant_id": "acme_corp", + "sub_tenant_id": "team_docs", + "title": "Tone preference", + "type": "memory", + "metadata": { "user_id": "alex" } } ], "total": 1, @@ -194,7 +203,10 @@ Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadat -When `type=memory`, `data` is a `ListUserMemoriesResponse` instead - same idea but with a `data.user_memories[]` array of memory items. +When `type=memory`, `data` is a `ListUserMemoriesResponse` instead: same idea but with a `data.user_memories[]` array of memory items, keyed by `memory_id`, with the same fields as knowledge rows. The memory text and `inferred_content` are not included; read them with [Inspect Context](/api-reference/v2/endpoint/fetch-content). + +- **Scope fields:** every row carries `database` and `collection`, plus the deprecated `tenant_id` and `sub_tenant_id` with the same values. +- **Memory listing unavailable:** if the memory store cannot be read, the call still returns `200` with an empty `user_memories` and `message` set to `"Memories temporarily unavailable"`. Check `message` before you treat an empty list as "no memories". **Use the canonical v2 names.** Prefer `filters.metadata` and `filters.additional_metadata`. Legacy `filters.tenant_metadata` and `filters.document_metadata` are accepted for back-compat, with canonical keys winning on conflicts. diff --git a/api-reference/v2/endpoint/list-sub-tenants.mdx b/api-reference/v2/endpoint/list-sub-tenants.mdx index ca8e9288..dd0e4ebf 100644 --- a/api-reference/v2/endpoint/list-sub-tenants.mdx +++ b/api-reference/v2/endpoint/list-sub-tenants.mdx @@ -4,13 +4,13 @@ description: "List collection IDs inside a database." openapi: "api-reference/v2/openapi.json GET /databases/collections" --- -1. The default collection is not created until the first write - no collection exists until then. Once you ingest without an explicit `collection`, the default collection is created and stores all context written without a `collection`. Create additional collections at any time to scope data to users, teams, or projects. +1. The default collection is not created until the first write; no collection exists until then. Once you ingest without an explicit `collection`, the default collection is created and stores all context written without a `collection`. Create additional collections at any time to scope data to users, teams, or projects. 2. **Implicit creation.** Collections are auto-created when ingestion writes data under a new `collection`. The returned list grows organically as your application writes data under new values. ```python Python SDK -response = client.databases.collections(database="your database id") +response = client.databases.collections(database="my_first_database") ``` ```typescript TypeScript SDK @@ -81,7 +81,7 @@ curl -X GET 'https://api.hydradb.com/databases/collections?database=my_first_dat **Related Resources** - - **Inspect content:** [List Documents](/api-reference/v2/endpoint/list-documents) - scoped to a `collection` + - **Inspect content:** [List Documents](/api-reference/v2/endpoint/list-documents), scoped to a `collection` - **Delete a collection:** [Delete Collection](/api-reference/v2/endpoint/delete-collection) - **Inspect usage:** [Database Stats](/api-reference/v2/endpoint/tenant-stats) diff --git a/api-reference/v2/endpoint/list-tenants.mdx b/api-reference/v2/endpoint/list-tenants.mdx index fe12f3d7..c666bf86 100644 --- a/api-reference/v2/endpoint/list-tenants.mdx +++ b/api-reference/v2/endpoint/list-tenants.mdx @@ -1,10 +1,10 @@ --- title: "List Databases" -description: "List all databases created. " +description: "List all databases created." openapi: "api-reference/v2/openapi.json GET /databases" --- -The response separates active or provisioning databases (in `data.databases`) from databases whose provisioning failed (in `data.failed_databases`). Use [Database Status](https://docs.hydradb.com/api-reference/v2/endpoint/tenant-status) to confirm readiness before ingestion. +The response separates active or provisioning databases (in `data.databases`) from databases whose provisioning failed (in `data.failed_databases`). Use [Database Status](/api-reference/v2/endpoint/tenant-status) to confirm readiness before ingestion. @@ -94,5 +94,5 @@ curl -X GET 'https://api.hydradb.com/databases' \ - **Inspect:** [Database Status](/api-reference/v2/endpoint/tenant-status) - **Inspect:** [Database Stats](/api-reference/v2/endpoint/tenant-stats) - **Delete:** [Delete Database](/api-reference/v2/endpoint/delete-tenant) - - **Read more:** [Concepts → Multi-Tenant Support](/essentials/v2/multi-tenant) + - **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) \ No newline at end of file diff --git a/api-reference/v2/endpoint/list-webhook-deliveries.mdx b/api-reference/v2/endpoint/list-webhook-deliveries.mdx index 500ac22a..a4acb509 100644 --- a/api-reference/v2/endpoint/list-webhook-deliveries.mdx +++ b/api-reference/v2/endpoint/list-webhook-deliveries.mdx @@ -16,4 +16,6 @@ Filter by `status` to isolate failures, and page through results with `limit` an | `failed` | An attempt failed and will be retried. | | `permanently_failed` | Retries are exhausted. No further attempts. | +`limit` is 1 to 100 (default 20). To get the next page, pass the response's `next_cursor` as `cursor`. `next_cursor` is `null` on the last page. A `limit` outside that range returns `422`, and an unknown `status` returns `400`. + See [Webhooks](/essentials/v2/webhooks) for retry behaviour. diff --git a/api-reference/v2/endpoint/query-overview.mdx b/api-reference/v2/endpoint/query-overview.mdx index 9fd25715..bdcc3092 100644 --- a/api-reference/v2/endpoint/query-overview.mdx +++ b/api-reference/v2/endpoint/query-overview.mdx @@ -1,5 +1,5 @@ --- -title: "Query - Overview" +title: "Query: Overview" description: "Quick reference for query modes, type selection, and when to call each." --- @@ -35,18 +35,18 @@ linkStyle default stroke:#64748b,stroke-width:2px; | Parameter | Values | Use it for | |---|---|---| -| | `"knowledge"`, `"memory"`, `"all"` | Choose the collection. Use `"knowledge"` for shared docs/app sources, `"memory"` for user context, and `"all"` when an answer should use both. | +| | `"knowledge"`, `"memory"`, `"all"` | Choose what to query. Use `"knowledge"` for shared docs/app sources, `"memory"` for user context, and `"all"` when an answer should use both. | | | `"hybrid"`, `"text"` | Choose the matching method. Use `"hybrid"` by default and `"text"` for exact terms or phrases. | -| | `"fast"`, `"thinking"`, `"auto"` | Choose latency vs quality, or let HydraDB decide. Use `"fast"` for low-latency paths, `"thinking"` for multi-query retrieval, reranking, and forceful-relation context, and `"auto"` to score the query and route to one of the two automatically (defaults to `"thinking"` when the signal is inconclusive; also overrides `graph_context` to match - **the default if `mode` is omitted**). | -| | integer | Control prompt size. Start with `10`, reduce for tight context windows, increase only when you rerank or summarize downstream. | -| | `0.0`-`1.0` or `"auto"` | Tune hybrid query. Lower values favor BM25 keywords; higher values favor semantic similarity. | +| | `"fast"`, `"thinking"`, `"auto"` | Choose latency vs quality, or let HydraDB decide. Use `"fast"` for low-latency paths, `"thinking"` for multi-query retrieval, reranking, and forceful-relation context, and `"auto"` to score the query and route to one of the two automatically (defaults to `"thinking"` when the signal is inconclusive; **the default if `mode` is omitted**). | +| | integer | Control prompt size. Default `10`, maximum `250`. Start with `10`, reduce for tight context windows, increase only when you rerank or summarize downstream. | +| | `0.0`-`1.0` or `"auto"` | Tune hybrid query. Lower values favor BM25 keywords; higher values favor semantic similarity. Default `0.8`, which `"auto"` also resolves to. | | | object | Narrow candidates before ranking. Top-level keys match `metadata`; nested `additional_metadata` filters free-form per-source fields. | -| | `string[]` or weighted object | Query one or more user/workspace/team scopes. A list uses equal normalized weights; an object like `{ "workspace_42": 2, "user_alex": 1 }` applies relative ranking weights with at most one decimal place. Max 100 collections. | -| | boolean | Include entity/relation context with the chunks. On by default; set `false` for chunk-only responses. | -| | boolean | Adds app-aware retrieval while still querying the full selected knowledge scope. Use it for better app-source matching; it does not restrict retrieval to app sources only. | +| | `string[]` or weighted object | Query one or more user/workspace/team scopes. A list uses equal normalized weights; an object like `{ "workspace_42": 2, "user_alex": 1 }` applies relative ranking weights with at most one decimal place. Max 100 collections, and each must exist. | +| | boolean | Include entity/relation context with the chunks. On by default; set `false` with `mode: "fast"` for chunk-only responses (`thinking` always includes it). | +| | boolean | Adds app-aware retrieval while still querying the full selected knowledge scope. Use it for better app-source matching; it does not restrict retrieval to app sources only. On by default for knowledge queries; send `false` to turn it off. | -For filter design, read [Usage - Metadata](/essentials/v2/metadata) before creating database schemas. For exact request fields, defaults, and response shape, use [Query](/api-reference/v2/endpoint/query). +For filter design, read [Usage: Metadata](/essentials/v2/metadata) before creating database schemas. For exact request fields, defaults, and response shape, use [Query](/api-reference/v2/endpoint/query). ## Recommended configurations @@ -59,7 +59,7 @@ For filter design, read [Usage - Metadata](/essentials/v2/metadata) before creat | User preferences only | `type="memory"`, include `collection`, `query_by="hybrid"` | | Exact keyword or phrase | `type="knowledge"`, `query_by="text"`, `operator="phrase"` | | Recent operational updates | `query_by="hybrid"`, `recency_bias=0.2-0.4`, filter to the right document type | -| Mixed or unpredictable query complexity | `query_by="hybrid"`, `mode="auto"` - let HydraDB route each query to `fast` or `thinking` | +| Mixed or unpredictable query complexity | `query_by="hybrid"`, `mode="auto"`: let HydraDB route each query to `fast` or `thinking` | ## Typical patterns @@ -121,7 +121,7 @@ Use this when the same question should search several collection scopes and retu -Use `metadata_filters` when you already know the slice you want. Top-level keys match schema-backed `metadata` fields; declare hot filters in `database_metadata_schema` with `enable_match: true`. Free-form per-source fields go under `additional_metadata` (`document_metadata` is a legacy alias). Multiple filters are ANDed exact-match constraints. +Use `metadata_filters` when you already know the slice you want. Top-level keys match schema-backed `metadata` fields; declare hot filters in `database_metadata_schema`. Free-form per-source fields go under `additional_metadata` (`document_metadata` is a legacy alias). Multiple filters are ANDed constraints. ```json { @@ -163,7 +163,7 @@ Use text query when literal wording matters: legal clauses, SKUs, error codes, I ## Related sections -- [Query](/api-reference/v2/endpoint/query) - full endpoint reference -- [Usage - Query](/essentials/v2/query) - conceptual overview, retrieval modes, and ranking behavior -- [Usage - Metadata](/essentials/v2/metadata) - filtering with database and document metadata -- [Concepts - Context Graphs](/essentials/v2/context-graphs) - graph context and relation paths +- [Query](/api-reference/v2/endpoint/query): full endpoint reference +- [Usage: Query](/essentials/v2/query): conceptual overview, retrieval modes, and ranking behavior +- [Usage: Metadata](/essentials/v2/metadata): filtering with database and document metadata +- [Concepts: Context Graphs](/essentials/v2/context-graphs): graph context and relation paths diff --git a/api-reference/v2/endpoint/query.mdx b/api-reference/v2/endpoint/query.mdx index f5258821..2ac2318e 100644 --- a/api-reference/v2/endpoint/query.mdx +++ b/api-reference/v2/endpoint/query.mdx @@ -11,8 +11,8 @@ The single retrieval endpoint for everything. Use it any time you need to feed a Three independent dimensions control behavior: - **`type`** picks **what** to query: `"knowledge"`, `"memory"`, or `"all"` (both, merged and re-ranked together). -- **`query_by`** picks **how** to match: `"hybrid"` (semantic + BM25, the default) or `"text"` (BM25 only - pair with `operator`). -- **`mode`** picks **how** to rank results: `"fast"` (single-pass, low-latency), `"thinking"` (expands query, reranks, and can include forceful-relation context), or `"auto"` (scores the query and routes to `"fast"` or `"thinking"` automatically, defaulting to `"thinking"` when the signal is inconclusive - **the default if `mode` is omitted**). +- **`query_by`** picks **how** to match: `"hybrid"` (semantic + BM25, the default) or `"text"` (BM25 only; pair with `operator`). +- **`mode`** picks **how** to rank results: `"fast"` (single-pass, low-latency), `"thinking"` (expands query, reranks, and can include forceful-relation context), or `"auto"` (scores the query and routes to `"fast"` or `"thinking"` automatically, defaulting to `"thinking"` when the signal is inconclusive; **the default if `mode` is omitted**). Read more about choosing the perfect mode for your use case [here](/api-reference/v2/endpoint/query-overview#recommended-configurations). @@ -152,7 +152,7 @@ Use `collections` when one query should fan out across multiple user, workspace, } ``` -A list gives every collection equal normalized weight. An object treats values as positive relative ranking weights with at most one decimal place and normalizes them server-side. You can send at most 100 collections. When `max_results` is omitted, HydraDB uses up to 10 results per collection, capped at 1000 fanout candidates before the final ranked response is shaped. When `max_results` is set, it is the final global response cap across the merged fanout result set. +A list gives every collection equal normalized weight. An object treats values as positive relative ranking weights with at most one decimal place and normalizes them server-side. You can send at most 100 collections. When `max_results` is omitted, HydraDB uses up to 10 results per collection, capped at 1000 fanout candidates before the final ranked response is shaped. When `max_results` is set, it is the final global response cap across the merged fanout result set. Every listed collection must exist, or the call returns `400 INVALID_INPUT` naming the missing ones. > **Caching tip:** `collections` list order is not semantically significant for fanout selection. Sort list values before constructing cache keys; for weighted objects, sort keys and keep weights at the documented one-decimal precision so equivalent calls share the same cache entry. @@ -207,7 +207,7 @@ const context = buildString(result); ## Common use-cases and their configurations - + ```bash cURL @@ -253,7 +253,7 @@ result = client.query( - + ```bash cURL @@ -296,7 +296,7 @@ result = client.query( - + ```bash cURL @@ -339,7 +339,7 @@ result = client.query( - + ```bash cURL @@ -379,7 +379,7 @@ result = client.query( - + ```bash cURL @@ -418,7 +418,7 @@ result = client.query( - HydraDB scores the query before retrieval and routes it to `"fast"` or `"thinking"` - a query naming several distinct entities like this one is likely to route to `"thinking"`. Use `"auto"` for traffic where query complexity varies call-to-call and you don't want to hand-pick per request. This is also the default: an omitted `mode` field behaves exactly like `mode: "auto"`. Set `mode` to `"fast"` or `"thinking"` explicitly if you want a deterministic pipeline instead. + HydraDB scores the query before retrieval and routes it to `"fast"` or `"thinking"`: a query naming several distinct entities like this one is likely to route to `"thinking"`. Use `"auto"` for traffic where query complexity varies call-to-call and you don't want to hand-pick per request. This is also the default: an omitted `mode` field behaves exactly like `mode: "auto"`. Set `mode` to `"fast"` or `"thinking"` explicitly if you want a deterministic pipeline instead. @@ -430,24 +430,24 @@ result = client.query( | | Single collection scope. Required for per-user memory queries. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=default collection) | | | Multi-collection scope. Send a list of collection IDs for equal weighting, or an object mapping collection ID to a positive relative weight (at most one decimal place, e.g. `{"finance": 1.5, "legal": 0.8}`) to bias ranking. Up to 100 collections. Do not combine with `collection`/`sub_tenant_id`. Formerly `sub_tenant_ids`; the `sub_tenant_ids` alias is still accepted (deprecated since 2.0.1). | | | Query terms or natural-language question. Cannot be empty. | -| | What collection to query. `"all"` runs knowledge and memory in parallel and merges by `relevancy_score`. (default=`"knowledge"`) | +| | What to query. `"all"` runs knowledge and memory in parallel over the same scope and merges by `relevancy_score`. (default=`"knowledge"`) | | | Retrieval method. See [Query methods](#decision-matrix). (default=`"hybrid"`) | -| | Adds an app-aware retrieval lane for app sources while still querying the full selected knowledge scope. Set `true` for better app-source matching, thread/relation traversal, exact IDs, and actor/provider hints. It does **not** limit query to only app sources. (default=`false`) | -| | BM25 operator for `query_by: "text"`. Ignored for `hybrid`. (default=`"or"`) | -| | Retrieval pipeline. Applies to `hybrid` only; ignored for `text`. `"auto"` scores the query before retrieval and resolves it to `"fast"` or `"thinking"`, defaulting to `"thinking"` when the signal is inconclusive; it also overrides whatever `graph_context` you sent to match that resolved mode. (default=`"auto"`) | -| | Maximum chunks to return. Default `10`; maximum `50`. Start with `10`, use `5` for tight prompts, and increase only when reranking downstream. | -| | Hybrid weight (`1.0` = pure semantic, `0.0` = pure BM25). Applies to `query_by: "hybrid"` only. (default=`0.8`) | -| | Boost newer content. (default=`0.0`) | -| | When `true`, includes the entity/relation graph slice in the response under `graph_context`. Set to `false` when you only need ranked chunks. Relations you supplied via [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) appear here identically to extracted ones. (default=`true`) - **under `mode: "auto"`, this value is overridden by the resolved mode regardless of what you send.** | -| | Pull author-declared related sources into `additional_context`. **Only takes effect when `mode` resolves to `"thinking"`** - silently ignored in `fast` mode, and under `mode: "auto"` whether it takes effect depends on the automatic routing decision. (default=`true`) | +| | Applies to knowledge queries only. Adds an app-aware retrieval lane for app sources while still querying the full selected knowledge scope, for better app-source matching, thread/relation traversal, exact IDs, and actor/provider hints. It does **not** limit query to only app sources. Send `false` to turn the lane off. (default=`true`) | +| | BM25 operator for `query_by: "text"`. `"and"` or `"phrase"` with any other `query_by` returns `400`. (default=`"or"`) | +| | Retrieval pipeline. Applies to `hybrid` only; ignored for `text`. `"auto"` scores the query before retrieval and resolves it to `"fast"` or `"thinking"`, defaulting to `"thinking"` when the signal is inconclusive. (default=`"auto"`) | +| | Maximum chunks to return. Default `10`; maximum `250`. Start with `10`, use `5` for tight prompts, and increase only when reranking downstream. | +| | Hybrid weight (`1.0` = pure semantic, `0.0` = pure BM25). Applies to `query_by: "hybrid"` only. `"auto"` currently resolves to `0.8`. (default=`0.8`) | +| | Boost newer content. Omitted applies a mild `0.4` tilt; send `0` to turn it off. (default=`0.4`) | +| | When `true`, includes the entity/relation graph slice in the response under `graph_context`. Set to `false` when you only need ranked chunks. Relations you supplied via [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) appear here identically to extracted ones. (default=`true`). **`false` only takes effect in `fast` mode: `thinking`, and `auto` when it resolves to thinking, always include the graph slice.** | +| | Pull author-declared related sources into `additional_context`. **Only takes effect when `mode` resolves to `"thinking"`;** silently skipped in `fast` mode, and under `mode: "auto"` whether it takes effect depends on the automatic routing decision. (default=`true`) | | | Request-time hint to guide retrieval (e.g., "user is on the billing page"). This is different from the response `additional_context` map. (default=`null`) | | | Deterministic narrowing before ranking. See [Filters](#decision-matrix). Each list holds at most 500 values, and the whole object is capped at 64 KiB of compact JSON, measured after operator objects are reduced to their values; over either returns `400`. (default=`null`) | **Tuning heuristics.**
    -
  • alpha: start at 0.8. Lower toward 0.3–0.5 when the query contains literal tokens (error codes, SKUs, product names). Raise toward 0.9 for conceptual questions. Use "auto" when query shape varies.
  • -
  • recency_bias: leave at 0 for static reference material. Set 0.2–0.4 for mixed content, 0.6–0.8 for changelogs, news, or status updates.
  • +
  • alpha: start at 0.8. Lower toward 0.3 to 0.5 when the query contains literal tokens (error codes, SKUs, product names). Raise toward 0.9 for conceptual questions.
  • +
  • recency_bias: set to 0 for static reference material. Set 0.2 to 0.4 for mixed content, 0.6 to 0.8 for changelogs, news, or status updates.
  • max_results: start at 10. Drop to 5 for tight context windows; raise to 20 if you rerank downstream.
@@ -458,8 +458,7 @@ result = client.query( | Value | Queries | Best for | |---|---|---| - | `"knowledge"` *(default)* | Knowledge documents, files, and app sources | Document Q&A, RAG context. | - | `"query_apps=true"` | Full selected knowledge scope plus app-aware retrieval | App-specific Q&A that should still query non-app knowledge documents. | + | `"knowledge"` *(default)* | Knowledge documents, files, and app sources, plus an app-aware lane unless `query_apps` is `false` | Document Q&A, RAG context. | | `"memory"` | User memories | Personalization and user preferences. | | `"all"` | Both, merged in one ranked result set | Personalized answers grounded in both shared and user-specific context. | @@ -478,13 +477,13 @@ result = client.query( |---|---|---| | `"fast"` | Single query pass | Real-time chat, autocomplete, simple lookups. | | `"thinking"` | Multi-query expansion + reranking + forceful-relation context | Complex queries, customer-facing answers, anything where quality matters. | - | `"auto"` *(default if `mode` is omitted)* | Scores the query before retrieval and routes to `"fast"` or `"thinking"`; defaults to `"thinking"` when the signal is inconclusive. Also overrides `graph_context` to match whichever mode it picks. | Mixed or unpredictable query traffic where you don't want to hand-pick per request. | + | `"auto"` *(default if `mode` is omitted)* | Scores the query before retrieval and routes to `"fast"` or `"thinking"`; defaults to `"thinking"` when the signal is inconclusive. | Mixed or unpredictable query traffic where you don't want to hand-pick per request. | - `"auto"`'s resolved pipeline isn't reported back in the response, so budget latency as thinking-level in the worst case. Omitting `mode` behaves exactly like `mode: "auto"` - set it explicitly to `"fast"` or `"thinking"` if you want a deterministic pipeline instead. + `"auto"`'s resolved pipeline isn't reported back in the response, so budget latency as thinking-level in the worst case. Omitting `mode` behaves exactly like `mode: "auto"`; set it explicitly to `"fast"` or `"thinking"` if you want a deterministic pipeline instead.
- `metadata_filters` are hard exact-match constraints applied before ranking and re-checked after hydration. The shape combines two filter scopes: + `metadata_filters` are hard constraints applied before ranking and re-checked after hydration. The shape combines two filter scopes: ```json { @@ -501,7 +500,7 @@ result = client.query( | Where | What it matches | |---|---| - | **Top-level keys** (`department`, `region`) | The source's schema-backed `metadata`. Keys must be declared in the database's `database_metadata_schema` with `enable_match: true`, otherwise they are silently ignored. | + | **Top-level keys** (`department`, `region`) | The source's schema-backed `metadata`. The filter matches stored values, not the schema. In a database with a schema, ingest rejects undeclared keys, so filtering on one returns nothing, except `connector_id` and `provider`, which connectors always write. Declare filter keys in the database's `database_metadata_schema`. | | **Nested under `additional_metadata`** | The source's free-form per-document fields. No schema declaration required. `document_metadata` is accepted as a legacy alias. | Separate keys are ANDed. Each `metadata` (top-level) key takes an operator object naming the comparison: @@ -529,12 +528,12 @@ result = client.query(
- The bare forms still work and are unchanged, but are **deprecated** in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as `equals`, a bare array as `contains_any`, and a bare single-element array as `contains` - so `{"emails": "a@x"}` and `{"emails": ["a@x"]}` differ by one character and return different results. + The bare forms still work and are unchanged, but are **deprecated** in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as `equals`, a bare array as `contains_any`, and a bare single-element array as `contains`, so `{"emails": "a@x"}` and `{"emails": ["a@x"]}` differ by one character and return different results. `contains`, `contains_any` and lists are supported on `VARCHAR` fields only: any of them passed for a declared field of another type is rejected with `400 VALIDATION_ERROR`. `equals` works on every declared type, so `{"priority": {"equals": 7}}` is valid on an `INT64` field. - A known operator given the wrong operand type, or several operators in one object, is rejected with `400 VALIDATION_ERROR`. A **misspelled** operator is not: `{"contian": "x"}` is indistinguishable from a filter for a stored object with that key, so it is left alone and matches nothing. + A known operator given the wrong operand type, or several operators in one object, is rejected with `400 INVALID_INPUT`, and so is a `null` filter value. A **misspelled** operator is not: `{"contian": "x"}` is indistinguishable from a filter for a stored object with that key, so it is left alone and matches nothing. `contains`, `contains_any` and `equals` are reserved key names: an object built only from them is read as an operator and can no longer exact-match a stored object, and an object whose keys are ALL operator names is rejected with `400`. Mixing an operator name with any other key (`{"contains": "a", "other": 1}`) is unaffected. A caller needing the reserved shape must rename the nested key or the field. @@ -704,8 +703,8 @@ result = client.query( "success": false, "data": null, "error": { - "code": "INVALID_PARAMETERS", - "message": "query must not be empty" + "code": "INVALID_INPUT", + "message": "query cannot be empty" }, "meta": { "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", @@ -722,24 +721,24 @@ A zero-result query returns empty arrays/maps rather than an error, as shown in **Default Behaviors** -- **`mode` defaults to `"auto"`.** Omitting `mode` entirely behaves exactly like `mode: "auto"` - set it explicitly to `"fast"` or `"thinking"` if you want a deterministic pipeline. -- **`graph_context` is on by default.** Set it to `false` if you only need ranked chunks and want to drop the graph slice from the response. -- **`recency_bias` is off by default.** Defaults to `0.0` - no recency boost is applied unless you set it. +- **`mode` defaults to `"auto"`.** Omitting `mode` entirely behaves exactly like `mode: "auto"`; set it explicitly to `"fast"` or `"thinking"` if you want a deterministic pipeline. +- **`graph_context` is on by default.** Set it to `false` if you only need ranked chunks and want to drop the graph slice from the response. `false` only takes effect in `fast` mode. +- **`recency_bias` applies a mild tilt by default.** Defaults to `0.4`; send `0` to turn the recency boost off. **Important Considerations & Common Mistakes** -- **`query_forceful_relations` requires `mode` to resolve to `"thinking"`.** In `fast` mode the flag is silently ignored. The server does not error or warn - your `additional_context` will simply be empty. Under `mode: "auto"` this depends on that request's routing decision, not on what you asked for. -- **`mode: "auto"` overrides `graph_context`.** Whatever you send for `graph_context` is replaced to match the resolved mode - `true` if auto escalates to `thinking`, `false` if it resolves to `fast`. This also applies when `mode` is omitted, since it defaults to `"auto"`. Set `graph_context` explicitly only when calling `"fast"` or `"thinking"` directly. -- **Want a deterministic pipeline instead of automatic routing?** Set `mode` explicitly to `"fast"` or `"thinking"` - an omitted `mode` field now defaults to `"auto"`, not `"fast"`. -- **Relation `timestamp` is a Unix epoch float here.** In the `graph_context` slice returned by `/query` - and in the passthrough relations returned by [List Documents](/api-reference/v2/endpoint/list-documents) with `include_fields: ["relations"]` - each relation's `timestamp` is a Unix epoch value in seconds (a float, e.g. `1778573640.0`). The dedicated [Context Relations](/api-reference/v2/endpoint/source-relations) endpoint returns the same field as an ISO-8601 string instead. Normalize before comparing relation timestamps across endpoints. -- **Use the right metadata namespace.** Top-level `metadata_filters` keys match `metadata`; free-form per-document fields must be nested under `additional_metadata` (`document_metadata` is only a legacy alias). Declare hot top-level filter fields in `database_metadata_schema` with `enable_match: true`. -- **Common mistakes.** Check [Ingestion Status](/api-reference/v2/endpoint/source-status) for recently ingested documents before querying. If you omit `collection`, HydraDB queries the default collection; use [List Collections](/api-reference/v2/endpoint/list-sub-tenants) to discover available IDs. +- **`query_forceful_relations` requires `mode` to resolve to `"thinking"`.** In `fast` mode the flag is silently skipped. The server does not error or warn; your `additional_context` will simply be empty. Under `mode: "auto"` this depends on that request's routing decision, not on what you asked for. +- **`graph_context: false` only takes effect in `fast` mode.** `thinking` always includes the graph slice, so under `mode: "auto"` whether `false` is honored depends on that request's routing decision. This also applies when `mode` is omitted, since it defaults to `"auto"`. Call `"fast"` explicitly if you never want the graph slice. +- **Want a deterministic pipeline instead of automatic routing?** Set `mode` explicitly to `"fast"` or `"thinking"`; an omitted `mode` field now defaults to `"auto"`, not `"fast"`. +- **Relation `timestamp` is a Unix epoch float here.** In the `graph_context` slice returned by `/query` (and in the passthrough relations returned by [List Documents](/api-reference/v2/endpoint/list-documents) with `include_fields: ["relations"]`), each relation's `timestamp` is a Unix epoch value in seconds (a float, e.g. `1778573640.0`). The dedicated [Context Relations](/api-reference/v2/endpoint/source-relations) endpoint returns the same field as an ISO-8601 string instead. Normalize before comparing relation timestamps across endpoints. +- **Use the right metadata namespace.** Top-level `metadata_filters` keys match `metadata`; free-form per-document fields must be nested under `additional_metadata` (`document_metadata` is only a legacy alias). Declare hot top-level filter fields in `database_metadata_schema`. +- **Common mistakes.** Check [Ingestion Status](/api-reference/v2/endpoint/source-status) for recently ingested documents before querying. If you omit `collection`, HydraDB queries the default collection, and a query that names a collection does not read the default one; use [List Collections](/api-reference/v2/endpoint/list-sub-tenants) to discover available IDs. ## Errors -Common codes: `400 INVALID_PARAMETERS` (empty `query`), `404 DATABASE_NOT_FOUND`, `422 VALIDATION_ERROR`, `500 INTERNAL_ERROR`. See [Error Responses](/api-reference/v2/error-responses) for the full list. +Common codes: `400 INVALID_INPUT` (empty `query`, `operator: "and"` or `"phrase"` without `query_by: "text"`, a malformed filter operator, or a listed collection that does not exist), `400 VALIDATION_ERROR` (a metadata filter that does not fit the declared field type), `404 DATABASE_NOT_FOUND`, `422 TENANT_INFRA_NOT_READY` (the database is still provisioning; poll [Database Status](/api-reference/v2/endpoint/tenant-status) until `ready_for_ingestion` is `true`), `500 INTERNAL_ERROR`. See [Error Responses](/api-reference/v2/error-responses) for the full list. `400` also covers oversized filters: a `metadata_filters` list above 500 values, or a `metadata_filters` object above 64 KiB of compact JSON. The message names the @@ -751,12 +750,12 @@ offending key or reports the actual byte count. See **Related Resources** -- **Setup first:** [Ingest Context](/api-reference/v2/endpoint/ingest-context) - content must be indexed -- **Confirm indexing:** [Ingestion Status](/api-reference/v2/endpoint/source-status) - wait for `completed` (or `graph_creation`) -- **Graph follow-up:** [Context Relations](/api-reference/v2/endpoint/source-relations) - inspect relationships in detail -- **Concepts:** [Usage → Query](/essentials/v2/query) -- **Concepts:** [Concepts → Semantic Search](/essentials/v2/semantic-search) -- **Concepts:** [Concepts → Context Graphs](/essentials/v2/context-graphs) -- **Response handling:** [Usage → How to Use API Results](/essentials/v2/api-results) -- **Read more:** [Query - Overview](/api-reference/v2/endpoint/query-overview) +- **Setup first:** [Ingest Context](/api-reference/v2/endpoint/ingest-context): content must be indexed +- **Confirm indexing:** [Ingestion Status](/api-reference/v2/endpoint/source-status): wait for `completed` (or `graph_creation`) +- **Graph follow-up:** [Context Relations](/api-reference/v2/endpoint/source-relations): inspect relationships in detail +- **Concepts:** [Usage: Query](/essentials/v2/query) +- **Concepts:** [Concepts: Semantic Search](/essentials/v2/semantic-search) +- **Concepts:** [Concepts: Context Graphs](/essentials/v2/context-graphs) +- **Response handling:** [Usage: How to Use API Results](/essentials/v2/api-results) +- **Read more:** [Query: Overview](/api-reference/v2/endpoint/query-overview) diff --git a/api-reference/v2/endpoint/register-webhook.mdx b/api-reference/v2/endpoint/register-webhook.mdx index 27d0d798..7a48f0e4 100644 --- a/api-reference/v2/endpoint/register-webhook.mdx +++ b/api-reference/v2/endpoint/register-webhook.mdx @@ -10,4 +10,6 @@ Registers the endpoint HydraDB calls when ingested content reaches a terminal in Omitting `signing_secret` **preserves** any secret you already have. Editing the URL or the event list never changes your signing configuration. To disable signing, call `DELETE /webhooks/indexing/signing-secret` explicitly. See [Manage the signing secret](/essentials/v2/webhooks#manage-the-signing-secret). +To have HydraDB create a signing secret, send `generate_signing_secret: true`. The secret comes back once, in this response, and cannot be read again. Send `generate_signing_secret` or `signing_secret`, not both. + Your endpoint must be reachable over public HTTPS. Localhost and private network addresses are rejected. See [Webhooks](/essentials/v2/webhooks) for the payload format, retry behaviour, and receiver examples. diff --git a/api-reference/v2/endpoint/retry-webhook-delivery.mdx b/api-reference/v2/endpoint/retry-webhook-delivery.mdx index 836491eb..467ca14a 100644 --- a/api-reference/v2/endpoint/retry-webhook-delivery.mdx +++ b/api-reference/v2/endpoint/retry-webhook-delivery.mdx @@ -4,7 +4,7 @@ description: "Queue a failed webhook delivery to be attempted again." openapi: "api-reference/v2/openapi.json POST /webhooks/indexing/deliveries/{delivery_id}/retry" --- -Queues a failed delivery for another attempt. Use it after fixing the problem on your side, such as a receiver that was down or was rejecting valid signatures. +Queues a `failed` or `permanently_failed` delivery for another attempt. Use it after fixing the problem on your side, such as a receiver that was down or was rejecting valid signatures. For a delivery in any other state, the call returns `queued: false` and nothing is retried. The retry is signed with your **current** signing secret, not the one in force when the delivery was first attempted. If you have rotated since, your receiver must know the new secret. diff --git a/api-reference/v2/endpoint/source-relations.mdx b/api-reference/v2/endpoint/source-relations.mdx index ca819275..3eabe8aa 100644 --- a/api-reference/v2/endpoint/source-relations.mdx +++ b/api-reference/v2/endpoint/source-relations.mdx @@ -47,8 +47,9 @@ curl -G 'https://api.hydradb.com/context/relations' \ | | When provided, returns relations for that specific source. When omitted, returns all relations across the collection. (default=`null`) | | | Bucket selector. Use `"memory"` when `id` belongs to a memory item. (default=`"knowledge"`) | | | Collection scope. If omitted, the default collection is used. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`null`) | -| | Maximum relation groups to return. Range `1–10000`. (default=`5000`) | +| | Maximum relation groups to return. Range `1` to `10000`. (default=`5000`) | | | Opaque pagination cursor from a previous response's `next_cursor`. (default=`null`) | +| | Principals to answer as (document ACLs). Only relations from sources those principals may see are returned. Repeated (`acl=a&acl=b`) or comma-separated. Omit for no ACL scoping. See [Access Control](/essentials/v2/access-control). | @@ -63,14 +64,16 @@ curl -G 'https://api.hydradb.com/context/relations' \ "type": "Service", "namespace": "default", "entity_id": "entity_payments_worker", - "identifier": null + "identifier": null, + "provider": "" }, "target": { "name": "OrdersDB", "type": "Database", "namespace": "default", "entity_id": "entity_orders_db", - "identifier": null + "identifier": null, + "provider": "" }, "relations": [ { @@ -89,10 +92,12 @@ curl -G 'https://api.hydradb.com/context/relations' \ "chunk_id": "policy_main_chunk_3" } ], + "auxiliary_relations": [], + "auxiliary_truncated": false, "is_truncated": false, "next_cursor": null, "success": true, - "message": "Relations retrieved successfully" + "message": "Successfully fetched relations for source" }, "error": null, "meta": { @@ -112,30 +117,41 @@ curl -G 'https://api.hydradb.com/context/relations' \ "name": "PaymentsWorker", "type": "Service", "namespace": "default", - "entity_id": "entity_payments_worker" + "entity_id": "entity_payments_worker", + "identifier": null, + "provider": "" }, "target": { "name": "OrdersDB", "type": "Database", "namespace": "default", - "entity_id": "entity_orders_db" + "entity_id": "entity_orders_db", + "identifier": null, + "provider": "" }, "relations": [ { "canonical_predicate": "DEPENDS_ON", "raw_predicate": "depends on", "context": "PaymentsWorker depends on OrdersDB for transaction sync.", + "confidence": 0.88, + "temporal_details": null, + "timestamp": "2026-05-12T08:14:00Z", "relationship_id": "rel_payments_orders", - "confidence": 0.88 + "chunk_id": "policy_main_chunk_4", + "source_entity_id": "entity_payments_worker", + "target_entity_id": "entity_orders_db" } ], "chunk_id": "policy_main_chunk_4" } ], + "auxiliary_relations": [], + "auxiliary_truncated": false, "is_truncated": true, - "next_cursor": 0.88, + "next_cursor": 1778573640.0, "success": true, - "message": "Relations retrieved successfully" + "message": "Successfully fetched relations for source" }, "error": null, "meta": { @@ -150,8 +166,8 @@ curl -G 'https://api.hydradb.com/context/relations' \ "success": false, "data": null, "error": { - "code": "SOURCE_NOT_FOUND", - "message": "Source not found" + "code": "DATABASE_NOT_FOUND", + "message": "Database 'acme_corp' not found. Use GET /databases to list active databases." }, "meta": { "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", @@ -184,21 +200,27 @@ while True: ## Some additional notes - **Cursor opacity.** `next_cursor` is opaque (currently a numeric score). Don't construct it client-side or assume meaning - pass back exactly what the server returned. + **Cursor opacity:** `next_cursor` is opaque (currently the timestamp, in Unix seconds, of the last group returned). Don't construct it client-side or assume meaning: pass back exactly what the server returned. - **Collection-wide queries:** Omitting `id` returns relations across the entire collection. This is useful for full-graph exports; pair with a small `limit` and paginate. - **Knowledge vs memory:** If the `id` belongs to a memory, set `type=memory`; otherwise the endpoint searches the Knowledge graph. The two graphs are completely separate. -- **Ordering:** Treat `data.relations[]` as ranked by relevance within the response. Preserve order for display or LLM context, but do not compare ordering across unrelated queries as an absolute signal. +- **Ordering:** `data.relations[]` comes back newest first, by each group's latest relation `timestamp`, not by relevance. +- **Unknown `id`:** returns `200` with an empty `relations` list, not an error. So does an `id` the `acl` principals may not see. +- **`auxiliary_relations`:** the structural graph around the sources (who sent a message, which comments hang off it), in the same triplet shape. It does not count against `limit`; `auxiliary_truncated` is `true` when it was clipped. - **Graph completeness:** Source relations only fully populate once the source's `indexing_status` reaches `completed`. Items in `graph_creation` are searchable but their relations may still be in flight. -- **`timestamp` format differs by endpoint.** On this endpoint each relation's `timestamp` is an ISO-8601 string (e.g. `2026-05-12T08:14:00Z`). The same relations surfaced as passthrough on [Query](/api-reference/v2/endpoint/query) (in `graph_context`) and [List Documents](/api-reference/v2/endpoint/list-documents) carry `timestamp` as a Unix epoch float (seconds) instead. Normalize before comparing relation timestamps across endpoints. +- **`timestamp` format differs by endpoint:** On this endpoint each relation's `timestamp` is an ISO-8601 string (e.g. `2026-05-12T08:14:00Z`). The same relations surfaced as passthrough on [Query](/api-reference/v2/endpoint/query) (in `graph_context`) and [List Documents](/api-reference/v2/endpoint/list-documents) carry `timestamp` as a Unix epoch float (seconds) instead. Normalize before comparing relation timestamps across endpoints. + +## Errors + +Common codes: `400 INVALID_INPUT` (missing `database`, a `limit` outside `1` to `10000`, or a non-numeric `cursor`), `404 DATABASE_NOT_FOUND`, `500 INTERNAL_ERROR` (a transient graph read failure; retry the request). See [Error Responses](/api-reference/v2/error-responses) for the full list.
**Related Resources** - - **Indexing status:** [Ingestion Status](/api-reference/v2/endpoint/source-status) - confirm the graph is complete + - **Indexing status:** [Ingestion Status](/api-reference/v2/endpoint/source-status), to confirm the graph is complete - **Query with graph context:** [Query](/api-reference/v2/endpoint/query) with `graph_context: true` - - **Concepts:** [Concepts → Context Graphs](/essentials/v2/context-graphs) + - **Concepts:** [Concepts: Context Graphs](/essentials/v2/context-graphs) diff --git a/api-reference/v2/endpoint/source-status.mdx b/api-reference/v2/endpoint/source-status.mdx index 3d02557f..b217204e 100644 --- a/api-reference/v2/endpoint/source-status.mdx +++ b/api-reference/v2/endpoint/source-status.mdx @@ -8,7 +8,7 @@ import { Field } from "/snippets/field.jsx"; Since ingestion is asynchronous, use this endpoint to determine when context is ready to be retrieved. -Pass one or more IDs in `ids` to retrieve status. Works for documents, app sources, and memories. When passing multiple IDs on the query string, use either repeated params (`?ids=policy_main&ids=runbook_deploy`) or a single comma-joined value (`?ids=policy_main,runbook_deploy`); both forms are equivalent and can be mixed. Surrounding whitespace is trimmed and empty entries are dropped. For more information, see the [Knowledge](/essentials/v2/knowledge) and [Memories](/essentials/v2/memories) guides. +Pass one or more IDs in `ids` to retrieve status. Works for documents, app sources, and memories. When passing multiple IDs on the query string, use either repeated params (`?ids=policy_main&ids=runbook_deploy`) or a single comma-joined value (`?ids=policy_main,runbook_deploy`); both forms are equivalent and can be mixed. Surrounding whitespace is trimmed, empty entries are dropped, and duplicates are removed. For more information, see the [Knowledge](/essentials/v2/knowledge) and [Memories](/essentials/v2/memories) guides. **Prefer webhooks over polling?** Register a webhook for `indexing.status_changed` events and HydraDB will `POST` to your endpoint when content reaches a terminal state (`completed` or `errored`). See [Webhooks](/essentials/v2/webhooks) for setup and receiver examples. @@ -47,7 +47,8 @@ curl -G 'https://api.hydradb.com/context/status' \ | Name | Description | | --- | --- | -| | One or more `id` values returned at ingestion. Accepts IDs for documents, app sources, or memories. Pass either repeated params (`ids=a&ids=b`) or a single comma-joined value (`ids=a,b`). Source IDs never contain commas (they are rejected at ingest), so the comma-joined form always splits unambiguously. | +| | One or more `id` values returned at ingestion. Accepts IDs for documents, app sources, or memories. Pass either repeated params (`ids=a&ids=b`) or a single comma-joined value (`ids=a,b`). Source IDs never contain commas (they are rejected at ingest), so the comma-joined form always splits unambiguously. | +| | A single ID. Merged with `ids` when you send both. At least one ID across `id` and `ids` is required. | | | Database the items belong to. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | | | Collection scope. If omitted, the default collection is used. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`null`) | @@ -76,9 +77,9 @@ curl -G 'https://api.hydradb.com/context/status' \ "id": "typo_in_id", "indexing_status": "errored", "error_code": "FILE_NOT_FOUND", - "error_message": "ID not found", + "error_message": "", "success": false, - "message": "Processing status retrieved successfully" + "message": "ID not found" } ] }, @@ -116,22 +117,23 @@ Each entry in `data.statuses` describes one requested `id`: | `id` | string | The source, app-source, or memory ID you asked about (echoed back). | | `indexing_status` | string | One of the [status values](#status-values) below. `errored` is terminal. | | `error_code` | string | Machine-readable reason an entry is `errored`; **empty string (`""`) when the entry is not errored.** See [`error_code` values](#error-code-values). | -| `error_message` | string | Human-readable explanation that accompanies a non-empty `error_code`; empty otherwise. | -| `success` | boolean | `false` when `indexing_status` is `errored`, otherwise `true`. Describes the item, **not** the HTTP request - a `200` response can contain `errored` items. | -| `message` | string | Status of the *lookup* itself ("Processing status retrieved successfully"). It does **not** describe the ingestion outcome - read `indexing_status` / `error_code` for that. | +| `error_message` | string | Human-readable explanation that accompanies a pipeline `error_code`; empty otherwise, including for `FILE_NOT_FOUND`. | +| `success` | boolean | Deprecated. `false` when `indexing_status` is `errored`, otherwise `true`. Describes the item, **not** the HTTP request: a `200` response can contain `errored` items. Read `indexing_status` instead. | +| `message` | string | Status of the *lookup* itself ("Processing status retrieved successfully", or "ID not found" for an unknown `id`). It does **not** describe the ingestion outcome. Read `indexing_status` / `error_code` for that. | ### Error code values -`error_code` is the field that lets you tell a **caller mistake** apart from a **real ingestion failure** - a distinction you cannot make from `indexing_status: "errored"` alone. It is empty on any non-errored entry. +`error_code` is the field that lets you tell a **caller mistake** apart from a **real ingestion failure**, a distinction you cannot make from `indexing_status: "errored"` alone. It is empty on any non-errored entry. | `error_code` | Meaning | What to do | | --- | --- | --- | -| `FILE_NOT_FOUND` | No source with this `id` exists in the given `database`/`collection` - usually a typo or an `id` that was never ingested (or whose status has expired). | Fix the `id`, or (re-)ingest the source. Not a processing failure - retrying the status call will not change it. | -| `INVALID_FILE_ID` | The `id` was empty or blank. | Send a non-empty `id`. | -| *ingestion-pipeline codes* | A genuine processing failure (e.g. `PARSE_FAILED`, `UNSUPPORTED_FORMAT`, `PROCESSING_FAILED`, `EMBEDDING_FAILED`, …). | Act on the specific code - see the [Error Responses reference](/api-reference/v2/error-responses#common-error-codes). Many are re-ingest-and-retry; some are terminal (unsupported format, empty content). | +| `FILE_NOT_FOUND` | No source with this `id` exists in the given `database`/`collection`: usually a typo, an `id` that was never ingested, or a deleted source. | Fix the `id`, or (re-)ingest the source. Not a processing failure, so retrying the status call will not change it. | +| *ingestion-pipeline codes* | A genuine processing failure, as a numeric `E####` code (e.g. `E1001` parse failed, `E2002` no text extracted, `E4001` embedding failed). | Act on the specific code. See the [Error Responses reference](/api-reference/v2/error-responses#ingestion-error-codes). Many are re-ingest-and-retry; some are terminal (unsupported format, empty content). | + +Status records can expire. A source that completed still reports `completed` after its record expires. A source that never finished before its record expired reports `errored` with `E9004`; ingest it again. - Branch on `error_code`, not on the text in `message` or `error_message`. `message` describes the lookup, not the ingestion result, and human-readable text may change. The full list of codes an `errored` entry can carry is in the [Error Responses reference](/api-reference/v2/error-responses#common-error-codes). + Branch on `error_code`, not on the text in `message` or `error_message`. `message` describes the lookup, not the ingestion result, and human-readable text may change. The full list of codes an `errored` entry can carry is in the [Error Responses reference](/api-reference/v2/error-responses#ingestion-error-codes). ## Status values @@ -171,7 +173,7 @@ flowchart LR | `completed` | Yes | Fully indexed and graphed. Ready for all retrieval modes. | | `errored` | No | Processing failed. Inspect `error_code` and `error_message`. | -The normal progression is `queued` → `processing` → `graph_creation` → `completed`. Treat `errored` as terminal. +The normal progression is `queued`, then `processing`, then `graph_creation`, then `completed`. Treat `errored` as terminal. ## Polling patterns @@ -246,8 +248,8 @@ while True: Typical processing time: - **Memories** (text, markdown, conversation pairs): seconds -- **Small documents** (under 50 pages): 1–5 minutes -- **Large documents** (50\+ pages): 5–15 minutes +- **Small documents** (under 50 pages): 1 to 5 minutes +- **Large documents** (50\+ pages): 5 to 15 minutes ## Behavior notes @@ -255,21 +257,21 @@ Typical processing time: **`graph_creation` is searchable.** Items in this state are already retrievable via `/query`. Wait for `completed` only when you specifically need full graph traversal (`graph_context: true`). -- **Unknown IDs return as `errored`:** If you pass an ID that does not exist (e.g., a typo), HydraDB returns an entry with `indexing_status: "errored"` and `error_code: "FILE_NOT_FOUND"` rather than silently dropping it. Use `error_code` to distinguish this from a genuine ingestion failure - see [`error_code` values](#error-code-values). +- **Unknown IDs return as `errored`:** If you pass an ID that does not exist (e.g., a typo), HydraDB returns an entry with `indexing_status: "errored"` and `error_code: "FILE_NOT_FOUND"` rather than silently dropping it. Use `error_code` to distinguish this from a genuine ingestion failure. See [`error_code` values](#error-code-values). ## Errors -Common codes: `400 INVALID_PARAMETERS`, `404 DATABASE_NOT_FOUND`, `422 VALIDATION_ERROR`. See [Error Responses](/api-reference/v2/error-responses) for the full list. +Common codes: `400 INVALID_INPUT` (for example `database` and `tenant_id` both sent with different values), `404 DATABASE_NOT_FOUND`, `422 VALIDATION_ERROR` (missing `database`, or no ID). See [Error Responses](/api-reference/v2/error-responses) for the full list.
**Related Resources** - - **Before this:** [Ingest Context](/api-reference/v2/endpoint/ingest-context) - to get the IDs + - **Before this:** [Ingest Context](/api-reference/v2/endpoint/ingest-context), to get the IDs - **After completion:** [Query](/api-reference/v2/endpoint/query) - **After completion:** [Fetch Content](/api-reference/v2/endpoint/fetch-content) - **After completion:** [Context Relations](/api-reference/v2/endpoint/source-relations) - - **Read more:** [Usage → Knowledge](/essentials/v2/knowledge) - - **Read more:** [Usage → Memories](/essentials/v2/memories) + - **Read more:** [Usage: Knowledge](/essentials/v2/knowledge) + - **Read more:** [Usage: Memories](/essentials/v2/memories) diff --git a/api-reference/v2/endpoint/sources-overview.mdx b/api-reference/v2/endpoint/sources-overview.mdx index aba256b2..c64401bf 100644 --- a/api-reference/v2/endpoint/sources-overview.mdx +++ b/api-reference/v2/endpoint/sources-overview.mdx @@ -1,5 +1,5 @@ --- -title: "Context Management - Overview" +title: "Context Management: Overview" description: "Quick reference for context management endpoints, their lifecycle, and when to use which." --- @@ -13,8 +13,10 @@ description: "Quick reference for context management endpoints, their lifecycle, | Poll indexing progress | `/context/status` | | Browse stored sources or memories | `/context/list` | | Inspect original source content | `/context/inspect` | +| Update a source's metadata without re-ingesting | `PATCH /context/{id}/metadata` | | Delete sources or memories | `DELETE /context` | | Inspect graph relations | `/context/relations` | +| Get everything connected to one item | `/context/{id}/subgraph` | ## Lifecycle @@ -54,22 +56,22 @@ flowchart LR ``` - **Why both** `type=knowledge `**and** `app_knowledge`**?** They have a theoretical differentiation. + **Why both** `type=knowledge` **and** `app_knowledge`**?** They have a theoretical differentiation. - `type` picks the **bucket**: `knowledge` (shared documents) or `memory` (per-user context). It routes the ingest to the right store. - - Within `type=knowledge`, you pick the **payload shape**: `documents` (binary documents HydraDB will parse - PDFs, DOCX, CSV) or `app_knowledge` (a JSON array of already-extracted content from your app - Slack messages, Notion pages, web pages). You can send both in the same request. + - Within `type=knowledge`, you pick the **payload shape**: `documents` (binary documents HydraDB will parse: PDFs, DOCX, CSV) or `app_knowledge` (a JSON array of already-extracted content from your app: Slack messages, Notion pages, web pages). You can send both in the same request. ## Core Ingestion Concepts - **Knowledge vs. Memories**: [Knowledge](/essentials/v2/knowledge) is shared, database-wide content (documents, app pages, Slack messages). [Memories](/essentials/v2/memories) are user-specific preferences and conversational traits scoped by `collection`. Both can be searched together via `type: "all"` on `POST /query`. -- **IDs**: Unique identifiers returned by `/context/ingest`. You can assign custom IDs using `id` in metadata or `id` in `app_knowledge` items. Use them for polling status, inspecting content, and deleting context. +- **IDs**: Unique identifiers returned by `/context/ingest`. You can assign custom IDs using `id` in `document_metadata`, `app_knowledge`, or `memories` items. Use them for polling status, inspecting content, and deleting context. - **Metadata Filtering**: You can scope queries using `metadata` (structured fields defined in your database schema) or `additional_metadata` (free-form per-document JSON). For detailed guidelines on structuring metadata, see the [Scoping using metadata](/essentials/v2/metadata) guide. - **Forceful Relations**: Relationships between sources can be declared at ingestion time to construct a robust knowledge graph. For more details on the graph layer, see the [Context Graphs](/essentials/v2/context-graphs) guide. ## Forceful relations and metadata -Forceful relations let you pre-wire document relationships at ingestion time so that relevant documents surface together during retrieval - even before the graph layer discovers connections organically. Think of them as explicit "see also" links between your documents. +Forceful relations let you pre-wire document relationships at ingestion time so that relevant documents surface together during retrieval, even before the graph layer discovers connections organically. Think of them as explicit "see also" links between your documents. Paired with document-level metadata, you get deterministic control over how results are filtered and ranked. @@ -89,13 +91,13 @@ result = client.context.ingest( ## Related sections -- [Usage - Forceful Relations](/essentials/v2/knowledge) - linking sources at ingestion (see §7) -- [Query](/api-reference/v2/endpoint/query-overview) - retrieve ingested content +- [Usage: Forceful Relations](/essentials/v2/knowledge): linking sources at ingestion (see section 7) +- [Query](/api-reference/v2/endpoint/query-overview): retrieve ingested content Related Resources - - [Usage - Memories](/essentials/v2/memories) - memories vs knowledge, when to use which + - [Usage: Memories](/essentials/v2/memories): memories vs knowledge, when to use which - - [Usage - Metadata](/essentials/v2/metadata) - database-level vs document-level metadata + - [Usage: Metadata](/essentials/v2/metadata): database-level vs document-level metadata diff --git a/api-reference/v2/endpoint/subgraph.mdx b/api-reference/v2/endpoint/subgraph.mdx index e8c9a8d9..84052cb5 100644 --- a/api-reference/v2/endpoint/subgraph.mdx +++ b/api-reference/v2/endpoint/subgraph.mdx @@ -8,10 +8,28 @@ import { Field } from "/snippets/field.jsx"; This endpoint returns the **connected subgraph** of one ingested item: every item reachable from it through item-level relations, traversed breadth-first up to `depth` hops, together with the relations among those members and the structural graph around them (entities, comments, attachments, people). -It answers a different question from [Inspecting Context Relations](/api-reference/v2/endpoint/source-relations). Relations are the entity-and-predicate triplets *extracted from text* (`PaymentsWorker → depends_on → OrdersDB`). The subgraph is about *items*: which Slack message replies to which, which page links to which, which ticket a comment belongs to. Use it after [Query](/api-reference/v2/endpoint/query) or [List Documents](/api-reference/v2/endpoint/list-documents) when a single result is not enough and you need what surrounds it. +It answers a different question from [Inspecting Context Relations](/api-reference/v2/endpoint/source-relations). Relations are the entity-and-predicate triplets *extracted from text* (`PaymentsWorker` `depends_on` `OrdersDB`). The subgraph is about *items*: which Slack message replies to which, which page links to which, which ticket a comment belongs to. Use it after [Query](/api-reference/v2/endpoint/query) or [List Documents](/api-reference/v2/endpoint/list-documents) when a single result is not enough and you need what surrounds it. +```python Python SDK +subgraph = client.context.subgraph( + id="slack_C0BE77_1788320073", + database="acme_corp", + collection="eng_slack", + depth=3, +) +``` + +```typescript TypeScript SDK +const subgraph = await client.context.subgraph({ + id: "slack_C0BE77_1788320073", + database: "acme_corp", + collection: "eng_slack", + depth: 3, +}); +``` + ```bash cURL curl -G 'https://api.hydradb.com/context/slack_C0BE77_1788320073/subgraph' \ -H "Authorization: Bearer " \ @@ -32,15 +50,11 @@ hydradb --output json subgraph slack_C0BE77_1788320073 | jq '.sources[].source_i - - The Python and TypeScript SDKs gain `context.subgraph()` with their next release, generated from this spec. Until then call the endpoint directly as above; the CLI and the MCP server already do. - - ## Path parameters | Name | Description | | --- | --- | -| | The item to start from. Any `id` returned by Query, List Documents or Ingest. URL-encode it if it contains reserved characters. An id containing a literal `/` cannot be written as one path segment; pass those as `GET /context/subgraph?id=...` instead. | +| | The item to start from. Any `id` returned by Query, List Documents or Ingest. URL-encode it if it contains reserved characters. An id containing a literal `/` cannot be written as one path segment; pass those as `GET /context/subgraph?id=...` instead (the form the SDKs use for every id). | ## Query parameters @@ -49,8 +63,8 @@ hydradb --output json subgraph slack_C0BE77_1788320073 | jq '.sources[].source_i | | Owning database. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | | | Collection scope. If omitted, the default collection is used. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=`null`) | | | Which graph the `id` belongs to. The two graphs are completely separate. (default=`"knowledge"`) | -| | Maximum traversal depth in hops. Range `1–10`. (default=`5`) | -| | Maximum number of members returned. Range `1–1000`. When this clips the traversal, `is_truncated` is `true`. (default=`200`) | +| | Maximum traversal depth in hops. Range `1` to `10`. (default=`5`) | +| | Maximum number of members returned. Range `1` to `1000`. When this clips the traversal, `is_truncated` is `true`. (default=`200`) | | | Principals to answer as (document ACLs). The subgraph then contains only items those principals may see, filtered at every hop. Repeated (`acl=a&acl=b`) or comma-separated. Omit for no ACL scoping. | ## How items connect @@ -117,7 +131,7 @@ Traversal is breadth-first, so `depth` on each member is its distance from the s "is_truncated": false, "max_depth_reached": 1, "success": true, - "message": "Subgraph fetched successfully" + "message": "Successfully fetched source subgraph" }, "error": null, "meta": { @@ -139,7 +153,7 @@ Traversal is breadth-first, so `depth` on each member is its distance from the s "is_truncated": false, "max_depth_reached": 0, "success": true, - "message": "Subgraph fetched successfully" + "message": "Successfully fetched source subgraph" }, "error": null, "meta": { @@ -180,18 +194,22 @@ Traversal is breadth-first, so `depth` on each member is its distance from the s - **An item nothing links to** comes back as a one-member subgraph: itself, at depth `0`, with `max_depth_reached: 0`. That is a real answer ("this stands alone"), distinct from an unknown id, which has no members. -- **Bounding the traversal.** Threads and hierarchies can be large. `depth` bounds how far the walk goes; `max_sources` bounds how many members it returns. When `max_sources` clips it, `is_truncated` is `true` and the members you have are the ones closest to the start item. `auxiliary_truncated` reports the same for the structural graph. -- **Knowledge vs memory.** If the `id` belongs to a memory, set `type=memory`; the two graphs never connect to each other. -- **Completeness.** An item's links populate once its `indexing_status` reaches `completed`. Items still in `graph_creation` may appear with fewer connections than they will have. -- **Cost.** One request fans out into a bounded series of graph reads, so it is rate-limited like a Query, not like a status poll. +- **Bounding the traversal:** Threads and hierarchies can be large. `depth` bounds how far the walk goes; `max_sources` bounds how many members it returns. When `max_sources` clips it, `is_truncated` is `true` and the members you have are the ones closest to the start item. `auxiliary_truncated` reports the same for the structural graph. +- **Knowledge vs memory:** If the `id` belongs to a memory, set `type=memory`; the two graphs never connect to each other. +- **Completeness:** An item's links populate once its `indexing_status` reaches `completed`. Items still in `graph_creation` may appear with fewer connections than they will have. +- **Cost:** One request fans out into a bounded series of graph reads, so it is rate-limited like a Query, not like a status poll. + +## Errors + +Common codes: `400 INVALID_INPUT` (missing `database` or `id`, `depth` outside `1` to `10`, or `max_sources` outside `1` to `1000`), `404 DATABASE_NOT_FOUND`, `500 INTERNAL_ERROR` (a transient graph read failure; retry the request). See [Error Responses](/api-reference/v2/error-responses) for the full list.
**Related Resources** - - **Entity relations:** [Inspecting Context Relations](/api-reference/v2/endpoint/source-relations) - the triplets extracted from text + - **Entity relations:** [Inspecting Context Relations](/api-reference/v2/endpoint/source-relations), the triplets extracted from text - **Full content of a member:** [Fetch Content](/api-reference/v2/endpoint/fetch-content) - **Query with graph context:** [Query](/api-reference/v2/endpoint/query) with `graph_context: true` - - **Concepts:** [Concepts → Context Graphs](/essentials/v2/context-graphs) + - **Concepts:** [Concepts: Context Graphs](/essentials/v2/context-graphs) diff --git a/api-reference/v2/endpoint/submit-feedback.mdx b/api-reference/v2/endpoint/submit-feedback.mdx index 11c352ca..a25cbb84 100644 --- a/api-reference/v2/endpoint/submit-feedback.mdx +++ b/api-reference/v2/endpoint/submit-feedback.mdx @@ -6,13 +6,13 @@ openapi: "api-reference/v2/openapi.json POST /feedback" import { Field } from "/snippets/field.jsx"; -Report back on a query that already ran - what was missing, what was wrong, or that it was exactly right. Feedback feeds retrieval-quality work; it does **not** change the result of the query it refers to. +Report back on a query that already ran: what was missing, what was wrong, or that it was exactly right. Feedback feeds retrieval-quality work; it does **not** change the result of the query it refers to. Both people and agents can submit. An agent that can tell a retrieval was unhelpful is often the best source of signal you have, so `source` labels which one it was. ## Linking feedback to a query -Every HydraDB response carries a `request_id` in `meta`, and the same value in the `X-Request-ID` header. Send that id back and we can line your comment up with the exact query it is about - the text queried, what came back, how long it took. +Every HydraDB response carries a `request_id` in `meta`, and the same value in the `X-Request-ID` header. Send that id back and we can line your comment up with the exact query it is about: the text queried, what came back, how long it took. ```json Query response {6} { @@ -28,10 +28,10 @@ Every HydraDB response carries a `request_id` in `meta`, and the same value in t ``` -Send `request_id` back **exactly as you received it**. It must be the UUID from `meta.request_id` (or the `X-Request-ID` header) - any other value is rejected with `400`. +Send `request_id` back **exactly as you received it**. It must be the UUID from `meta.request_id` (or the `X-Request-ID` header). Any other value is rejected with `400`. -Submit feedback for queries that **returned**. If the query itself failed, handle the error instead - there is no retrieval to judge, and the fix is in the request rather than in the index. +Submit feedback for queries that **returned**. If the query itself failed, handle the error instead: there is no retrieval to judge, and the fix is in the request rather than in the index. ## Fields @@ -42,7 +42,7 @@ The `request_id` from the target query's `meta`. Must be a UUID. What was right or wrong, in your own words. Up to 8000 characters. -Required **unless** you send `ground_truth` - every submission needs at least one of the two. +Required **unless** you send `ground_truth`; every submission needs at least one of the two. @@ -53,17 +53,17 @@ What you already know the right answer to be. See [Ground truth](#ground-truth). The response you expected. Up to 8000 characters. - IDs of the sources that actually contain the answer. Up to 100. + IDs of the sources that actually contain the answer. Up to 100, each up to 256 characters. -`positive`, `negative`, or `neutral`. Optional - leaving it out is not the same as `neutral`; it records that you sent a comment without a rating. +`positive`, `negative`, or `neutral`. Optional. Leaving it out is not the same as `neutral`; it records that you sent a comment without a rating. -`user` _(default)_ or `agent` - who is submitting. +`user` _(default)_ or `agent`: who is submitting. @@ -71,7 +71,7 @@ Optional. Scopes the feedback to a database. Must be one your API key can reach. -Optional. Requires `database` - a collection is scoped to a database, so sending it alone returns `400`. +Optional. Requires `database`: a collection is scoped to a database, so sending it alone returns `400`. @@ -102,8 +102,12 @@ const result = await client.query({ query: "What is our refund policy?", }); +// Response fields are optional in the TypeScript types. +const requestId = result.meta?.requestId; +if (!requestId) throw new Error("query response has no request_id"); + await client.feedback.submit({ - requestId: result.meta.requestId, + requestId, feedback: "Returned the 2023 policy - the current one is in the Q3 handbook.", rating: "negative", source: "agent", @@ -184,7 +188,7 @@ curl -X POST 'https://api.hydradb.com/feedback' \ ## Ground truth -If you already know the right answer - you are running an evaluation set, or you know which document the user needed - send it. It is a much stronger signal than a comment, because we can score it without a human reading it. +If you already know the right answer (you are running an evaluation set, or you know which document the user needed), send it. It is a much stronger signal than a comment, because we can score it without a human reading it. ```json "ground_truth": { @@ -193,13 +197,13 @@ If you already know the right answer - you are running an evaluation set, or y } ``` -- **`answer`** - the response you expected. -- **`source_ids`** - the sources that actually contain the answer. This is the one that grades retrieval: it tells us whether the query surfaced those documents, and where they ranked. +- **`answer`:** the response you expected. +- **`source_ids`:** the sources that actually contain the answer. This is the one that grades retrieval: it tells us whether the query surfaced those documents, and where they ranked. -Send either on its own or both together. If `ground_truth` is your only signal, at least one of the two has to carry something - values that are empty or all whitespace are treated as not sent. +Send either on its own or both together. If `ground_truth` is your only signal, at least one of the two has to carry something; values that are empty or all whitespace are treated as not sent. -When you send `ground_truth`, the `feedback` comment becomes optional - an evaluation run with an answer key does not need prose for every row. A submission with neither is rejected. +When you send `ground_truth`, the `feedback` comment becomes optional: an evaluation run with an answer key does not need prose for every row. A submission with neither is rejected. ```python Evaluation run @@ -221,26 +225,28 @@ for case in eval_set: At eval volumes you may brush the rate limit, so keep the submission from ending the loop: an unguarded call means a single `429` loses every remaining case, not just the one it failed on. -Duplicate `source_ids` are collapsed and blank entries dropped, so you do not need to de-duplicate or filter your answer key first - a list that still has one real id in it is scored on that id. +Duplicate `source_ids` are collapsed and blank entries dropped, so you do not need to de-duplicate or filter your answer key first. A list that still has one real id in it is scored on that id. ## Submitting more than once -Each submission is stored separately - a second comment about the same query does not replace the first. Send several as your understanding of a bad result develops, and file feedback from more than one user on the same query. +Each submission is stored separately: a second comment about the same query does not replace the first. Send several as your understanding of a bad result develops, and file feedback from more than one user on the same query. ## Rate limit -100 submissions per minute per organization. Over that, you get `429` with a `Retry-After` header and a message naming the seconds to wait - it is safe to retry after waiting. +100 submissions per minute per organization. Over that, you get `429` with a `Retry-After` header and a message naming the seconds to wait; it is safe to retry after waiting. The ceiling is well above normal use; an agent reporting on every query it makes will stay comfortably under it. ## Errors +A successful submission returns `201`. + | Status | When | | --- | --- | -| `400` | `request_id` missing or not a UUID; no usable signal - `feedback` blank or absent **and** `ground_truth` absent, empty, or blank; `feedback` too long; unknown `rating`/`source`; `collection` without `database` | +| `400` | `request_id` missing or not a UUID; no usable signal (`feedback` blank or absent **and** `ground_truth` absent, empty, or blank); `feedback` too long; unknown `rating`/`source`; `collection` without `database`; or any other limit in [Fields](#fields) exceeded | | `401` | Missing or invalid API key | | `404` | `database` does not exist or is not reachable by this key | -| `429` | Over the rate limit - see `Retry-After` | -| `500` | Feedback could not be stored. Nothing was recorded; retrying is safe | +| `429` | Over the rate limit; see `Retry-After` | +| `500` | Feedback could not be recorded. Retrying is safe | -A `500` means the submission was **not** saved, so a retry cannot create a duplicate of something already stored. +Each submission is its own row, so the worst case of retrying after a `500` is a duplicate report. diff --git a/api-reference/v2/endpoint/tenant-stats.mdx b/api-reference/v2/endpoint/tenant-stats.mdx index 153e604a..51a4b499 100644 --- a/api-reference/v2/endpoint/tenant-stats.mdx +++ b/api-reference/v2/endpoint/tenant-stats.mdx @@ -5,14 +5,14 @@ description: "Retrieve usage statistics for a database." openapi: "api-reference/v2/openapi.json GET /databases/stats" --- -Get detailed metrics including storage growth and object counts. +Get detailed metrics including object counts. The response splits stats into two database-wide collections: `data.knowledge_collection` (Knowledge) and `data.memory_collection` (User memories), each with its own row count. Counts aggregate across all collections in the database. ```python Python SDK -stats = client.databases.stats(database="your database id") +stats = client.databases.stats(database="my_first_database") ``` ```typescript TypeScript SDK @@ -103,5 +103,5 @@ curl -X GET 'https://api.hydradb.com/databases/stats?database=my_first_database' - **List sources/memories:** [List Documents](/api-reference/v2/endpoint/list-documents) - **List collections:** [List Collections](/api-reference/v2/endpoint/list-sub-tenants) - **Check provisioning:** [Database Status](/api-reference/v2/endpoint/tenant-status) - - **Read more:** [Concepts → Multi-Tenant Support](/essentials/v2/multi-tenant) + - **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) diff --git a/api-reference/v2/endpoint/tenant-status.mdx b/api-reference/v2/endpoint/tenant-status.mdx index ce1eb16c..e2556529 100644 --- a/api-reference/v2/endpoint/tenant-status.mdx +++ b/api-reference/v2/endpoint/tenant-status.mdx @@ -4,12 +4,12 @@ description: "Check the readiness of a database's infrastructure." openapi: "api-reference/v2/openapi.json GET /databases/status" --- -Database creation is asynchronous, check if your database is ready before executing ingestion or any queries. +Database creation is asynchronous; check if your database is ready before executing ingestion or any queries. -The response includes `infra.ready_for_ingestion`, a convenience flag derived from the vectorstore fields below. For a single check, read that field. The underlying signals are: +The response includes `infra.ready_for_ingestion`, a convenience flag derived from the four signals below. For a single check, read that field. The underlying signals are: -- `infra.ready_for_ingestion === true` (derived - true once both vectorstores are provisioned) -- `infra.scheduler_status === true` (the background indexing scheduler) +- `infra.ready_for_ingestion === true` (derived: true only when all four signals below are true) +- `infra.scheduler_status === true` (lifecycle provisioning has finished; `false` while the database is still being created, even if some collections already exist) - `infra.graph_status === true` (the graph layer) - `infra.vectorstore_status.knowledge === true` - `infra.vectorstore_status.memories === true` @@ -28,7 +28,7 @@ const response = await client.databases.status({ database: "my_first_database", }); -if (response.data?.infra.readyForIngestion) { +if (response.data?.infra?.readyForIngestion) { console.log("Database is ready."); } ``` @@ -74,7 +74,7 @@ curl -X GET 'https://api.hydradb.com/databases/status?database=my_first_database "data": null, "error": { "code": "DATABASE_NOT_FOUND", - "message": "Database not found" + "message": "database my_first_database does not exist" }, "meta": { "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", @@ -87,6 +87,8 @@ curl -X GET 'https://api.hydradb.com/databases/status?database=my_first_database **Stale database IDs:** If `database` does not exist, the call returns `404 DATABASE_NOT_FOUND`. Always verify that the database was created successfully before polling. +**Querying too early:** a `POST /query` sent before `ready_for_ingestion` is `true` returns `422 TENANT_INFRA_NOT_READY`. + **Common mistake:** `row_count` from [Database Stats](/api-reference/v2/endpoint/tenant-stats) counts individual chunks, not documents. For distinct source or memory counts, use [List Documents](/api-reference/v2/endpoint/list-documents) and read `pagination.total` from the response. @@ -96,6 +98,6 @@ curl -X GET 'https://api.hydradb.com/databases/status?database=my_first_database **Related Resources** - - **Before this:** [Create Database](/api-reference/v2/endpoint/create-tenant) - kicks off provisioning - - **After this:** [Ingest Context](/api-reference/v2/endpoint/ingest-context) - once status is ready + - **Before this:** [Create Database](/api-reference/v2/endpoint/create-tenant): kicks off provisioning + - **After this:** [Ingest Context](/api-reference/v2/endpoint/ingest-context): once status is ready diff --git a/api-reference/v2/endpoint/tenants-overview.mdx b/api-reference/v2/endpoint/tenants-overview.mdx index 81d9d111..b0e5a1b4 100644 --- a/api-reference/v2/endpoint/tenants-overview.mdx +++ b/api-reference/v2/endpoint/tenants-overview.mdx @@ -1,5 +1,5 @@ --- -title: "Databases - Overview" +title: "Databases: Overview" description: "Quick reference for all databases endpoints, their lifecycle, and when to call each." --- @@ -16,6 +16,7 @@ Databases are physically isolated spaces for storing context. In most integratio | [`/databases/stats`](/api-reference/v2/endpoint/tenant-stats) | `GET` | `databases.stats` | Monitor database load | No | | [`/databases/collections`](/api-reference/v2/endpoint/list-sub-tenants) | `GET` | TypeScript: `databases.collections`
Python: `databases.collections` | List active collections | No | | [`/databases/collections`](/api-reference/v2/endpoint/delete-collection) | `DELETE` | TypeScript: `databases.deleteCollection`
Python: `databases.delete_collection` | Permanently remove one collection | Yes | +| [`/databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema) | `PATCH` | TypeScript: `databases.updateMetadataSchema`
Python: `databases.update_metadata_schema` | Add metadata schema fields | No | ## Typical call sequence @@ -39,6 +40,6 @@ GET /databases/stats -> check database health & growth ## Key concepts -- **Database** - A top-level isolated space. For example - you can dedicate one database to one enterprise customer. -- **Collection** - Partitions within a database for per-user separation. The first collection is created implicitly at ingestion. Collections are useful when you need to scope data per user, team, or customer within a single database. -- **Database Metadata & Schema** - Structured fields defined at database creation to enable query-time filtering. \ No newline at end of file +- **Database:** A top-level isolated space. For example, you can dedicate one database to one enterprise customer. +- **Collection:** Partitions within a database for per-user separation. The first collection is created implicitly at ingestion. Collections are useful when you need to scope data per user, team, or customer within a single database. +- **Database Metadata & Schema:** Structured fields defined at database creation to enable query-time filtering. You can add fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema). \ No newline at end of file diff --git a/api-reference/v2/endpoint/test-webhook.mdx b/api-reference/v2/endpoint/test-webhook.mdx index 98c41a7e..da47f8f3 100644 --- a/api-reference/v2/endpoint/test-webhook.mdx +++ b/api-reference/v2/endpoint/test-webhook.mdx @@ -6,7 +6,7 @@ openapi: "api-reference/v2/openapi.json POST /webhooks/indexing/test" Sends a synthetic `indexing.status_changed` payload to your registered endpoint so you can confirm it is reachable and that your signature verification works. -The test payload carries `"test": true`, and is signed exactly like a real delivery when a signing secret is configured. Use it to check your verifier before relying on it in production. +The test payload carries `"test": true`, and is signed exactly like a real delivery when a signing secret is configured. Use it to check your verifier before relying on it in production. The response reports whether your endpoint accepted it (`delivered`) and the HTTP `status_code` it returned. If no webhook is registered, the call returns `404`. Test deliveries do not appear in the delivery history returned by [List Deliveries](/api-reference/v2/endpoint/list-webhook-deliveries). diff --git a/api-reference/v2/endpoint/update-connector.mdx b/api-reference/v2/endpoint/update-connector.mdx index 6b60fa9c..f1055051 100644 --- a/api-reference/v2/endpoint/update-connector.mdx +++ b/api-reference/v2/endpoint/update-connector.mdx @@ -10,6 +10,8 @@ Updates a connector in place. Every field is optional and only the fields you se - **`sync_interval_seconds`** sets how often scheduled syncs run. The allowed range depends on the provider and comes back in the response as `min_sync_interval_seconds` and `max_sync_interval_seconds`. Values outside it are rejected, and `0` resets to the provider default. The next sync is rescheduled right away, so a shorter interval takes effect immediately. - **`credentials`** reconnects the connector, for example after a token expired or was revoked. Send the provider's full credential set, the same as on create. The connector keeps its id, resources and sync positions, and a reconnect-required state is cleared. +Only these three fields are read. A connector's `name`, `database`, `collection` and `provider_account_scope` cannot be changed after create; sending them has no effect. + ```bash cURL diff --git a/api-reference/v2/endpoint/update-metadata-schema.mdx b/api-reference/v2/endpoint/update-metadata-schema.mdx index f03c55ad..652a8c98 100644 --- a/api-reference/v2/endpoint/update-metadata-schema.mdx +++ b/api-reference/v2/endpoint/update-metadata-schema.mdx @@ -6,10 +6,10 @@ openapi: "api-reference/v2/openapi.json PATCH /databases/{database}/metadata-sch import { Field } from "/snippets/field.jsx"; -Use this endpoint to add new fields to a database's `database_metadata_schema`. +Use this endpoint to add new fields to a database's `database_metadata_schema`. SDK methods: `client.databases.update_metadata_schema()` (Python) and `client.databases.updateMetadataSchema()` (TypeScript). `GET /databases/{database}/metadata-schema` returns the current schema in the same field shape. - This endpoint is additive only. It cannot delete fields, rename fields, or change the type/flags of existing fields. + This endpoint is additive only. It cannot delete fields, rename fields, change the type/flags of existing fields, or add dense/sparse search lanes. @@ -23,14 +23,11 @@ curl -X PATCH 'https://api.hydradb.com/databases/acme_corp/metadata-schema' \ "add_fields": [ { "name": "region", - "data_type": "VARCHAR", - "enable_match": true + "data_type": "VARCHAR" }, { - "name": "summary_label", - "data_type": "VARCHAR", - "enable_dense_embedding": true, - "enable_sparse_embedding": true + "name": "priority", + "data_type": "INT64" } ] }' @@ -48,13 +45,8 @@ response = requests.patch( }, json={ "add_fields": [ - {"name": "region", "data_type": "VARCHAR", "enable_match": True}, - { - "name": "summary_label", - "data_type": "VARCHAR", - "enable_dense_embedding": True, - "enable_sparse_embedding": True, - }, + {"name": "region", "data_type": "VARCHAR"}, + {"name": "priority", "data_type": "INT64"}, ] }, ) @@ -70,13 +62,8 @@ const response = await fetch("https://api.hydradb.com/databases/acme_corp/metada }, body: JSON.stringify({ add_fields: [ - { name: "region", data_type: "VARCHAR", enable_match: true }, - { - name: "summary_label", - data_type: "VARCHAR", - enable_dense_embedding: true, - enable_sparse_embedding: true, - }, + { name: "region", data_type: "VARCHAR" }, + { name: "priority", data_type: "INT64" }, ], }), }); @@ -96,42 +83,42 @@ const response = await fetch("https://api.hydradb.com/databases/acme_corp/metada | Name | Description | | --- | --- | -| | New metadata schema fields to append. Must contain at least one field. | +| | New metadata schema fields to append. Must contain at least one field. A field that already exists with an identical definition is skipped, so you can resend your full field list. | Each `add_fields[]` item uses the same field shape as `database_metadata_schema` on [Create Database](/api-reference/v2/endpoint/create-tenant): | Field | Description | | --- | --- | -| | New metadata key. Must start with a letter or `_`, contain only letters/numbers/underscores, and not be a reserved system name. | +| | New metadata key. Must start with a letter, contain only letters/numbers/underscores (up to 255 characters), and not be a reserved system name. | | | `VARCHAR`, `BOOL`, `INT8`, `INT16`, `INT32`, `INT64`, `FLOAT`, `DOUBLE`, `JSON`, or friendly aliases such as `string`, `integer`, `float`, `boolean`, `object`. Defaults to `VARCHAR`. `ARRAY` is not supported and is rejected with `400`; for multi-value fields declare `VARCHAR` and store the values comma-joined. | | | Max length for `VARCHAR`. Default `1024`; maximum `65535`. | -| | Enables the intended exact-match metadata filtering path for this field. | -| | Adds a dense semantic-search lane for a `VARCHAR` metadata field. | -| | Adds a sparse/BM25 search lane for a `VARCHAR` metadata field. | -| | Backward-compatible shorthand for `enable_match: true`. Prefer `enable_match`. | +| | Rejected with `400` on a new field. Dense semantic-search lanes can only be declared at [database creation](/api-reference/v2/endpoint/create-tenant). | +| | Rejected with `400` on a new field. Sparse/BM25 search lanes can only be declared at database creation. | ## Rules - Additions only. -- Existing field names cannot be reused, case-insensitively. +- Existing field names cannot be reused with a different definition, case-insensitively. A field that differs from the existing one in type, `max_length`, or flags returns `409`, as does one name declared twice in a request with different definitions. Resending a field with its identical definition is a no-op. - Existing fields cannot be deleted or changed. -- Total custom database metadata fields cannot exceed 32. +- Total custom database metadata fields cannot exceed 32. Skipped fields do not count toward the new total. - Reserved names such as `source_id`, `chunk_id`, `metadata`, and `document_metadata` are rejected. -- Dense/sparse embedding flags are only valid on `VARCHAR` fields. -- MongoDB indexes for `enable_match` fields are created before the merged schema is persisted. +- Dense/sparse embedding flags are rejected on new fields. Resending an existing field that already has those flags is a no-op. - This endpoint persists the updated schema and MongoDB filter indexes. It does not yet alter/backfill existing Milvus collections for newly added dense/sparse metadata fields. Create the desired semantic metadata fields before ingesting, or migrate/re-ingest into a database with the final schema if those fields must participate in semantic/BM25 metadata search. + This endpoint persists the updated schema and MongoDB filter indexes. It does not alter existing Milvus collections, so it cannot add dense/sparse metadata fields. Declare the desired semantic metadata fields at database creation, or migrate/re-ingest into a database with the final schema if those fields must participate in semantic/BM25 metadata search. ## Response +A success returns `database`, the deprecated `tenant_id`, and `added_fields` directly, without the `success`/`data` envelope. `added_fields` lists only the fields this call added, so an empty list means the database already had everything you sent. Errors use the standard [error envelope](/api-reference/v2/error-responses). + ```json Success { "database": "acme_corp", - "added_fields": ["region", "summary_label"] + "tenant_id": "acme_corp", + "added_fields": ["region", "priority"] } ``` @@ -140,8 +127,8 @@ Each `add_fields[]` item uses the same field shape as `database_metadata_schema` "success": false, "data": null, "error": { - "code": "CONFLICT", - "message": "field \"region\" already exists in the schema" + "code": "INTERNAL_ERROR", + "message": "Schema conflict: field \"region\" already exists in the schema with a different definition. Resubmitting a field with its existing definition is accepted as a no-op, but an existing field cannot be modified. See https://docs.hydradb.com/api-reference/v2/endpoint/patch-metadata-schema for usage details. Re-send the field with its existing definition to make this a no-op, or use a different field name." }, "meta": { "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", @@ -156,10 +143,14 @@ Each `add_fields[]` item uses the same field shape as `database_metadata_schema` | Status | When it happens | | --- | --- | -| `400` | Invalid request body, empty `add_fields`, invalid field name/type, too many fields, embedding enabled on a non-`VARCHAR` field. | -| `404` | Database not found. | -| `409` | Field already exists or the update conflicts with stored database mapping/schema state. | -| `500` | Backend persistence or index creation failed. | +| `400` | Invalid request body (`INVALID_INPUT`), empty `add_fields`, invalid field name/type, `ARRAY`, too many fields, or an embedding flag on a new field. | +| `404` | Database not found (`DATABASE_NOT_FOUND`). | +| `409` | A field conflicts with an existing field or with another entry in the same request, or the database changed during the request (retry this last case). | +| `500` | Backend persistence or index creation failed. Retry. | + + + Apart from a malformed body and a missing database, this endpoint currently returns `error.code: "INTERNAL_ERROR"` for `400`, `409`, and `500` alike. Branch on the HTTP status and read `error.message`, which names the field and the rule it broke. + ## Related diff --git a/api-reference/v2/endpoint/update-source-metadata.mdx b/api-reference/v2/endpoint/update-source-metadata.mdx index 2ab9c162..79046c1f 100644 --- a/api-reference/v2/endpoint/update-source-metadata.mdx +++ b/api-reference/v2/endpoint/update-source-metadata.mdx @@ -13,7 +13,7 @@ PATCH /context/{id}/metadata ``` - The legacy route `PATCH /context/sources/{source_id}/metadata` still works but is deprecated - migrate to the route above. Both dispatch to the same handler; `source_id` and `id` name the same value. + The legacy route `PATCH /context/sources/{source_id}/metadata` still works but is deprecated. Migrate to the route above. Both dispatch to the same handler; `source_id` and `id` name the same value. @@ -38,12 +38,14 @@ curl -X PATCH 'https://api.hydradb.com/context/policy_main/metadata' \ ``` ```python Python +import os + import requests response = requests.patch( "https://api.hydradb.com/context/policy_main/metadata", headers={ - "Authorization": f"Bearer {HYDRA_DB_API_KEY}", + "Authorization": f"Bearer {os.environ['HYDRA_DB_API_KEY']}", "API-Version": "2", "Content-Type": "application/json", }, @@ -103,11 +105,12 @@ const response = await fetch("https://api.hydradb.com/context/policy_main/metada | | Collection that contains the source. This endpoint does not default it. (deprecated alias: `sub_tenant_id`) | | | Schema-backed metadata fields to merge into the source's `metadata`. Keys must satisfy the tenant metadata schema when one exists. (deprecated alias: `tenant_metadata`) | | | Free-form metadata fields to merge into the source's `additional_metadata`. | +| | Replaces the source's access-control list, so send the complete new list. `[]` or `null` revokes access for everyone. Omit it to keep the stored list. See [Access Control](/essentials/v2/access-control). | -At least one of `database_metadata` or `additional_metadata` is required. +At least one of `database_metadata`, `additional_metadata`, or `acl` is required. - This edit endpoint uses `database_metadata` for schema-backed source metadata (deprecated alias: `tenant_metadata` - still accepted, but the canonical field wins if both are sent). The shorter `metadata` field used by ingestion/list examples is not accepted in this PATCH body. `document_metadata` is also not accepted; use `additional_metadata`. + This edit endpoint uses `database_metadata` for schema-backed source metadata (deprecated alias: `tenant_metadata`, still accepted, but the canonical field wins if both are sent). The shorter `metadata` field used by ingestion/list examples is not accepted in this PATCH body. `document_metadata` is also not accepted; use `additional_metadata`. ## Behavior @@ -115,11 +118,12 @@ At least one of `database_metadata` or `additional_metadata` is required. - The update is a **merge/upsert**: - keys present in the request are inserted or overwritten - keys omitted from the request are preserved + - `acl` is the exception: when sent, it replaces the stored list - The source must already exist. This endpoint does not create sources. - The endpoint edits one source at a time. Bulk metadata edits are not supported. - Updated metadata is visible to [`/query`](/api-reference/v2/endpoint/query) metadata filters and [`/context/list`](/api-reference/v2/endpoint/list-documents) filters. - If an edited tenant metadata field has `enable_dense_embedding` or `enable_sparse_embedding`, HydraDB synchronously refreshes the relevant vector store metadata search lane. -- If the edited fields are `enable_match`-only, the edit remains MongoDB-only and `vector_sync_required` is `false`. +- If no edited field has an embedding flag, the edit remains MongoDB-only and `vector_sync_required` is `false`. ## Response @@ -130,6 +134,8 @@ At least one of `database_metadata` or `additional_metadata` is required. "success": true, "data": { "id": "policy_main", + "database": "acme_corp", + "collection": "team_docs", "tenant_id": "acme_corp", "sub_tenant_id": "team_docs", "updated": true, @@ -154,6 +160,8 @@ At least one of `database_metadata` or `additional_metadata` is required. "success": true, "data": { "id": "policy_main", + "database": "acme_corp", + "collection": "team_docs", "tenant_id": "acme_corp", "sub_tenant_id": "team_docs", "updated": true, @@ -182,7 +190,7 @@ At least one of `database_metadata` or `additional_metadata` is required. "success": false, "data": null, "error": { - "code": "BAD_REQUEST", + "code": "INVALID_INPUT", "message": "invalid metadata edit: tenant_metadata.department must be of type VARCHAR" }, "meta": { @@ -197,15 +205,21 @@ At least one of `database_metadata` or `additional_metadata` is required. | Field | Description | | --- | --- | | | Updated source ID. | -| | Public tenant ID. | -| | Sub-tenant that contained the source. | +| | The database you named. | +| | Collection that contained the source. | +| | Deprecated alias for `database`, same value. | +| | Deprecated alias for `collection`, same value. | | | `true` when the source metadata was updated. | | | Database metadata keys included in the request. | | | Deprecated alias for `database_metadata_keys`; still emitted for backward compatibility. | | | Additional metadata keys included in the request. | | | `true` when at least one changed tenant metadata field has dense/sparse embedding enabled. | -| | Present when sync was required. `true` means the sync completed. | +| | Present (`true`) when a required sync completed. Absent when no sync was required or the sync failed. | | | Number of chunk rows synced to the vector store when sync was required. | +| | Present when a required sync failed after the metadata was saved. Retry the same edit. | +| | Present and `true` when the edit replaced the source's `acl`. | +| | Present and `true` when an `acl` edit also reached the vector store. If absent after an `acl` edit, the new list is still enforced, but an added principal may miss this source in results until it is re-indexed. | +| | Present when the edit was saved in one store but a later write failed. Retry the same edit. | | | Deprecated alias for `vector_sync_required`; still emitted for backward compatibility. | | | Deprecated alias for `vector_synced`; still emitted for backward compatibility. | | | Deprecated alias for `vector_rows_synced`; still emitted for backward compatibility. | @@ -218,25 +232,27 @@ At least one of `database_metadata` or `additional_metadata` is required. | --- | --- | | `400` | Missing `database`, missing `collection`, empty metadata payload, `document_metadata` supplied, unknown tenant metadata key when a schema exists, wrong type, reserved key, over-size payload, too-deep nesting, or `null` for a dense/sparse-enabled field. | | `404` | Source does not exist for the `(database, collection, id)` scope. | -| `500` | Metadata was written to MongoDB but dense/sparse vector store sync failed. Retry the same idempotent edit to converge. | +| `500` | The edit failed. Retry the same idempotent edit to converge. | + +If the metadata is written to MongoDB but the dense/sparse vector store sync fails, the call still returns `200`, without `vector_synced` and with `vector_sync_error` explaining why. Retry the same idempotent edit to converge. ### Size limits `database_metadata` (and its still-accepted `tenant_metadata` alias) is capped at **16 KiB**; `additional_metadata` at **1 KiB**. Each cap applies to the whole map, -measured on its compact JSON encoding in UTF-8 bytes - keys, quotes and +measured on its compact JSON encoding in UTF-8 bytes: keys, quotes and punctuation count toward the budget, so budget in bytes rather than in characters of content. `document_metadata` has no size limit here because it is **not accepted on this - endpoint at all** - any non-null value returns `400`, whatever its size. It is a + endpoint at all**: any non-null value returns `400`, whatever its size. It is a valid alias for `additional_metadata` on [`/context/ingest`](/api-reference/v2/endpoint/ingest-context), but not on this one. Send `additional_metadata`. -The cap is checked against the payload in **this** request, before the merge - not +The cap is checked against the payload in **this** request, before the merge, not against the stored map the merge produces. A small edit to an already-large map is therefore accepted, so treat the cap as a per-request budget rather than a guarantee about the final stored size. Over-cap fails the whole edit with `400` and @@ -253,7 +269,7 @@ reports both numbers: } ``` -See [Scoping using metadata → Size limits](/essentials/v2/metadata#size-limits). +See [Scoping using metadata: Size limits](/essentials/v2/metadata#size-limits). ## Related diff --git a/api-reference/v2/error-responses.mdx b/api-reference/v2/error-responses.mdx index 41d79bdc..c189b546 100644 --- a/api-reference/v2/error-responses.mdx +++ b/api-reference/v2/error-responses.mdx @@ -3,9 +3,9 @@ title: "Error Responses" description: "Response envelope, HTTP status codes, error codes, and retry patterns." --- -### Response envelope +## Response envelope -HydraDB core endpoints (`/databases`, `/context/*`, and `/query`) use the same top-level envelope for successful and failed requests. Webhook management endpoints (`/webhooks/indexing*`) return their documented response object directly and do not include this envelope. +HydraDB core endpoints (`/databases`, `/context/*`, `/query`, `/feedback`, and `/webhooks/indexing*`) use the same top-level envelope for successful and failed requests. Connector endpoints (`/connectors*`) and `PATCH /databases/{database}/metadata-schema` return their success body without this envelope, but their errors use it too. ```json { @@ -23,15 +23,15 @@ HydraDB core endpoints (`/databases`, `/context/*`, and `/query`) use the same t ``` - -Field | Description | +| Field | Description | |---|---| | `success` | `false` for errors. | -| `data` | Always `null` for error responses. | +| `data` | `null` for error responses. One exception: a failed `DELETE /context` keeps its per-ID results under `data`. | | `error.code` | Machine-readable code for programmatic handling. | | `error.message` | Human-readable explanation of what failed. | | `meta.request_id` | Request identifier. Include it when contacting support. | -| `meta.latency_ms` | Server-side processing time in milliseconds. +| `meta.latency_ms` | Server-side processing time in milliseconds. | +| `detail` | Deprecated copy of the error (`detail.error_code`, `detail.message`) kept for older clients. Read `error` instead. | Use `error.code` for branching and log `meta.request_id` for every failed request. The HTTP status tells you the class of failure; the error code tells you what to do. @@ -43,10 +43,13 @@ Use `error.code` for branching and log `meta.request_id` for every failed reques |---|---|---| | `400` | Invalid parameters or malformed request | No | | `401` | Missing, expired, or invalid API key | No | -| `403` | Authenticated, but not permitted for the resource | No | +| `402` | Your plan does not allow the request, for example one connector too many | No. Upgrade | +| `403` | Authenticated, but not permitted for the resource, or the plan's database limit is reached | No | | `404` | Database, source, memory, or related resource was not found | No | -| `409` | Conflict, usually an existing database or context item ID | Usually no | -| `422` | Well-formed request that failed validation | No | +| `409` | Conflict: the database already exists, a source is still indexing when you delete it, or a schema field conflicts | Only `SOURCE_PROCESSING`, after `Retry-After` | +| `413` | Request body or upload too large | No | +| `415` | Unsupported `Content-Type`, such as a `POST /context/ingest` body that is neither a multipart form nor JSON | No | +| `422` | Well-formed request that failed validation, or a query sent before the database finished provisioning | Only `TENANT_INFRA_NOT_READY`, once the database is ready | | `429` | Rate limit exceeded | Yes, with backoff | | `500` | Internal server error | Yes, with backoff | | `503` | Temporary service unavailability | Yes, with backoff | @@ -55,15 +58,20 @@ Use `error.code` for branching and log `meta.request_id` for every failed reques | Code | Typical status | Meaning | |---|---|---| -| `INVALID_PARAMETERS` | `400` | A required parameter is missing, malformed, or mutually incompatible with another parameter. | +| `INVALID_INPUT` | `400` | The default for a bad request: a required parameter is missing, malformed, or mutually incompatible with another parameter. Also used for `413` and `415`. | | `UNAUTHORIZED` | `401` | The `Authorization: Bearer ` header is missing or invalid. | -| `FORBIDDEN` | `403` | The API key is valid but does not have access to the requested resource, or the account/plan limit prevents the operation. | +| `PAYMENT_REQUIRED` | `402` | A plan limit refused the write, for example the connector count. | +| `FREE_PLAN_DEPRECATED` | `402` | The workspace is on the deprecated Free plan. The message and `detail.upgrade_url` carry the upgrade link. | +| `FORBIDDEN` | `403` | The API key is valid but does not have access to the requested route or resource, or the plan's database limit prevents the operation. | | `DATABASE_ALREADY_EXISTS` | `409` | `POST /databases` received a `database` (formerly `tenant_id`) that is already in use. | | `DATABASE_NOT_FOUND` | `404` | The requested database does not exist or is not visible to the current API key. | -| `SOURCE_NOT_FOUND` | `404` | The requested source or memory ID does not exist in the selected database/collection. | -| `VALIDATION_ERROR` | `422` | The request shape was valid JSON/form data, but one or more fields failed semantic validation. | +| `NOT_FOUND` | `404` | The requested resource does not exist, for example a source ID in `DELETE /context`. | +| `SOURCE_PROCESSING` | `409` | `DELETE /context` hit a source that is still indexing. Retry after `Retry-After`. | +| `VALIDATION_ERROR` | `400` or `422` | One or more fields failed semantic validation: `400` on `POST /query` for a metadata filter that does not fit the field's type, `422` on connectors for credentials that fail validation. | +| `TENANT_INFRA_NOT_READY` | `422` | `POST /query` reached a database that is still provisioning. Poll [Database Status](/api-reference/v2/endpoint/tenant-status) until `ready_for_ingestion` is `true`, then retry. | +| `CORPUS_TYPE_UNSUPPORTED` | `400` | The `type` value does not fit this endpoint or database, for example `type: "all"` on ingest. | | `RATE_LIMITED` | `429` | The API key exceeded its current rate limit. | -| `INTERNAL_ERROR` | `500` | HydraDB hit an unexpected server-side error. | +| `INTERNAL_ERROR` | `500` | HydraDB hit an unexpected server-side error. A rare auth-layer failure reports `INTERNAL_SERVER_ERROR` instead. | | `SERVICE_UNAVAILABLE` | `503` | A dependency is temporarily unavailable or the service is under load. | @@ -71,14 +79,18 @@ Endpoint pages list the most common codes for that operation. New codes may be a -**Missing database returns `DATABASE_NOT_FOUND`.** A request for a database that does not exist returns `DATABASE_NOT_FOUND` on both the canonical `/databases` routes and the deprecated `/tenants` routes. The deprecated `/tenants` routes otherwise keep their pre-rename error contract for backward compatibility: a duplicate on `POST /tenants` returns `INVALID_INPUT` (not `DATABASE_ALREADY_EXISTS`), whereas the canonical `POST /databases` returns `DATABASE_ALREADY_EXISTS` as shown above. The HTTP status is identical on both. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). +`PATCH /databases/{database}/metadata-schema` currently returns `INTERNAL_ERROR` for its `400` and `409` errors too, so branch on the HTTP status there. + + + +**Deprecated `/tenants` routes keep their pre-rename error codes.** For backward compatibility, a request for a database that does not exist returns `NOT_FOUND` on the deprecated `/tenants` routes (not `DATABASE_NOT_FOUND`), and a duplicate on `POST /tenants` returns `INVALID_INPUT` (not `DATABASE_ALREADY_EXISTS`), whereas the canonical `/databases` routes return `DATABASE_NOT_FOUND` and `DATABASE_ALREADY_EXISTS` as shown above. The HTTP status is identical on both. The route decides the code, so the old `tenant_id` field sent to a `/databases` route still gets the new codes. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). ## Ingestion error codes Asynchronous ingestion failures surface a numeric `E####` code in the `error_code` field of [`GET /context/status`](/api-reference/v2/endpoint/source-status) responses and `indexing.status_changed` [webhook](/essentials/v2/webhooks) payloads. Unlike the HTTP `error.code` values above (which describe why a *request* was rejected), these describe why a specific *item* failed to index. -Many storage- and capacity-related ingestion errors are **transient**: the pipeline retries them automatically with backoff, and they typically self-resolve within minutes. A code appearing in `error_code` does not by itself mean the item has failed permanently - only treat an item as a real failure once it reaches the terminal `errored` status. +Many storage- and capacity-related ingestion errors are **transient**: the pipeline retries them automatically with backoff, and they typically self-resolve within minutes. A code appearing in `error_code` does not by itself mean the item has failed permanently; only treat an item as a real failure once it reaches the terminal `errored` status. | Code | Meaning | Severity | |---|---|---| @@ -90,12 +102,12 @@ Many storage- and capacity-related ingestion errors are **transient**: the pipel -`E6001` is **transient**, not terminal. If you observe it on an in-flight item, keep polling [`/context/status`](/api-reference/v2/endpoint/source-status) - the item normally advances to `graph_creation` / `completed` on a subsequent retry with no action on your part. Only contact support if the item is still reported as `errored` after retries are exhausted. +`E6001` is **transient**, not terminal. If you observe it on an in-flight item, keep polling [`/context/status`](/api-reference/v2/endpoint/source-status): the item normally advances to `graph_creation` / `completed` on a subsequent retry with no action on your part. Only contact support if the item is still reported as `errored` after retries are exhausted. ## Retry pattern -Retry only transient failures: `429`, `500`, and `503`. Use exponential backoff with jitter and keep retries bounded. +Retry only transient failures: `429`, `500`, and `503`. Use exponential backoff with jitter and keep retries bounded. Two other codes clear on their own, so wait instead of backing off blindly: `409 SOURCE_PROCESSING` (wait for `Retry-After`) and `422 TENANT_INFRA_NOT_READY` (wait until the database is ready). ```typescript TypeScript SDK @@ -188,8 +200,12 @@ try { }); } catch (error) { if (error instanceof HydraDBError) { - const code = error.body?.error?.code; - const requestId = error.body?.meta?.request_id; + // error.body is typed `unknown`; it holds the raw snake_case error envelope. + const body = error.body as + | { error?: { code?: string }; meta?: { request_id?: string } } + | undefined; + const code = body?.error?.code; + const requestId = body?.meta?.request_id; if (code === "DATABASE_NOT_FOUND") { // Create or select a valid database before ingesting. @@ -243,7 +259,7 @@ Also send `API-Version: 2` on raw HTTP requests. The official SDKs set the versi ### Database not found after creation -Database creation is asynchronous. After `POST /databases`, poll [`GET /databases/status`](/api-reference/v2/endpoint/tenant-status) until `infra.scheduler_status`, `infra.graph_status`, `infra.vectorstore_status.knowledge`, and `infra.vectorstore_status.memories` are all `true`. +Database creation is asynchronous. After `POST /databases`, poll [`GET /databases/status`](/api-reference/v2/endpoint/tenant-status) until `infra.scheduler_status`, `infra.graph_status`, `infra.vectorstore_status.knowledge`, and `infra.vectorstore_status.memories` are all `true` (`infra.ready_for_ingestion` combines them). A query sent earlier returns `422 TENANT_INFRA_NOT_READY`. ### Ingestion validation errors @@ -261,10 +277,10 @@ Empty results are not always errors. Check these first: - Context status may still be `queued` or `processing`; poll [`GET /context/status`](/api-reference/v2/endpoint/source-status). - `metadata_filters` may be too restrictive or may target the wrong metadata namespace. - The query may be scoped to the wrong `database` or `collection` (formerly `tenant_id` / `sub_tenant_id`). -- The `type` value may exclude the collection you need. Use `type: "all"` when combining knowledge and memories in `POST /query`. +- The `type` value may exclude the store you need. Use `type: "all"` when combining knowledge and memories in `POST /query`. ## Related sections -- [API Reference](/api-reference/v2) - endpoint inventory and conventions -- [Ingestion Status](/api-reference/v2/endpoint/source-status) - async ingestion state -- [Query](/api-reference/v2/endpoint/query) - retrieval parameters and response shape +- [API Reference](/api-reference/v2): endpoint inventory and conventions +- [Ingestion Status](/api-reference/v2/endpoint/source-status): async ingestion state +- [Query](/api-reference/v2/endpoint/query): retrieval parameters and response shape diff --git a/api-reference/v2/index.mdx b/api-reference/v2/index.mdx index f46336a6..5b841ebc 100644 --- a/api-reference/v2/index.mdx +++ b/api-reference/v2/index.mdx @@ -6,7 +6,7 @@ description: "Single reference to all HydraDB endpoints" ## Quick links - **New to HydraDB?** Start with the [Quickstart](/get-started/v2/quickstart) -- **Prefer SDKs?** See [SDKs - Node and Python](/api-reference/v2/sdks) +- **Prefer SDKs?** See [SDKs (Node and Python)](/api-reference/v2/sdks) - **Authentication:** Every endpoint requires `Authorization: Bearer ` - **Base URL:** `https://api.hydradb.com` - **Errors:** See [Error Responses](/api-reference/v2/error-responses) @@ -16,9 +16,11 @@ description: "Single reference to all HydraDB endpoints" | Group | Purpose | When to reach for it | |---|---|---| -| [Databases](/api-reference/v2/endpoint/tenants-overview) | Create, monitor, and manage isolated workspaces | First step in any integration - and any time you need usage stats, provisioning status, or to tear down a workspace | -| [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB - document uploads, app sources, user memories, and lifecycle ops | -| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query | At query time - the only endpoint you call to feed an LLM | +| [Databases](/api-reference/v2/endpoint/tenants-overview) | Create, monitor, and manage isolated workspaces | First step in any integration, and any time you need usage stats, provisioning status, or to tear down a workspace | +| [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB: document uploads, app sources, user memories, and lifecycle ops | +| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | At query time: the only endpoint you call to feed an LLM | +| [Connectors](/api-reference/v2/endpoint/connectors-overview) | Connect, configure, sync, and manage app connectors such as Slack, GitHub, Google Drive, and Supabase | When you want app data synced into a database without writing ingest code | +| [Webhooks](/essentials/v2/webhooks) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | ## Core concepts @@ -87,27 +89,53 @@ const client = new HydraDBClient({ -SDK methods mirror the API: `client..()` maps to the corresponding endpoint. The SDKs are generated from the OpenAPI contract and set `API-Version: 2` automatically. +SDK methods mirror the API: `client..()` maps to the corresponding endpoint. The SDKs are generated from the OpenAPI contract and set `API-Version: 2` automatically. The SDK method names below are the Python names; TypeScript camelCases multi-word names (for example, `delete_collection` is `deleteCollection`). ## Full endpoint inventory | Endpoint | Method | SDK method | Purpose | Use when | |---|---|---|---|---| | [`/databases`](/api-reference/v2/endpoint/create-tenant) | `POST` | `databases.create` | Create a database | You are setting up a new isolated workspace and optional metadata schema. | -| [`/databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema) | `PATCH` | REST | Add metadata schema fields | You need to add filterable metadata fields after database creation. | +| [`/databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema) | `PATCH` | `databases.update_metadata_schema` | Add metadata schema fields | You need to add filterable metadata fields after database creation. | | [`/databases`](/api-reference/v2/endpoint/list-tenants) | `GET` | `databases.list` | List databases | You need to discover database IDs available to the current API key. | | [`/databases`](/api-reference/v2/endpoint/delete-tenant) | `DELETE` | `databases.delete` | Delete a database | You need to permanently remove a workspace and its data. | | [`/databases/status`](/api-reference/v2/endpoint/tenant-status) | `GET` | `databases.status` | Check provisioning readiness | You just created a database and need to wait before ingesting data. | -| [`/databases/collections`](/api-reference/v2/endpoint/list-sub-tenants) | `GET` | `databases.collections` / `databases.collections` | List active collections | You partition data by user, team, customer, or account and need to inspect those partitions. | +| [`/databases/collections`](/api-reference/v2/endpoint/list-sub-tenants) | `GET` | `databases.collections` | List active collections | You partition data by user, team, customer, or account and need to inspect those partitions. | +| [`/databases/collections`](/api-reference/v2/endpoint/delete-collection) | `DELETE` | `databases.delete_collection` | Delete a collection | You need to permanently remove one collection and its data without deleting the database. | | [`/databases/stats`](/api-reference/v2/endpoint/tenant-stats) | `GET` | `databases.stats` | Get usage statistics | You want to monitor object counts for a database. | | [`/context/ingest`](/api-reference/v2/endpoint/ingest-context) | `POST` | `context.ingest` | Ingest knowledge or memories | You are uploading documents, app sources, or user memories. | | [`/context/status`](/api-reference/v2/endpoint/source-status) | `GET` | `context.status` | Check processing status | You have IDs from ingestion and need to know when they are queryable. | | [`/context/inspect`](/api-reference/v2/endpoint/fetch-content) | `GET` | `context.inspect` | Inspect original source content or presigned URL | You need to display or inspect the original ingested content. | | [`/context/list`](/api-reference/v2/endpoint/list-documents) | `POST` | `context.list` | Browse knowledge or memories | You need pagination, filters, field projection, or a specific subset by `ids`. | -| [`/context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata) | `PATCH` | REST | Update source metadata | You need to merge `metadata` or `additional_metadata` onto one existing source without re-ingesting. | +| [`/context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata) | `PATCH` | `context.update_source_metadata` | Update source metadata | You need to merge `metadata` or `additional_metadata` onto one existing source without re-ingesting. | | [`/context`](/api-reference/v2/endpoint/delete-source) | `DELETE` | `context.delete` | Delete sources or memories | You need to remove one or more knowledge sources or memories by ID. | | [`/context/relations`](/api-reference/v2/endpoint/source-relations) | `GET` | `context.relations` | Inspect entity relationships | You need graph relations for a source or collection. | +| [`/context/{id}/subgraph`](/api-reference/v2/endpoint/subgraph) | `GET` | `context.subgraph` | Get the connected subgraph | You need the connected subgraph around one source. | | [`/query`](/api-reference/v2/endpoint/query) | `POST` | `query` | Unified query over knowledge, memories, or both | You need retrieval with `hybrid` or `text` query across `type: "knowledge"`, `type: "memory"`, or `type: "all"`. | +| [`/feedback`](/api-reference/v2/endpoint/submit-feedback) | `POST` | `feedback.submit` | Submit query feedback | You want to tell HydraDB whether a query returned what you needed. | +| [`/connectors/providers`](/api-reference/v2/endpoint/list-connector-providers) | `GET` | `list_providers` | List connector providers | You need the providers you can connect, or the credentials one provider needs. | +| [`/connectors`](/api-reference/v2/endpoint/create-connector) | `POST` | `connectors.create` | Create a connector | You are storing credentials for one provider account. | +| [`/connectors/{id}/discover`](/api-reference/v2/endpoint/discover-connector-resources) | `GET` | `connectors.discover` | Discover resources | You need the resources the connector's credentials can reach. | +| [`/connectors/{id}/configure`](/api-reference/v2/endpoint/configure-connector) | `POST` | `connectors.configure` | Configure a connector | You are activating resources and starting the first sync. | +| [`/connectors/{id}/status`](/api-reference/v2/endpoint/get-connector-status) | `GET` | `connectors.status` | Get connector status | You need to know whether the connector and each resource are working. | +| [`/connectors`](/api-reference/v2/endpoint/list-connectors) | `GET` | `connectors.list` | List connectors | You need your connectors and their sync state. | +| [`/connectors/{id}`](/api-reference/v2/endpoint/get-connector) | `GET` | `connectors.get` | Get a connector | You need one connector's settings. | +| [`/connectors/{id}/resources`](/api-reference/v2/endpoint/connector-resources) | `GET` | `connectors.list_resources` | List connector resources | You need each resource's settings. | +| [`/connectors/{id}/sync`](/api-reference/v2/endpoint/sync-connector) | `POST` | `connectors.sync` | Sync a connector | You want to sync now instead of waiting for the schedule. | +| [`/connectors/{id}/pause`](/api-reference/v2/endpoint/pause-connector) | `POST` | `connectors.pause` | Pause a connector | You want to turn scheduled syncs off. | +| [`/connectors/{id}/resume`](/api-reference/v2/endpoint/resume-connector) | `POST` | `connectors.resume` | Resume a connector | You want to turn scheduled syncs back on. | +| [`/connectors/{id}`](/api-reference/v2/endpoint/update-connector) | `PATCH` | `connectors.update` | Update a connector | You need to change the sync interval, connector-level instructions, or credentials. | +| [`/connectors/{id}/resources/{resource_id}`](/api-reference/v2/endpoint/update-connector-resource) | `PATCH` | `connectors.update_resource_acl` | Update a connector resource | You need to change one resource's instructions or access rule. | +| [`/connectors/{id}/resources`](/api-reference/v2/endpoint/add-connector-resource) | `POST` | `connectors.create_resource` | Add a connector resource | You want to add one new resource to an existing connector. | +| [`/connectors/{id}/resources/{resource_id}`](/api-reference/v2/endpoint/delete-connector-resource) | `DELETE` | `connectors.delete_resource` | Delete a connector resource | You want to stop syncing one resource. | +| [`/connectors/{id}`](/api-reference/v2/endpoint/delete-connector) | `DELETE` | `connectors.delete` | Delete a connector | You need to remove a connector. | +| [`/webhooks/indexing`](/api-reference/v2/endpoint/register-webhook) | `POST` | `webhooks.register` | Register a webhook | You want a `POST` to your URL when indexing finishes, instead of polling. | +| [`/webhooks/indexing`](/api-reference/v2/endpoint/get-webhook) | `GET` | `webhooks.get` | Get the webhook | You need the current webhook configuration. | +| [`/webhooks/indexing`](/api-reference/v2/endpoint/delete-webhook) | `DELETE` | `webhooks.delete` | Delete the webhook | You want to stop receiving indexing notifications. | +| [`/webhooks/indexing/test`](/api-reference/v2/endpoint/test-webhook) | `POST` | `webhooks.test` | Send a test delivery | You want to check that your endpoint receives and verifies deliveries. | +| [`/webhooks/indexing/deliveries`](/api-reference/v2/endpoint/list-webhook-deliveries) | `GET` | `webhooks.list_deliveries` | List deliveries | You need recent delivery attempts and their status. | +| [`/webhooks/indexing/deliveries/{delivery_id}`](/api-reference/v2/endpoint/get-webhook-delivery) | `GET` | `webhooks.get_delivery` | Get a delivery | You need the details of one delivery. | +| [`/webhooks/indexing/deliveries/{delivery_id}/retry`](/api-reference/v2/endpoint/retry-webhook-delivery) | `POST` | `webhooks.retry_delivery` | Retry a delivery | You want to resend a `failed` or `permanently_failed` delivery. | ## Conventions @@ -128,7 +156,7 @@ curl -X POST 'https://api.hydradb.com/query' \ }' ``` -**Response envelope:** Core v2 endpoints (`/databases`, `/context/*`, and `/query`) return a consistent envelope. Endpoint pages show the full envelope; the resource-specific payload lives under `data`. Webhook management endpoints (`/webhooks/indexing*`) are the exception: they return their documented response object directly. +**Response envelope:** Core v2 endpoints (`/databases`, `/context/*`, `/query`, `/feedback`, and `/webhooks/indexing*`) return a consistent envelope. Endpoint pages show the full envelope; the resource-specific payload lives under `data`. Connector endpoints (`/connectors*`) and `PATCH /databases/{database}/metadata-schema` are the exception: they return their success body without the envelope. ```json { @@ -142,7 +170,7 @@ curl -X POST 'https://api.hydradb.com/query' \ } ``` -Errors use the same envelope with `success: false`, `data: null`, and an `error` object containing `code` and `message`. +Errors use the same envelope on every endpoint, including the exceptions above, with `success: false`, `data: null`, and an `error` object containing `code` and `message`. `meta` may also include a `deprecation` list when a request uses a legacy `/tenants` route or a deprecated field (`tenant_id`/`sub_tenant_id`, or `sub_tenant_ids` on `/query`); each entry carries `deprecated`, a `message`, and `deprecated_since` (field-level notices also add `deprecated_field` and `preferred_field`). It is a non-breaking migration nudge (the status code is unchanged) and is accompanied by a `Deprecation: true` response header. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). @@ -150,11 +178,11 @@ Errors use the same envelope with `success: false`, `data: null`, and an `error` - **Database scoping.** Most database-scoped endpoints require a `database` (formerly `tenant_id`). Many source and query endpoints also accept an optional `collection` (formerly `sub_tenant_id`) for finer-grained scoping. If omitted, the default collection is used. The old `tenant_id`/`sub_tenant_id` names (and the old `/tenants` routes) remain accepted as deprecated aliases; sending a canonical name and its alias with **different** values returns `400`. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). -- **Async operations.** Database creation, deletion, and content ingestion are asynchronous. They return immediately after queuing. Use the relevant status endpoint to confirm completion before downstream operations. +- **Async operations.** Database creation, database and collection deletion, and content ingestion are asynchronous. They return immediately after queuing. Use the relevant status endpoint to confirm completion before downstream operations. - **Pagination.** Listing endpoints (`/context/list`) return pagination fields for browsing large result sets. -- **Parameter casing.** The REST API uses snake_case (`database`). The TypeScript SDK accepts the same snake_case keys; method names are camelCase when generated for TypeScript. The Python SDK uses snake_case throughout. +- **Parameter casing.** The REST API uses snake_case (`database`, `max_results`). The TypeScript SDK uses camelCase keys and method names (`maxResults`, `deleteCollection`) and ignores request keys it does not recognize, including snake_case ones. The Python SDK uses snake_case throughout. See [SDKs](/api-reference/v2/sdks#naming-conventions). - **Query modes.** `POST /query` supports `query_by: "hybrid"` or `"text"` and `type: "knowledge"`, `"memory"`, or `"all"`. The same `type` enum is used across ingestion, listing, deletion, and query; query additionally accepts `"all"`. @@ -166,10 +194,13 @@ Errors use the same envelope with `success: false`, `data: null`, and an `error` | `202` | Accepted (async operation queued) | | `400` | Invalid parameters | | `401` | Authentication required | +| `402` | Your plan does not allow the request | | `403` | Forbidden | | `404` | Resource not found | | `409` | Conflict (e.g., database already exists) | -| `422` | Validation error | +| `413` | Request body too large | +| `415` | Unsupported content type | +| `422` | Validation error, or a query sent before the database is ready | | `429` | Rate limit exceeded | | `500` | Internal server error | | `503` | Service unavailable | diff --git a/api-reference/v2/sdks.mdx b/api-reference/v2/sdks.mdx index 71c42030..5ef5850c 100644 --- a/api-reference/v2/sdks.mdx +++ b/api-reference/v2/sdks.mdx @@ -1,9 +1,9 @@ --- -title: "SDKs – Python and Node" +title: "SDKs: Python and Node" description: "Official Python and TypeScript/Node.js SDKs for the HydraDB API." --- -The SDKs wrap every endpoint in the [API Reference](/api-reference/v2) with typed methods and IDE autocomplete. They automatically set the `API-Version: 2` header on every request - you do not need to send it manually. +The SDKs wrap every endpoint in the [API Reference](/api-reference/v2) with typed methods and IDE autocomplete. They automatically set the `API-Version: 2` header on every request, so you do not need to send it manually. ## Installation @@ -43,7 +43,7 @@ const client = new HydraDBClient({ -**Python:** Both synchronous (`HydraDB`) and asynchronous (`AsyncHydraDB`) clients are available. They share an identical surface - choose based on your application's concurrency model. +**Python:** Both synchronous (`HydraDB`) and asynchronous (`AsyncHydraDB`) clients are available. They share an identical surface; choose based on your application's concurrency model. ## Versioning @@ -54,40 +54,43 @@ You can verify which version your client is using by inspecting the response hea ## Naming conventions -The REST API uses **snake_case** for all request and response fields, and both SDKs preserve those field names for request/response objects. TypeScript only camelCases multi-word **method names**. +The REST API uses **snake_case** for all request and response fields, and the Python SDK preserves those field names for request/response objects. TypeScript camelCases multi-word **method names** and all request and response **fields**. | | Method naming | Parameter field naming | Example | |---|---|---|---| -| **Raw HTTP / cURL** | - | snake_case | `database`, `collection`, `query_by` | +| **Raw HTTP / cURL** | N/A | snake_case | `database`, `collection`, `query_by` | | **Python SDK** | snake_case | snake_case | method: `query()`, params: `database`, `query_by` | -| **TypeScript SDK** | camelCase for multi-word methods | snake_case | method: `query()`, params: `database`, `query_by` | +| **TypeScript SDK** | camelCase for multi-word methods | camelCase | method: `query()`, params: `database`, `queryBy` | -**Use snake_case request fields in TypeScript examples too** (e.g., `client.context.list({ database: "...", page_size: 50 })`). The SDK sends the same names on the wire, matching the OpenAPI schema. +**Use camelCase request fields in TypeScript** (e.g., `client.context.list({ database: "...", pageSize: 50 })`). The SDK converts them to snake_case on the wire and silently drops keys it does not recognize, such as `page_size: 50`. JSON-string values (`documentMetadata`, `memories`, `appKnowledge`) and free-form maps (`metadataFilters`) are sent as written, so keep snake_case inside them. The Python SDK also uses snake_case throughout (e.g., `client.context.list(database="...", page_size=50)`). `database` and `collection` are the current field names (formerly `tenant_id` and `sub_tenant_id`). The old names remain accepted as deprecated aliases for full backward compatibility. -Your IDE's autocomplete and type checking work directly off the API contract - if a field is optional in the API, it's optional in the SDK. +Your IDE's autocomplete and type checking work directly off the API contract: if a field is optional in the API, it's optional in the SDK. ## SDK method structure -SDK methods are grouped under three top-level namespaces - one per `/api-reference/v2` group: +SDK methods are grouped under top-level namespaces, one per `/api-reference/v2` group: | URL prefix | SDK group | Purpose | |---|---|---| -| `/context/*` and `/context` | `client.context` | Ingest, status, fetch, list, delete, graph relations | +| `/context/*` and `/context` | `client.context` | Ingest, status, fetch, list, delete, metadata updates, graph relations, subgraph | | `/query` | `client.query` | Unified retrieval over knowledge, memories, or both | -| `/databases/*` and `/databases` | `client.databases` | Create, list, delete, status, collections, stats. Legacy `/tenants/*` paths remain as deprecated aliases. | +| `/feedback` | `client.feedback` | Feedback on query results | +| `/databases/*` and `/databases` | `client.databases` | Create, list, delete, status, collections, stats, collection deletion, metadata schema. Legacy `/tenants/*` paths remain as deprecated aliases. | +| `/connectors/*` and `/connectors` | `client.connectors` | Create, configure, sync, and manage app connectors. `GET /connectors/providers` is `client.list_providers()`. | +| `/webhooks/indexing*` | `client.webhooks` | Register, inspect, test, and delete the indexing webhook; list and retry deliveries | ## Method reference ### Context -`client.context.*` covers every flow around content lifecycle - document uploads, app sources, memories, polling, fetching, listing, deletion, and graph inspection. +`client.context.*` covers every flow around content lifecycle: document uploads, app sources, memories, polling, fetching, listing, deletion, and graph inspection. | Method | Endpoint | |---|---| @@ -97,6 +100,8 @@ SDK methods are grouped under three top-level namespaces - one per `/api-refer | `client.context.list()` | `POST /context/list` | | `client.context.delete()` | `DELETE /context` | | `client.context.relations()` | `GET /context/relations` | +| `client.context.update_source_metadata()` | `PATCH /context/{id}/metadata` | +| `client.context.subgraph()` | `GET /context/{id}/subgraph` | ### Query @@ -106,6 +111,12 @@ A single method covers all retrieval. Pick `type` (`"knowledge"`, `"memory"`, `" |---|---| | `client.query()` | `POST /query` | +### Feedback + +| Method | Endpoint | +|---|---| +| `client.feedback.submit()` | `POST /feedback` | + ### Databases | Method | Endpoint | @@ -116,12 +127,48 @@ A single method covers all retrieval. Pick `type` (`"knowledge"`, `"memory"`, `" | `client.databases.status()` | `GET /databases/status` | | `client.databases.collections()` | `GET /databases/collections` | | `client.databases.stats()` | `GET /databases/stats` | +| `client.databases.delete_collection()` | `DELETE /databases/collections` | +| `client.databases.update_metadata_schema()` | `PATCH /databases/{database}/metadata-schema` | +| `client.databases.get_metadata_schema()` | `GET /databases/{database}/metadata-schema` | + +### Connectors + +| Method | Endpoint | +|---|---| +| `client.list_providers()` | `GET /connectors/providers` | +| `client.connectors.create()` | `POST /connectors` | +| `client.connectors.discover()` | `GET /connectors/{id}/discover` | +| `client.connectors.configure()` | `POST /connectors/{id}/configure` | +| `client.connectors.status()` | `GET /connectors/{id}/status` | +| `client.connectors.list()` | `GET /connectors` | +| `client.connectors.get()` | `GET /connectors/{id}` | +| `client.connectors.list_resources()` | `GET /connectors/{id}/resources` | +| `client.connectors.sync()` | `POST /connectors/{id}/sync` | +| `client.connectors.pause()` | `POST /connectors/{id}/pause` | +| `client.connectors.resume()` | `POST /connectors/{id}/resume` | +| `client.connectors.update()` | `PATCH /connectors/{id}` | +| `client.connectors.update_resource_acl()` | `PATCH /connectors/{id}/resources/{resource_id}` | +| `client.connectors.create_resource()` | `POST /connectors/{id}/resources` | +| `client.connectors.delete_resource()` | `DELETE /connectors/{id}/resources/{resource_id}` | +| `client.connectors.delete()` | `DELETE /connectors/{id}` | + +### Webhooks + +| Method | Endpoint | +|---|---| +| `client.webhooks.register()` | `POST /webhooks/indexing` | +| `client.webhooks.get()` | `GET /webhooks/indexing` | +| `client.webhooks.delete()` | `DELETE /webhooks/indexing` | +| `client.webhooks.test()` | `POST /webhooks/indexing/test` | +| `client.webhooks.list_deliveries()` | `GET /webhooks/indexing/deliveries` | +| `client.webhooks.get_delivery()` | `GET /webhooks/indexing/deliveries/{delivery_id}` | +| `client.webhooks.retry_delivery()` | `POST /webhooks/indexing/deliveries/{delivery_id}/retry` | -The method-reference tables above use the Python (snake_case) method names. TypeScript keeps the same method names but camelCases any that are multi-word (for example, the connector method `list_resources()` in Python is `listResources()` in TypeScript). Request parameters stay snake_case in both SDKs. +The method-reference tables above use the Python (snake_case) method names. TypeScript keeps the same method names but camelCases any that are multi-word (for example, the connector method `list_resources()` in Python is `listResources()` in TypeScript). Request parameters are snake_case in Python and camelCase in TypeScript. ## Migrating from v1 -The v2 SDKs are a deliberate consolidation. A few common v1 → v2 method swaps: +The v2 SDKs are a deliberate consolidation. A few common v1 to v2 method swaps: | v1 | v2 | |---|---| @@ -156,7 +203,7 @@ A database is an isolated workspace. Most organizations create one database tota response = client.databases.create( database="my_first_database", database_metadata_schema=[ - {"name": "department", "data_type": "VARCHAR", "enable_match": True}, + {"name": "department", "data_type": "VARCHAR"}, ], ) ``` @@ -174,7 +221,7 @@ response = asyncio.run(create_database()) const response = await client.databases.create({ database: "my_first_database", databaseMetadataSchema: [ - { name: "department", dataType: "VARCHAR", enableMatch: true }, + { name: "department", dataType: "VARCHAR" }, ], }); ``` @@ -203,12 +250,13 @@ while (true) { const status = await client.databases.status({ database: "my_first_database", }); - const { schedulerStatus, graphStatus, vectorstoreStatus } = status.data.infra; + // Response fields are optional in the TypeScript types. + const infra = status.data?.infra; if ( - schedulerStatus && - graphStatus && - vectorstoreStatus.knowledge && - vectorstoreStatus.memories + infra?.schedulerStatus && + infra.graphStatus && + infra.vectorstoreStatus?.knowledge && + infra.vectorstoreStatus?.memories ) break; await new Promise((r) => setTimeout(r, 2000)); } @@ -339,7 +387,7 @@ Wait until each `indexing_status` is `completed` (or `graph_creation` if you don ### Query -A single method covers every retrieval pattern - switch `type` and `query_by` to control behavior: +A single method covers every retrieval pattern: switch `type` and `query_by` to control behavior: ```python Python SDK @@ -404,7 +452,7 @@ const exact = await client.query({ ``` -For the full parameter reference, see [Query – Overview](/api-reference/v2/endpoint/query-overview). +For the full parameter reference, see [Query: Overview](/api-reference/v2/endpoint/query-overview). ### Browse, fetch, and inspect @@ -471,7 +519,7 @@ await client.context.delete({ ## Response envelope -All responses are wrapped in a consistent envelope: +All responses, except connector endpoints and `PATCH /databases/{database}/metadata-schema`, are wrapped in a consistent envelope: ```json { @@ -503,10 +551,10 @@ Both SDKs are fully typed: The SDKs provide exact type parity with the API specification: -- **Request parameters** - every field documented in the API reference is reflected in method signatures -- **Response objects** - return types match the JSON schema for each endpoint -- **Error types** - exception structures mirror error response formats -- **Nested objects** - complex parameters and responses keep their full structure +- **Request parameters:** every field documented in the API reference is reflected in method signatures +- **Response objects:** return types match the JSON schema for each endpoint +- **Error types:** exception structures mirror error response formats +- **Nested objects:** complex parameters and responses keep their full structure ## Error handling @@ -534,11 +582,6 @@ except ApiError as exc: raise ``` - -The Python SDK raises a typed exception per status from `hydra_db.errors`. Each subclasses `ApiError` and carries `status_code`, `headers`, and the parsed response `body`, from which you can read `body["error"]["code"]`. - -Not every status has its own class - `401` and `429` do not. Catch the base `ApiError` and branch on `status_code`, as above, whenever you need to handle those. - ```typescript TypeScript SDK import { HydraDBError } from "@hydradb/sdk"; @@ -550,7 +593,8 @@ try { }); } catch (error) { if (error instanceof HydraDBError) { - const code = error.body?.error?.code; + // error.body is typed `unknown`; it holds the raw snake_case error envelope. + const code = (error.body as { error?: { code?: string } } | undefined)?.error?.code; if (code === "DATABASE_NOT_FOUND") { // Handle missing database @@ -566,22 +610,28 @@ try { ``` + +The Python SDK raises a typed exception per status from `hydra_db.errors`. Each subclasses `ApiError` and carries `status_code`, `headers`, and the parsed response `body`, from which you can read `body["error"]["code"]`. + +Not every status has its own class: `401` and `429` do not. Catch the base `ApiError` and branch on `status_code`, as above, whenever you need to handle those. + + For the full list of error codes and retry patterns, see [Error Responses](/api-reference/v2/error-responses). ## IDE-driven discovery Whether you're using TypeScript, Python, VS Code, PyCharm, or any modern IDE, the workflow is the same: -1. Type the method name → see all available methods -2. Open the parentheses → see all required and optional parameters -3. Press `Cmd+Space` (macOS) or `Ctrl+Space` (Windows/Linux) → get inline documentation +1. Type the method name to see all available methods +2. Open the parentheses to see all required and optional parameters +3. Press `Cmd+Space` (macOS) or `Ctrl+Space` (Windows/Linux) to get inline documentation This works because the SDKs are fully typed with comprehensive parameter docs sourced from the OpenAPI spec. ## Related sections -- [API Reference](/api-reference/v2) - complete endpoint documentation -- [Error Responses](/api-reference/v2/error-responses) - HTTP codes, error codes, retry patterns -- [Quickstart](/get-started/v2/quickstart) - build your first integration in five minutes -- [Query](/essentials/v2/query) - conceptual overview of `POST /query` -- [Knowledge](/essentials/v2/knowledge) and [Memories](/essentials/v2/memories) - ingestion deep-dives +- [API Reference](/api-reference/v2): complete endpoint documentation +- [Error Responses](/api-reference/v2/error-responses): HTTP codes, error codes, retry patterns +- [Quickstart](/get-started/v2/quickstart): build your first integration in five minutes +- [Query](/essentials/v2/query): conceptual overview of `POST /query` +- [Knowledge](/essentials/v2/knowledge) and [Memories](/essentials/v2/memories): ingestion deep-dives diff --git a/essentials/v2/api-results.mdx b/essentials/v2/api-results.mdx index 5b52b65a..6ebce124 100644 --- a/essentials/v2/api-results.mdx +++ b/essentials/v2/api-results.mdx @@ -5,8 +5,6 @@ description: "Turn a query response into an LLM prompt." `POST /query` returns structured JSON. Before you can pass it to an LLM, you need to convert the retrieval payload into a plain string. This page shows how. -The same pattern works regardless of whether you set `type: "knowledge"`, `type: "memory"`, or `type: "all"` - the response shape is identical. - --- ## 1. What the response looks like @@ -75,12 +73,12 @@ The retrieval payload has the same core shape regardless of `type` or `query_by` Four things matter for prompt construction: -- **`chunks`** - the primary retrieval output. Ranked by relevance; preserve the order HydraDB returns. -- **`graph_context.query_paths`** - entity traversal paths derived from your query. Useful for relational reasoning. See [Context Graphs](/essentials/v2/context-graphs). -- **`graph_context.chunk_relations`** + **`chunk_id_to_group_ids`** - per-chunk graph relations grouped by `group_id`, so you can attach the right triplets to each chunk. -- **`additional_context`** - a map keyed by `chunk_uuid`. When a chunk includes `extra_context_ids`, use those IDs to look up related chunks here. +- **`chunks`**: the primary retrieval output. Ranked by relevance; preserve the order HydraDB returns. +- **`graph_context.query_paths`**: entity traversal paths derived from your query. Useful for relational reasoning. See [Context Graphs](/essentials/v2/context-graphs). +- **`graph_context.chunk_relations`** + **`chunk_id_to_group_ids`**: per-chunk graph relations grouped by `group_id`, so you can attach the right triplets to each chunk. +- **`additional_context`**: a map keyed by `chunk_uuid`. When a chunk includes `extra_context_ids`, use those IDs to look up related chunks here. -The raw object has too much noise for an LLM - IDs, timestamps, metadata. Section 2 shows how to convert it into a clean string. +The raw object has too much noise for an LLM: IDs, timestamps, metadata. Section 2 shows how to convert it into a clean string. --- @@ -285,7 +283,9 @@ asyncio.run(main()) ## 4. Combining Knowledge and Memories -The simplest path is one `POST /query` with `type: "all"` - HydraDB queries both stores in parallel and returns one merged, ranked result set. +The simplest path is one `POST /query` with `type: "all"`. HydraDB queries both stores in parallel and returns one merged, ranked result set. + +`type: "all"` reads both stores from the same scope, so this example assumes the shared documents are in `user_123` too. If they live in their own collection, send `collections: ["company_docs", "user_123"]` instead of `collection`. ```python Python SDK @@ -399,12 +399,13 @@ from hydra_db import AsyncHydraDB from hydra_db.helpers import build_string hydra = AsyncHydraDB(token="YOUR_HYDRA_DB_API_KEY") +question = "What is our refund policy?" async def main(): knowledge_result, memory_result = await asyncio.gather( hydra.query( database="acme_corp", - query="refund policy", + query=question, type="knowledge", query_by="hybrid", mode="thinking", @@ -433,11 +434,12 @@ import { HydraDBClient } from "@hydradb/sdk"; import { buildString } from "@hydradb/sdk/helpers"; const hydra = new HydraDBClient({ token: process.env.HYDRA_DB_API_KEY }); +const question = "What is our refund policy?"; const [knowledgeResult, memoryResult] = await Promise.all([ hydra.query({ database: "acme_corp", - query: "refund policy", + query: question, type: "knowledge", queryBy: "hybrid", mode: "thinking", @@ -458,7 +460,7 @@ const prompt = ``` -If memory query fails or times out, fall back to the knowledge-only prompt. +If the memory query fails or times out, fall back to the knowledge-only prompt. --- @@ -467,9 +469,8 @@ If memory query fails or times out, fall back to the knowledge-only prompt. - **Preserve server order.** Don't re-sort chunks client-side. - **Start small on chunks.** `max_results: 10` is a reasonable default. Drop to 5 if you hit token limits, raise to 20 if you rerank downstream. - **Use graph context selectively.** It improves relational queries and bloats simple lookups. See [Context Graphs](/essentials/v2/context-graphs). -- **Always give the model a grounding instruction.** A system prompt like "answer only from the provided context" prevents the model from inventing answers when retrieval is thin. +- **Give the model a grounding instruction:** Use a system prompt like "answer only from the provided context" to reduce unsupported answers when retrieval is thin. - **Format consistently.** Whatever section delimiters you choose (`=== CONTEXT ===`, `Chunk N`, `Source:`), keep them stable across calls so the model learns the structure. -- **`type: "all"` is the simplest path when you need both knowledge and memories.** One call, one result set, one `build_string` call. --- @@ -482,18 +483,18 @@ If memory query fails or times out, fall back to the knowledge-only prompt. | Including too many chunks | Token overflow or answer quality drops | Start at `max_results: 10`; reduce if needed. | | Re-sorting chunks client-side | Overrides HydraDB's ranking | Preserve the server-returned order. | | Missing a grounding instruction | The model invents answers when retrieval is thin | System prompt: answer only from the provided context. | -| Setting `graph_context: false` and expecting graph fields | Graph context will be omitted | `graph_context` is on by default; only set it to `false` when you explicitly don't want graph data. | -| Using `query_forceful_relations` with `mode: "fast"` | Flag is silently ignored | `query_forceful_relations` only takes effect when `mode: "thinking"`. | +| Setting `graph_context: false` and expecting graph fields | Graph context will be omitted (in fast mode; thinking always returns it) | `graph_context` is on by default; only set it to `false` when you explicitly don't want graph data. | +| Using `query_forceful_relations` with `mode: "fast"` | Flag is ignored | `query_forceful_relations` only takes effect when `mode: "thinking"` (or `"auto"` routed to thinking). | | Looking up `chunk_relations` without `chunk_id_to_group_ids` | The right relations don't attach to the right chunk | Use `chunk_id_to_group_ids[chunk_uuid]` to find the `group_id`s, then filter `chunk_relations` by them. | --- ## Related -- [Query](/essentials/v2/query) - request parameters and response shape -- [Memories](/essentials/v2/memories) - what `POST /query` with `type: "memory"` queries -- [Knowledge](/essentials/v2/knowledge) - what `POST /query` with `type: "knowledge"` queries -- [Context Graphs](/essentials/v2/context-graphs) - what `graph_context` fields contain +- [Query](/essentials/v2/query): request parameters and response shape +- [Memories](/essentials/v2/memories): what `POST /query` with `type: "memory"` queries +- [Knowledge](/essentials/v2/knowledge): what `POST /query` with `type: "knowledge"` queries +- [Context Graphs](/essentials/v2/context-graphs): what `graph_context` fields contain --- From 62b470f7ea08cee98edc13ab2eef9946271e5973 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Tue, 6 Oct 2026 00:47:14 +0530 Subject: [PATCH 02/11] docs(v2): line-by-line pass over the API Reference pages Read every API Reference page and api-results line by line, keeping the original voice and cutting only repetition and wrong claims: - Query: drop the alias Note and the Default Behaviors block that restated the field table; split the indexing and collection-scope advice; link Recommended configurations by name. - Query overview: new description, removed the duplicate Tip, and the personalized recipe now says to list both collections when shared docs live elsewhere. Thinking mode always includes graph context. - Submit Feedback: removed a Note that repeated the field rules. - Error Responses: removed the E6001 Note that repeated the table, and the troubleshooting bullet no longer names the deprecated tenant_metadata. - Labels end with colons throughout. Connector and webhook pages were read and left as they are. Refs PRO-2457 Co-Authored-By: Claude Opus 5.5 (1M context) Signed-off-by: SohamRatnaparkhi --- api-reference/v2/endpoint/create-tenant.mdx | 10 +-- .../v2/endpoint/delete-collection.mdx | 5 -- api-reference/v2/endpoint/delete-source.mdx | 8 +-- api-reference/v2/endpoint/delete-tenant.mdx | 7 +- api-reference/v2/endpoint/fetch-content.mdx | 16 +---- api-reference/v2/endpoint/ingest-context.mdx | 54 ++++++++------- api-reference/v2/endpoint/list-documents.mdx | 15 ++-- .../v2/endpoint/list-sub-tenants.mdx | 3 +- api-reference/v2/endpoint/list-tenants.mdx | 7 +- api-reference/v2/endpoint/query-overview.mdx | 12 ++-- api-reference/v2/endpoint/query.mdx | 32 +++------ .../v2/endpoint/source-relations.mdx | 7 +- api-reference/v2/endpoint/source-status.mdx | 14 +--- .../v2/endpoint/sources-overview.mdx | 56 +++++++-------- api-reference/v2/endpoint/subgraph.mdx | 3 +- api-reference/v2/endpoint/submit-feedback.mdx | 4 -- api-reference/v2/endpoint/tenant-stats.mdx | 10 ++- api-reference/v2/endpoint/tenant-status.mdx | 4 -- .../v2/endpoint/tenants-overview.mdx | 15 ++-- .../v2/endpoint/update-metadata-schema.mdx | 10 +-- .../v2/endpoint/update-source-metadata.mdx | 18 ++--- api-reference/v2/error-responses.mdx | 6 +- api-reference/v2/index.mdx | 18 +++-- api-reference/v2/sdks.mdx | 68 +++---------------- essentials/v2/api-results.mdx | 8 +-- 25 files changed, 136 insertions(+), 274 deletions(-) diff --git a/api-reference/v2/endpoint/create-tenant.mdx b/api-reference/v2/endpoint/create-tenant.mdx index 279feae1..84fae23d 100644 --- a/api-reference/v2/endpoint/create-tenant.mdx +++ b/api-reference/v2/endpoint/create-tenant.mdx @@ -1,6 +1,6 @@ --- title: "Create Database" -description: "Creates a space for storing context." +description: "Create an isolated database, with an optional metadata schema." openapi: "api-reference/v2/openapi.json POST /databases" --- @@ -126,18 +126,18 @@ Always check if a database is ready before using it. Use [Database Status](/api- ## What happens after database creation? -1. Create the database with `POST /databases` +1. **Wait for provisioning:** creation is asynchronous. Poll [Database Status](/api-reference/v2/endpoint/tenant-status) until `infra.ready_for_ingestion` is `true`. 2. **Default collection:** No collection exists until your first write. The first time you ingest without an explicit `collection`, HydraDB creates the database's default collection, which then stores all context written without a `collection`. Create additional collections at any time to scope data to users, teams, or projects. 3. **Retry failed databases:** If a database appears in `data.failed_databases` from [List Databases](/api-reference/v2/endpoint/list-tenants), re-create that database with `POST /databases` after addressing the reported issue. Poll status again before ingestion. -4. Start [ingesting context](/api-reference/v2/endpoint/ingest-context) once databases are ready -5. Check status of [ingestion](/api-reference/v2/endpoint/source-status). Start querying the database once the recently ingested sources show `graph_creation` (searchable) or `completed` +4. **Ingest:** start [ingesting context](/api-reference/v2/endpoint/ingest-context) once the database is ready. +5. **Query:** check [ingestion status](/api-reference/v2/endpoint/source-status), and start querying once sources show `graph_creation` (searchable) or `completed`. --- ## Defining metadata schema - Schema field names are **immutable** after database creation. You can add per-document free-form metadata fields at ingestion time, and add new database-level fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but updates are additive only: no delete, rename, or type change. Dense and sparse metadata lanes (`enable_dense_embedding`, `enable_sparse_embedding`) can only be declared here, at creation. Plan your schema carefully before creating the database. + Schema field names are **immutable** after database creation. You can add per-document free-form metadata fields at ingestion time, and add new database-level fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but updates are additive only: no delete, rename, or type change. Dense and sparse metadata lanes (`enable_dense_embedding`, `enable_sparse_embedding`) can only be declared here, at creation. Plan your schema before you create the database. You can define a custom schema at database creation to declare the `metadata` fields you filter on, and to enable semantic/BM25 search over metadata text fields (`enable_dense_embedding` / `enable_sparse_embedding`). Each dense or sparse flag adds one vector field, so a field with both uses two; a database can have at most 6. Going over returns `400`, as does declaring an `ARRAY` field. diff --git a/api-reference/v2/endpoint/delete-collection.mdx b/api-reference/v2/endpoint/delete-collection.mdx index b7b5f5c9..eda6c2a4 100644 --- a/api-reference/v2/endpoint/delete-collection.mdx +++ b/api-reference/v2/endpoint/delete-collection.mdx @@ -84,11 +84,6 @@ You do not need one to reuse the name safely. Ingestion creates a missing collec ## Behavior notes - -**Irreversible action.** Ingested documents, memories, embeddings, graph nodes, and storage objects for this collection are permanently removed. Other collections in the same database are not touched. There is no recovery window. - - -- **Async cleanup:** The endpoint returns immediately after accepting the request. Cleanup of vector stores, graphs, and storage objects runs in the background. - **Repeat calls are the retry path:** Deleting the same collection again is idempotent. A duplicate call while cleanup is still running joins the delete in progress rather than starting a second one. If a cleanup fails part-way, the collection stays fenced and re-issuing the same `DELETE` re-runs it. - **Stopping work first is still kinder:** The API cancels this collection's in-flight ingestion for you, but a job cancelled mid-run is reported as failed to whatever started it. Draining your own writers first avoids that noise. - **Dashboard:** Owners can also expand a database on the Databases page and delete a collection from the inline list. diff --git a/api-reference/v2/endpoint/delete-source.mdx b/api-reference/v2/endpoint/delete-source.mdx index df29ca80..c604d1f8 100644 --- a/api-reference/v2/endpoint/delete-source.mdx +++ b/api-reference/v2/endpoint/delete-source.mdx @@ -13,7 +13,7 @@ Specify the resource category with the `type` parameter: Pass one or more IDs in `ids`. Send `database`, `collection`, `ids`, and `type` as top-level fields in the request body. -### Knowledge deletion +### Knowledge deletion Use `type: "knowledge"` and pass knowledge `ids` in the `ids` array. Include the same `collection` you used when ingesting the knowledge; omitting it targets the default collection. @@ -76,7 +76,7 @@ curl -X DELETE 'https://api.hydradb.com/context' \ -### Memory deletion +### Memory deletion Use `type: "memory"` and pass memory `ids` in the `ids` array. Include the same `collection` you used when ingesting the memories; omitting it targets the default collection. @@ -251,7 +251,7 @@ outcomes on a failure exactly as you would on success: The header always wins. Without it, the server default applies. -| Request | Behaviour | +| Request | Behavior | | --- | --- | | No header | The server default, currently `legacy`, so `200` for every outcome. | | `X-HydraDB-Delete-Status: strict` | Honest `404` / `409` / `500`. | @@ -270,7 +270,7 @@ The header always wins. Without it, the server default applies. surfaces. -## Some additional notes +## Notes - **Partial-success semantics:** For `type=knowledge`, each ID is reported independently in `results[]`. A failure on one ID does not stop the rest. For `type=memory`, the response reports an aggregate `user_memory_deleted` reflecting _all_ listed IDs. One exception: if any source in the request is still indexing, the whole request is refused and nothing is deleted (reported as `409` in strict mode, and as a `200` with `deleted_count: 0` by default). - **Retrieval drops the source immediately:** Even before background cleanup finishes, deleted IDs disappear from `/query` and `/context/list` responses. diff --git a/api-reference/v2/endpoint/delete-tenant.mdx b/api-reference/v2/endpoint/delete-tenant.mdx index 2d5405b6..b318f70b 100644 --- a/api-reference/v2/endpoint/delete-tenant.mdx +++ b/api-reference/v2/endpoint/delete-tenant.mdx @@ -70,18 +70,13 @@ curl -X DELETE 'https://api.hydradb.com/databases?database=database_to_delete' \ ## Deletion completion -Deletion is asynchronous. Treat deletion as complete when the database no longer appears in `GET /databases`, or when `GET /databases/status?database=...` returns `DATABASE_NOT_FOUND`. +Deletion is asynchronous: the endpoint returns once the database is deregistered, and cleanup of vector stores, graphs, and storage objects runs in the background for a few minutes. Treat deletion as complete when the database no longer appears in `GET /databases`, or when `GET /databases/status?database=...` returns `DATABASE_NOT_FOUND`. After deletion completes, the same `database` can be used in a new `POST /databases` request. Until then, avoid recreating the database or retrying ingestion/query against it. ## Behavior notes - -**Irreversible action.** Ingested documents, memories, embeddings, graph nodes, metadata schema, and storage objects are permanently removed. There is no recovery window, so ensure you have a backup if the content matters. - - - **Stop in-flight work first:** Stop all ingestion, polling, query, and background jobs targeting this database before deleting. Calls made after deregistration can fail with `DATABASE_NOT_FOUND` even while infrastructure cleanup is still running. -- **Async cleanup:** The endpoint returns immediately after deregistering the database. Infrastructure cleanup of vector stores, graphs, and storage objects runs in the background and may take a few minutes to complete. - **Repeat calls:** Deleting an already-deleted database returns `404 DATABASE_NOT_FOUND`. Deleting a database that is still deleting returns `200` again. Deleting a database that is still provisioning, or whose provisioning or earlier deletion failed, is treated as a request to tear down that database. ## Errors diff --git a/api-reference/v2/endpoint/fetch-content.mdx b/api-reference/v2/endpoint/fetch-content.mdx index 4531f9c7..c2f320d3 100644 --- a/api-reference/v2/endpoint/fetch-content.mdx +++ b/api-reference/v2/endpoint/fetch-content.mdx @@ -1,12 +1,12 @@ --- title: "Inspect Context" -description: "Inspect the content of a knowledge or memory source." +description: "Read the original content of a knowledge source or memory, or get a download URL." openapi: "api-reference/v2/openapi.json GET /context/inspect" --- import { Field } from "/snippets/field.jsx"; -Specify the `id` of the knowledge or memory you want to retrieve. +Returns the original content of one knowledge source or memory, a presigned download URL, or both. Pass its `id`. @@ -56,16 +56,12 @@ curl -G 'https://api.hydradb.com/context/inspect' \ | Mode | Returns | Use when | | --- | --- | --- | -| `content` | The stored bytes. Valid UTF-8 (plain text, Markdown, CSV, memory text, app sources) comes back as text in `content`; anything else (PDF, DOCX, images) comes back base64-encoded in `content_base64`. The other field is `null`. | You want to render the content in-app or feed it to another model. | +| `content` | The stored bytes. Valid UTF-8 (plain text, Markdown, CSV, memory text, app sources) comes back as text in `content`; anything else (PDF, DOCX, images) comes back base64-encoded in `content_base64`. The other field is `null`. | You want the content inline. It is the stored original, so a PDF comes back as base64, not extracted text. Check both fields for unknown types. | | `url` | Presigned URL (`presigned_url`) valid for `expiry_seconds`. `content`, `content_base64`, and `inferred_content` are `null`. | You want a client or a service to download the original file directly without proxying through your backend. | | `both`_(default)_ | Everything `content` returns, **and** presigned URL. | UI flows that show text inline plus a "Download original" link. | In `content` and `both` modes, the response also includes `inferred_content` when available: the model-derived text for the source (for memories, the inferred memory statement; for knowledge, derived/normalized content). It is `null` when the source has no inferred content. - -Use `mode=url` when you need the original file. Use `mode=content` when you only need the content inline for display, summarization, or prompting. It is the stored original, so a PDF comes back as base64 bytes, not extracted text. - - ### Mode examples @@ -234,12 +230,6 @@ Use `mode=url` when you need the original file. Use `mode=content` when you only ## Behavior notes - - **Text vs binary handling:** In `mode=content` and `mode=both`, you get the stored original, not parsed text. Valid UTF-8 sources (TXT, MD, CSV, memory text, app sources) populate `content` with text, while binary files (PDF, DOCX, images) come back in `content_base64`; check both fields when handling unknown content types. - - -- **`inferred_content`:** Alongside the raw `content`, the response carries `inferred_content`: the model-derived text for the source. For **memories** this is the inferred memory statement (e.g. `"User prefers concise answers and dark mode."`); for **knowledge** sources it is typically `null` unless derived content exists. It is returned in `content` and `both` modes; `mode=url` returns it as `null`. - - **Recently ingested sources:** Fetching immediately after ingestion returns the stored original, but `inferred_content` stays `null` until processing produces it. For reliable reads, use [Ingestion Status](/api-reference/v2/endpoint/source-status) first. - **Presigned URL TTL:** The URL is valid only for `expiry_seconds`. Anyone with the URL can download the file during that window, so treat it as a short-lived secret. - **Memory items:** Fetching a memory's `id` returns its raw text content (a conversation memory returns its stored JSON, with `user_name` and `pairs`). Memories are stored like files, so `mode=url` and `mode=both` also return a `presigned_url`. diff --git a/api-reference/v2/endpoint/ingest-context.mdx b/api-reference/v2/endpoint/ingest-context.mdx index 58c970d1..4a3b5faa 100644 --- a/api-reference/v2/endpoint/ingest-context.mdx +++ b/api-reference/v2/endpoint/ingest-context.mdx @@ -15,10 +15,6 @@ When context is of `type=memory`: Use `memories` for per-user content, scoped with `collection`. Set `infer: true` to let HydraDB extract preferences from raw signals, or `infer: false` to store the text verbatim. Read more about [ingesting memories](/essentials/v2/memories). - -`database` and `collection` are the current field names (formerly `tenant_id` and `sub_tenant_id`). The old names remain accepted as deprecated aliases for full backward compatibility. - - ```python Python SDK @@ -322,9 +318,9 @@ response.raise_for_status() | | Memory items, **memory only**. Required and non-empty when `type=memory`. Use plural `memories` for the form field, even though `type` is singular `memory`. See the `memories` item shape below. | -1. **`id` must not contain a comma (`,`).** The comma is reserved as the id separator on [Ingestion Status](/api-reference/v2/endpoint/source-status) (`GET /context/status?ids=a,b`), so an `id` containing a comma cannot be looked up unambiguously. This applies to every `id` you supply: `document_metadata`, `app_knowledge`, and `memories` items. Ingesting an item whose `id` contains a comma is rejected with a `400`. +1. **`id` must not contain a comma (`,`):** The comma is reserved as the id separator on [Ingestion Status](/api-reference/v2/endpoint/source-status) (`GET /context/status?ids=a,b`), so an `id` containing a comma cannot be looked up unambiguously. This applies to every `id` you supply: `document_metadata`, `app_knowledge`, and `memories` items. Ingesting an item whose `id` contains a comma is rejected with a `400`. -2. **`202 Accepted` means queued, not indexed.** Ingestion is asynchronous. A successful response only confirms your sources were accepted, not that they are ready to query. Before querying, poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with the returned IDs until each source reaches `graph_creation` (searchable), `completed`, or `errored`. Alternatively, register a webhook for `indexing.status_changed` events (see [Webhooks](/essentials/v2/webhooks)). +2. **`202 Accepted` means queued, not indexed:** Ingestion is asynchronous. A successful response only confirms your sources were accepted, not that they are ready to query. Before querying, poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with the returned IDs until each source reaches `graph_creation` (searchable), `completed`, or `errored`. Alternatively, register a webhook for `indexing.status_changed` events (see [Webhooks](/essentials/v2/webhooks)). @@ -341,7 +337,7 @@ Anything in this table can be uploaded through `documents` and HydraDB will read | Images | `.png` `.jpg` `.jpeg` `.tif` `.tiff` `.webp` `.gif` `.bmp` | | Plain text | `.txt` `.md` `.markdown` `.json` | -What HydraDB extracts differs by type, and it is worth knowing which one you are uploading: +What HydraDB extracts differs by type: | Type | What gets indexed | | --- | --- | @@ -507,11 +503,12 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in app_knowledge=json.dumps([ { "id": "slack_thread_001", - "database": "acme_corp", - "collection": "team_docs", "title": "Pricing discussion", "type": "slack", - "content": {"text": "We agreed on three tiers..."}, + "kind": "message", + "provider": "slack", + "external_id": "1716213600.000100", + "fields": {"kind": "message", "body": "We agreed on three tiers...", "author": "alice"}, "metadata": {"channel": "product"}, } ]), @@ -526,11 +523,12 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in appKnowledge: JSON.stringify([ { id: "slack_thread_001", - database: "acme_corp", - collection: "team_docs", title: "Pricing discussion", type: "slack", - content: { text: "We agreed on three tiers..." }, + kind: "message", + provider: "slack", + external_id: "1716213600.000100", + fields: { kind: "message", body: "We agreed on three tiers...", author: "alice" }, metadata: { channel: "product" }, }, ]), @@ -547,11 +545,12 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in -F 'app_knowledge=[ { "id": "slack_thread_001", - "database": "acme_corp", - "collection": "team_docs", "title": "Pricing discussion", "type": "slack", - "content": { "text": "We agreed on three tiers..." }, + "kind": "message", + "provider": "slack", + "external_id": "1716213600.000100", + "fields": { "kind": "message", "body": "We agreed on three tiers...", "author": "alice" }, "metadata": { "channel": "product" } } ]' @@ -562,18 +561,22 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in | Field | Description | | --- | --- | | | Context ID. Treated as the upsert key. Send an empty string to have one generated upstream. Must not contain a comma (`,`). It is reserved as the id separator on `/context/status?ids=`. | - | | Target database. Must match the form-level `database`. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | - | | Logical partition inside the database. Must match the form-level `collection`. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). | + | | Optional. If sent, must match the form-level `database`. Formerly `tenant_id` (deprecated alias). | + | | Optional. If sent, must match the form-level `collection`. Formerly `sub_tenant_id` (deprecated alias). | | | Short title or subject shown in search results. | | | Source category (`slack`, `notion`, `gmail`, `webpage`, etc.). Used for filtering and display. | | | Optional long-form description. | | | Canonical URL or reference link. | | | ISO-8601 timestamp (creation or last-updated). | - | | Content payload. Use `{ "text": "..." }` for plain text. Required for app sources. | + | | App object kind: `email`, `message`, `ticket`, `knowledge_base`, `comment`, `meeting`, or `custom`. Selects the parser for `fields`. | + | | App namespace such as `slack`, `gmail`, `jira`, or `notion`. | + | | The provider's ID for this exact item. Enables exact lookup and relation resolution. | + | | App-native content and structure (`body`, `description`, `title`, or `data`, by `kind`). `fields.kind` must match `kind`. See [App Sources](/essentials/v2/app-sources). | + | | Older generic payload, `{ "text": "..." }`. Prefer `kind` and `fields`, which let app-aware retrieval see authors, threads, and parents. | | | Database-schema fields. At most 16 KiB. (default=`{}`) | | | Free-form per-document fields. At most 1 KiB. (default=`{}`) | | | Optional related attachments. (default=`[]`) | - | | Forceful relations, same shape as on `document_metadata` items. | + | | Relations to other sources: `{ "ids": [...] }` as on `document_metadata` items, or the typed `[{ "predicate", "target" }]` list described in [App Sources](/essentials/v2/app-sources#6-relations). | | | Principals allowed to retrieve this source. Omit to leave it unrestricted. A malformed principal rejects the whole request with `400`. See [Access Control](/essentials/v2/access-control). | @@ -794,15 +797,14 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in | | Optional timing info (e.g. "since 2021", "in Q3"). | - **Per-source replace mode:** Each top-level key must match a `document_metadata` id or `app_knowledge` item id for `type=knowledge`, or a memory `id` for `type=memory`, in the same request; attach graphs to multiple sources at once. Extraction is skipped for keyed sources. Caps per graph: at most 5,000 entities, 10,000 relations, and 500 relations per entity; over-cap returns `400`. Graphs survive re-ingest (re-upload or connector re-sync re-applies the stored graph). + **Caps per graph:** at most 5,000 entities, 10,000 relations, and 500 relations per entity; over-cap returns `400`. Graphs survive re-ingest: a re-upload or connector re-sync re-applies the stored graph. -## Some important notes +## Notes -- **Async indexing:** `202 Accepted` means HydraDB queued the work, not that content is searchable. Poll [Ingestion Status](/api-reference/v2/endpoint/source-status) until `indexing_status` reaches `graph_creation` (searchable) or `completed` (graph-ready). - **Multipart, not JSON:** This endpoint uses `multipart/form-data`. Stringify all JSON arrays (`document_metadata`, `app_knowledge`, `memories`) before placing them in the form field. A body that is neither a form nor JSON returns `415`. -- **Declare hot schema fields upfront:** Put frequently filtered fields in `metadata`, define them in `database_metadata_schema`, and use `additional_metadata` for free-form display/bookkeeping fields. Define filterable fields when creating the database via [Create Database](/api-reference/v2/endpoint/create-tenant). Additive schema updates exist ([Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema)), but dense/sparse metadata lanes can only be declared at database creation. +- **Declare hot schema fields upfront:** put frequently filtered fields in `metadata` and declare them when you [create the database](/api-reference/v2/endpoint/create-tenant); use `additional_metadata` for free-form fields. You can add fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but dense and sparse lanes can only be declared at creation. - **Memory vs knowledge:** Use `type: "memory"` for memory ingestion, listing, and deletion. Use `type: "all"` on `POST /query` when results should combine both. The multipart field name for memories is always `memories`. - **Collection defaulting:** Omitting `collection` writes to the default collection, created on first write. List available collections with [List Collections](/api-reference/v2/endpoint/list-sub-tenants). @@ -813,7 +815,7 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in - **Always check** [ingestion status](/api-reference/v2/endpoint/source-status) to ensure context is ready to be retrieved - [Query](/api-reference/v2/endpoint/query) once context is ready - - **Inspect:** [List Documents](/api-reference/v2/endpoint/list-documents) helps you fetch titles and descriptions of ingested context - - **Inspect:** [Fetch Content](/api-reference/v2/endpoint/fetch-content) helps you fetch full context of a document, memory, knowledge item + - **Inspect:** [List Documents](/api-reference/v2/endpoint/list-documents) to browse what you ingested + - **Inspect:** [Fetch Content](/api-reference/v2/endpoint/fetch-content) to read one source's original content - **Cleanup:** [Delete Context](/api-reference/v2/endpoint/delete-source) diff --git a/api-reference/v2/endpoint/list-documents.mdx b/api-reference/v2/endpoint/list-documents.mdx index e6db571e..7d65063d 100644 --- a/api-reference/v2/endpoint/list-documents.mdx +++ b/api-reference/v2/endpoint/list-documents.mdx @@ -1,6 +1,6 @@ --- title: "List Context" -description: "Browse over knowledge or memories with optional filters. Results are paginated. " +description: "Page through knowledge sources or memories, with optional filters." openapi: "api-reference/v2/openapi.json POST /context/list" --- @@ -73,14 +73,14 @@ curl -X POST 'https://api.hydradb.com/context/list' \ | | When provided and non-empty, only items with these IDs are returned (pagination \+ filters still apply). At most `100` IDs. (default=`null`) | | | Page number (1-indexed). (default=`1`) | | | Items per page, from `1` to `100`. (default=`50`) | -| | Structured exact-match filters. See [Filters](#1-filters). (default=`null`) | +| | Structured exact-match filters. See [Filters](#filters). (default=`null`) | | | Field projection. Only the listed fields plus `id`, `database`, `collection` are populated. Only applies to `type=knowledge`. (default=`null`, all fields) | | | Principals to list as (document ACLs). Only knowledge sources those principals may see come back. Omitted, `[]`, or `["*"]` lists everything. Ignored for `type=memory`. See [Access Control](/essentials/v2/access-control). | -### 1. Filters +### Filters - `filters` is a structured object with three optional categories. Filters are exact-match constraints i.e. filtered values are matched against stored values as exact values (except `source_fields.title`, a case-insensitive prefix match). There are no range, contains, or OR operators on this endpoint; run multiple calls and merge client-side for OR behavior. A `null` filter value returns `400`. -- **AND/OR:** All filter pairs combine with a logical AND. To express OR semantics, run multiple calls and union them client-side. +- **AND:** all filter pairs combine with a logical AND. - **`ids` \+ filters:** When `ids` is non-empty, only those IDs are considered, but other `filters` still apply on top: useful for "show me items 1, 2, 3 that also belong to department=legal". ```json @@ -99,7 +99,7 @@ curl -X POST 'https://api.hydradb.com/context/list' \ | | Context item's `additional_metadata` payload | Free-form per-document JSON. No schema declaration required. `document_metadata` is accepted as a legacy alias. | | | Built-in source fields (`type`, `title`, `description`, `url`, `timestamp`) and connector fields (`app_provider`, `app_kind`, `app_external_id`, `app_parent_id`) | Use for app-source categories or quick title lookups. `title` must be a string and matches by case-insensitive prefix: `"standup"` finds "Standup notes 2026-05-12". External IDs are unique only per provider, so pair `app_external_id` or `app_parent_id` with `app_provider`. Any other key returns `400`. | -### 2. Including Fields for convenient data objects +### Field projection When you don't need every field on every row, pass `include_fields` to keep response payloads small. Only the listed fields are populated; omitted fields should be treated as unavailable in that response. `id`, `database`, and `collection` are always returned. @@ -109,8 +109,6 @@ Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadat **Projectable vs. fetchable fields:** `content`, `url`, and `attachments` are **not** valid `include_fields` values: they are stripped from list responses, and requesting one returns `400`. Fetch them per-source via [Inspect Context](/api-reference/v2/endpoint/fetch-content). -`include_fields` only applies to `type=knowledge`. It is ignored for `type=memory`. - ```json Success @@ -208,9 +206,6 @@ When `type=memory`, `data` is a `ListUserMemoriesResponse` instead: same idea bu - **Scope fields:** every row carries `database` and `collection`, plus the deprecated `tenant_id` and `sub_tenant_id` with the same values. - **Memory listing unavailable:** if the memory store cannot be read, the call still returns `200` with an empty `user_memories` and `message` set to `"Memories temporarily unavailable"`. Check `message` before you treat an empty list as "no memories". - - **Use the canonical v2 names.** Prefer `filters.metadata` and `filters.additional_metadata`. Legacy `filters.tenant_metadata` and `filters.document_metadata` are accepted for back-compat, with canonical keys winning on conflicts. -
diff --git a/api-reference/v2/endpoint/list-sub-tenants.mdx b/api-reference/v2/endpoint/list-sub-tenants.mdx index dd0e4ebf..074873a4 100644 --- a/api-reference/v2/endpoint/list-sub-tenants.mdx +++ b/api-reference/v2/endpoint/list-sub-tenants.mdx @@ -4,8 +4,7 @@ description: "List collection IDs inside a database." openapi: "api-reference/v2/openapi.json GET /databases/collections" --- -1. The default collection is not created until the first write; no collection exists until then. Once you ingest without an explicit `collection`, the default collection is created and stores all context written without a `collection`. Create additional collections at any time to scope data to users, teams, or projects. -2. **Implicit creation.** Collections are auto-created when ingestion writes data under a new `collection`. The returned list grows organically as your application writes data under new values. +Collections are created by writing to them. A new database has none: the first ingest without a `collection` creates the default collection, and the first ingest under a new `collection` value creates that one. The list grows as your application writes. diff --git a/api-reference/v2/endpoint/list-tenants.mdx b/api-reference/v2/endpoint/list-tenants.mdx index c666bf86..dce7f5de 100644 --- a/api-reference/v2/endpoint/list-tenants.mdx +++ b/api-reference/v2/endpoint/list-tenants.mdx @@ -1,6 +1,6 @@ --- title: "List Databases" -description: "List all databases created." +description: "List the databases your API key can see, including any that failed to provision." openapi: "api-reference/v2/openapi.json GET /databases" --- @@ -80,10 +80,9 @@ curl -X GET 'https://api.hydradb.com/databases' \ -## Retry notes +## Retrying a failed database -- If provisioning failed for a database, `data.failed_databases` contains diagnostic entries as shown in the **Provisioning issue** tab. -- **Retry failed databases:** If a database appears in `data.failed_databases`, re-create that database with `POST /databases` after addressing the reported issue. Poll status again before ingestion. +A database in `data.failed_databases` carries diagnostic entries (see the **Provisioning issue** tab). Fix the reported issue, re-create it with `POST /databases`, and poll status again before ingesting.
diff --git a/api-reference/v2/endpoint/query-overview.mdx b/api-reference/v2/endpoint/query-overview.mdx index bdcc3092..cc067607 100644 --- a/api-reference/v2/endpoint/query-overview.mdx +++ b/api-reference/v2/endpoint/query-overview.mdx @@ -1,6 +1,6 @@ --- title: "Query: Overview" -description: "Quick reference for query modes, type selection, and when to call each." +description: "Choose type, query_by, and mode for a query, with recipes for common cases." --- import { Field } from "/snippets/field.jsx"; @@ -39,23 +39,19 @@ linkStyle default stroke:#64748b,stroke-width:2px; | | `"hybrid"`, `"text"` | Choose the matching method. Use `"hybrid"` by default and `"text"` for exact terms or phrases. | | | `"fast"`, `"thinking"`, `"auto"` | Choose latency vs quality, or let HydraDB decide. Use `"fast"` for low-latency paths, `"thinking"` for multi-query retrieval, reranking, and forceful-relation context, and `"auto"` to score the query and route to one of the two automatically (defaults to `"thinking"` when the signal is inconclusive; **the default if `mode` is omitted**). | | | integer | Control prompt size. Default `10`, maximum `250`. Start with `10`, reduce for tight context windows, increase only when you rerank or summarize downstream. | -| | `0.0`-`1.0` or `"auto"` | Tune hybrid query. Lower values favor BM25 keywords; higher values favor semantic similarity. Default `0.8`, which `"auto"` also resolves to. | +| | `0.0` to `1.0`, or `"auto"` | Tune hybrid query. Lower values favor BM25 keywords; higher values favor semantic similarity. Default `0.8`, which `"auto"` also resolves to. | | | object | Narrow candidates before ranking. Top-level keys match `metadata`; nested `additional_metadata` filters free-form per-source fields. | | | `string[]` or weighted object | Query one or more user/workspace/team scopes. A list uses equal normalized weights; an object like `{ "workspace_42": 2, "user_alex": 1 }` applies relative ranking weights with at most one decimal place. Max 100 collections, and each must exist. | | | boolean | Include entity/relation context with the chunks. On by default; set `false` with `mode: "fast"` for chunk-only responses (`thinking` always includes it). | | | boolean | Adds app-aware retrieval while still querying the full selected knowledge scope. Use it for better app-source matching; it does not restrict retrieval to app sources only. On by default for knowledge queries; send `false` to turn it off. | - -For filter design, read [Usage: Metadata](/essentials/v2/metadata) before creating database schemas. For exact request fields, defaults, and response shape, use [Query](/api-reference/v2/endpoint/query). - - ## Recommended configurations | User intent | Recommended config | |---|---| | Fast document RAG | `type="knowledge"`, `query_by="hybrid"`, `mode="fast"`, `max_results=5-10`, `graph_context=false` | -| Highest-quality document RAG | `type="knowledge"`, `query_by="hybrid"`, `mode="thinking"`, `graph_context=true`, `alpha="auto"` | -| Personalized answer | `type="all"`, include `collection`, `query_by="hybrid"`, `mode="thinking"` | +| Highest-quality document RAG | `type="knowledge"`, `query_by="hybrid"`, `mode="thinking"` (thinking always includes graph context) | +| Personalized answer | `type="all"`, the user's `collection`, `query_by="hybrid"`, `mode="thinking"`. If shared docs live in another collection, list both in `collections` | | User preferences only | `type="memory"`, include `collection`, `query_by="hybrid"` | | Exact keyword or phrase | `type="knowledge"`, `query_by="text"`, `operator="phrase"` | | Recent operational updates | `query_by="hybrid"`, `recency_bias=0.2-0.4`, filter to the right document type | diff --git a/api-reference/v2/endpoint/query.mdx b/api-reference/v2/endpoint/query.mdx index 2ac2318e..87245b4c 100644 --- a/api-reference/v2/endpoint/query.mdx +++ b/api-reference/v2/endpoint/query.mdx @@ -14,11 +14,8 @@ Three independent dimensions control behavior: - **`query_by`** picks **how** to match: `"hybrid"` (semantic + BM25, the default) or `"text"` (BM25 only; pair with `operator`). - **`mode`** picks **how** to rank results: `"fast"` (single-pass, low-latency), `"thinking"` (expands query, reranks, and can include forceful-relation context), or `"auto"` (scores the query and routes to `"fast"` or `"thinking"` automatically, defaulting to `"thinking"` when the signal is inconclusive; **the default if `mode` is omitted**). -Read more about choosing the perfect mode for your use case [here](/api-reference/v2/endpoint/query-overview#recommended-configurations). +[Recommended configurations](/api-reference/v2/endpoint/query-overview#recommended-configurations) maps common use cases to these three settings. - -`database` and `collection` are the current field names (formerly `tenant_id` and `sub_tenant_id`). The old names remain accepted as deprecated aliases for full backward compatibility. - @@ -444,7 +441,7 @@ result = client.query( | | Deterministic narrowing before ranking. See [Filters](#decision-matrix). Each list holds at most 500 values, and the whole object is capped at 64 KiB of compact JSON, measured after operator objects are reduced to their values; over either returns `400`. (default=`null`) | -**Tuning heuristics.** +**Tuning heuristics:**
  • alpha: start at 0.8. Lower toward 0.3 to 0.5 when the query contains literal tokens (error codes, SKUs, product names). Raise toward 0.9 for conceptual questions.
  • recency_bias: set to 0 for static reference material. Set 0.2 to 0.4 for mixed content, 0.6 to 0.8 for changelogs, news, or status updates.
  • @@ -479,7 +476,7 @@ result = client.query( | `"thinking"` | Multi-query expansion + reranking + forceful-relation context | Complex queries, customer-facing answers, anything where quality matters. | | `"auto"` *(default if `mode` is omitted)* | Scores the query before retrieval and routes to `"fast"` or `"thinking"`; defaults to `"thinking"` when the signal is inconclusive. | Mixed or unpredictable query traffic where you don't want to hand-pick per request. | - `"auto"`'s resolved pipeline isn't reported back in the response, so budget latency as thinking-level in the worst case. Omitting `mode` behaves exactly like `mode: "auto"`; set it explicitly to `"fast"` or `"thinking"` if you want a deterministic pipeline instead. + `"auto"`'s resolved pipeline isn't reported back in the response, so budget latency as thinking-level in the worst case. Set `mode` to `"fast"` or `"thinking"` when you need a predictable pipeline. @@ -719,31 +716,20 @@ A zero-result query returns empty arrays/maps rather than an error, as shown in ## Behavior notes - -**Default Behaviors** -- **`mode` defaults to `"auto"`.** Omitting `mode` entirely behaves exactly like `mode: "auto"`; set it explicitly to `"fast"` or `"thinking"` if you want a deterministic pipeline. -- **`graph_context` is on by default.** Set it to `false` if you only need ranked chunks and want to drop the graph slice from the response. `false` only takes effect in `fast` mode. -- **`recency_bias` applies a mild tilt by default.** Defaults to `0.4`; send `0` to turn the recency boost off. - - **Important Considerations & Common Mistakes** -- **`query_forceful_relations` requires `mode` to resolve to `"thinking"`.** In `fast` mode the flag is silently skipped. The server does not error or warn; your `additional_context` will simply be empty. Under `mode: "auto"` this depends on that request's routing decision, not on what you asked for. -- **`graph_context: false` only takes effect in `fast` mode.** `thinking` always includes the graph slice, so under `mode: "auto"` whether `false` is honored depends on that request's routing decision. This also applies when `mode` is omitted, since it defaults to `"auto"`. Call `"fast"` explicitly if you never want the graph slice. -- **Want a deterministic pipeline instead of automatic routing?** Set `mode` explicitly to `"fast"` or `"thinking"`; an omitted `mode` field now defaults to `"auto"`, not `"fast"`. -- **Relation `timestamp` is a Unix epoch float here.** In the `graph_context` slice returned by `/query` (and in the passthrough relations returned by [List Documents](/api-reference/v2/endpoint/list-documents) with `include_fields: ["relations"]`), each relation's `timestamp` is a Unix epoch value in seconds (a float, e.g. `1778573640.0`). The dedicated [Context Relations](/api-reference/v2/endpoint/source-relations) endpoint returns the same field as an ISO-8601 string instead. Normalize before comparing relation timestamps across endpoints. -- **Use the right metadata namespace.** Top-level `metadata_filters` keys match `metadata`; free-form per-document fields must be nested under `additional_metadata` (`document_metadata` is only a legacy alias). Declare hot top-level filter fields in `database_metadata_schema`. -- **Common mistakes.** Check [Ingestion Status](/api-reference/v2/endpoint/source-status) for recently ingested documents before querying. If you omit `collection`, HydraDB queries the default collection, and a query that names a collection does not read the default one; use [List Collections](/api-reference/v2/endpoint/list-sub-tenants) to discover available IDs. +- **`query_forceful_relations` requires `mode` to resolve to `"thinking"`:** In `fast` mode the flag is silently skipped. The server does not error or warn; your `additional_context` will simply be empty. Under `mode: "auto"` this depends on that request's routing decision, not on what you asked for. +- **Relation `timestamp` is a Unix epoch float here:** In the `graph_context` slice returned by `/query` (and in the passthrough relations returned by [List Documents](/api-reference/v2/endpoint/list-documents) with `include_fields: ["relations"]`), each relation's `timestamp` is a Unix epoch value in seconds (a float, e.g. `1778573640.0`). The dedicated [Context Relations](/api-reference/v2/endpoint/source-relations) endpoint returns the same field as an ISO-8601 string instead. Normalize before comparing relation timestamps across endpoints. +- **Use the right metadata namespace:** Top-level `metadata_filters` keys match `metadata`; free-form per-document fields must be nested under `additional_metadata` (`document_metadata` is only a legacy alias). Declare hot top-level filter fields in `database_metadata_schema`. +- **Check indexing first:** New documents are not searchable until [Ingestion Status](/api-reference/v2/endpoint/source-status) reports them indexed. +- **Collection scope:** If you omit `collection`, HydraDB queries the default collection, and a query that names a collection does not read the default one. Use [List Collections](/api-reference/v2/endpoint/list-sub-tenants) to discover available IDs. ## Errors Common codes: `400 INVALID_INPUT` (empty `query`, `operator: "and"` or `"phrase"` without `query_by: "text"`, a malformed filter operator, or a listed collection that does not exist), `400 VALIDATION_ERROR` (a metadata filter that does not fit the declared field type), `404 DATABASE_NOT_FOUND`, `422 TENANT_INFRA_NOT_READY` (the database is still provisioning; poll [Database Status](/api-reference/v2/endpoint/tenant-status) until `ready_for_ingestion` is `true`), `500 INTERNAL_ERROR`. See [Error Responses](/api-reference/v2/error-responses) for the full list. -`400` also covers oversized filters: a `metadata_filters` list above 500 values, or -a `metadata_filters` object above 64 KiB of compact JSON. The message names the -offending key or reports the actual byte count. See -[Filter size limits](/essentials/v2/metadata#filter-size-limits). +`400` also covers oversized filters: a `metadata_filters` list above 500 values, or a `metadata_filters` object above 64 KiB of compact JSON. The message names the offending key or reports the actual byte count. See [Filter size limits](/essentials/v2/metadata#filter-size-limits).
    diff --git a/api-reference/v2/endpoint/source-relations.mdx b/api-reference/v2/endpoint/source-relations.mdx index 3eabe8aa..a15e6771 100644 --- a/api-reference/v2/endpoint/source-relations.mdx +++ b/api-reference/v2/endpoint/source-relations.mdx @@ -1,7 +1,7 @@ --- title: "Inspecting Context Relations" openapi: "api-reference/v2/openapi.json GET /context/relations" -description: "See and explore relationships that create the brain for your AI. " +description: "Read the entity-relation triplets extracted from your content, for one source or a whole collection." --- import { Field } from "/snippets/field.jsx"; @@ -197,14 +197,13 @@ while True: cursor = page.next_cursor ``` -## Some additional notes +## Notes **Cursor opacity:** `next_cursor` is opaque (currently the timestamp, in Unix seconds, of the last group returned). Don't construct it client-side or assume meaning: pass back exactly what the server returned. -- **Collection-wide queries:** Omitting `id` returns relations across the entire collection. This is useful for full-graph exports; pair with a small `limit` and paginate. -- **Knowledge vs memory:** If the `id` belongs to a memory, set `type=memory`; otherwise the endpoint searches the Knowledge graph. The two graphs are completely separate. +- **Full-graph exports:** omit `id`, use a small `limit`, and paginate. - **Ordering:** `data.relations[]` comes back newest first, by each group's latest relation `timestamp`, not by relevance. - **Unknown `id`:** returns `200` with an empty `relations` list, not an error. So does an `id` the `acl` principals may not see. - **`auxiliary_relations`:** the structural graph around the sources (who sent a message, which comments hang off it), in the same triplet shape. It does not count against `limit`; `auxiliary_truncated` is `true` when it was clipped. diff --git a/api-reference/v2/endpoint/source-status.mdx b/api-reference/v2/endpoint/source-status.mdx index b217204e..dcf4e91e 100644 --- a/api-reference/v2/endpoint/source-status.mdx +++ b/api-reference/v2/endpoint/source-status.mdx @@ -1,6 +1,6 @@ --- title: "Ingestion Status" -description: "Check the processing status of ingested documents. " +description: "Check whether ingested documents, app sources, or memories are ready to query." openapi: "api-reference/v2/openapi.json GET /context/status" --- @@ -8,11 +8,11 @@ import { Field } from "/snippets/field.jsx"; Since ingestion is asynchronous, use this endpoint to determine when context is ready to be retrieved. -Pass one or more IDs in `ids` to retrieve status. Works for documents, app sources, and memories. When passing multiple IDs on the query string, use either repeated params (`?ids=policy_main&ids=runbook_deploy`) or a single comma-joined value (`?ids=policy_main,runbook_deploy`); both forms are equivalent and can be mixed. Surrounding whitespace is trimmed, empty entries are dropped, and duplicates are removed. For more information, see the [Knowledge](/essentials/v2/knowledge) and [Memories](/essentials/v2/memories) guides. +Pass the IDs you got back from ingestion in `ids` (or a single `id`). It works for documents, app sources, and memories. Whitespace is trimmed, empty entries are dropped, and duplicates are removed. **Prefer webhooks over polling?** Register a webhook for `indexing.status_changed` events and HydraDB will `POST` to your endpoint when content reaches a terminal state (`completed` or `errored`). See [Webhooks](/essentials/v2/webhooks) for setup and receiver examples. - + @@ -251,14 +251,6 @@ Typical processing time: - **Small documents** (under 50 pages): 1 to 5 minutes - **Large documents** (50\+ pages): 5 to 15 minutes -## Behavior notes - - - **`graph_creation` is searchable.** Items in this state are already retrievable via `/query`. Wait for `completed` only when you specifically need full graph traversal (`graph_context: true`). - - -- **Unknown IDs return as `errored`:** If you pass an ID that does not exist (e.g., a typo), HydraDB returns an entry with `indexing_status: "errored"` and `error_code: "FILE_NOT_FOUND"` rather than silently dropping it. Use `error_code` to distinguish this from a genuine ingestion failure. See [`error_code` values](#error-code-values). - ## Errors Common codes: `400 INVALID_INPUT` (for example `database` and `tenant_id` both sent with different values), `404 DATABASE_NOT_FOUND`, `422 VALIDATION_ERROR` (missing `database`, or no ID). See [Error Responses](/api-reference/v2/error-responses) for the full list. diff --git a/api-reference/v2/endpoint/sources-overview.mdx b/api-reference/v2/endpoint/sources-overview.mdx index c64401bf..58d40326 100644 --- a/api-reference/v2/endpoint/sources-overview.mdx +++ b/api-reference/v2/endpoint/sources-overview.mdx @@ -1,6 +1,6 @@ --- title: "Context Management: Overview" -description: "Quick reference for context management endpoints, their lifecycle, and when to use which." +description: "Every context endpoint, the ingestion lifecycle, and which call to use for which content." --- ## Endpoint references @@ -56,48 +56,44 @@ flowchart LR ``` - **Why both** `type=knowledge` **and** `app_knowledge`**?** They have a theoretical differentiation. + **Why both `type=knowledge` and `app_knowledge`?** They answer different questions. - - `type` picks the **bucket**: `knowledge` (shared documents) or `memory` (per-user context). It routes the ingest to the right store. + - `type` picks the **store**: `knowledge` (shared documents) or `memory` (per-user context). It routes the ingest to the right store. - Within `type=knowledge`, you pick the **payload shape**: `documents` (binary documents HydraDB will parse: PDFs, DOCX, CSV) or `app_knowledge` (a JSON array of already-extracted content from your app: Slack messages, Notion pages, web pages). You can send both in the same request. -## Core Ingestion Concepts +## Core ingestion concepts -- **Knowledge vs. Memories**: [Knowledge](/essentials/v2/knowledge) is shared, database-wide content (documents, app pages, Slack messages). [Memories](/essentials/v2/memories) are user-specific preferences and conversational traits scoped by `collection`. Both can be searched together via `type: "all"` on `POST /query`. -- **IDs**: Unique identifiers returned by `/context/ingest`. You can assign custom IDs using `id` in `document_metadata`, `app_knowledge`, or `memories` items. Use them for polling status, inspecting content, and deleting context. -- **Metadata Filtering**: You can scope queries using `metadata` (structured fields defined in your database schema) or `additional_metadata` (free-form per-document JSON). For detailed guidelines on structuring metadata, see the [Scoping using metadata](/essentials/v2/metadata) guide. -- **Forceful Relations**: Relationships between sources can be declared at ingestion time to construct a robust knowledge graph. For more details on the graph layer, see the [Context Graphs](/essentials/v2/context-graphs) guide. +- **Knowledge vs. memories:** [Knowledge](/essentials/v2/knowledge) is shared content (documents, app pages, Slack messages). [Memories](/essentials/v2/memories) are one user's preferences and conversation history. Both live in a `collection`, and `type: "all"` on `POST /query` searches both from the same scope. +- **IDs:** Unique identifiers returned by `/context/ingest`. You can assign custom IDs using `id` in `document_metadata`, `app_knowledge`, or `memories` items. Use them for polling status, inspecting content, and deleting context. +- **Metadata filtering:** You can scope queries using `metadata` (structured fields defined in your database schema) or `additional_metadata` (free-form per-document JSON). For detailed guidelines on structuring metadata, see the [Scoping using metadata](/essentials/v2/metadata) guide. ## Forceful relations and metadata Forceful relations let you pre-wire document relationships at ingestion time so that relevant documents surface together during retrieval, even before the graph layer discovers connections organically. Think of them as explicit "see also" links between your documents. -Paired with document-level metadata, you get deterministic control over how results are filtered and ranked. +Linked sources come back in the query response's `additional_context` (thinking mode). Set relations on the same `document_metadata` item as the file's metadata: ```python Python SDK -result = client.context.ingest( - type="knowledge", - database="acme_corp", - documents=[("runbook.pdf", f, "application/pdf")], - document_metadata=json.dumps([{ - "id": "runbook_deploy", - "metadata": {"department": "ops"}, - "additional_metadata": {"owner": "platform-team"}, - "relations": {"ids": ["monitoring_guide"]}, - }]), -) +import json + +with open("runbook.pdf", "rb") as f: + result = client.context.ingest( + type="knowledge", + database="acme_corp", + documents=[("runbook.pdf", f, "application/pdf")], + document_metadata=json.dumps([{ + "id": "runbook_deploy", + "metadata": {"department": "ops"}, + "additional_metadata": {"owner": "platform-team"}, + "relations": {"ids": ["monitoring_guide"]}, + }]), + ) ``` -## Related sections +## Related -- [Usage: Forceful Relations](/essentials/v2/knowledge): linking sources at ingestion (see section 7) +- [Knowledge: forceful relations](/essentials/v2/knowledge#7-forceful-relations): linking sources at ingestion +- [Memories](/essentials/v2/memories): memories vs knowledge, when to use which +- [Metadata](/essentials/v2/metadata): database-level vs document-level metadata - [Query](/api-reference/v2/endpoint/query-overview): retrieve ingested content - - - Related Resources - - - [Usage: Memories](/essentials/v2/memories): memories vs knowledge, when to use which - - - [Usage: Metadata](/essentials/v2/metadata): database-level vs document-level metadata - diff --git a/api-reference/v2/endpoint/subgraph.mdx b/api-reference/v2/endpoint/subgraph.mdx index 84052cb5..d5f4f67c 100644 --- a/api-reference/v2/endpoint/subgraph.mdx +++ b/api-reference/v2/endpoint/subgraph.mdx @@ -187,7 +187,7 @@ Traversal is breadth-first, so `depth` on each member is its distance from the s - **`auxiliary_relations[]`** is the structural graph around the members: which person sent a message, which entities are mentioned in it, which comments and attachments hang off it. These are recorded from the item itself, not extracted from text, so their `context` is empty. - **Not included:** the chunk-level entity relations that Query returns as `graph_context`. Those are a different read. -## Some additional notes +## Notes **An unknown `id` is an empty subgraph, not an error.** The endpoint does not confirm or deny that an item exists; the same answer comes back for an id that was never ingested and for one the `acl` principals may not see. @@ -195,7 +195,6 @@ Traversal is breadth-first, so `depth` on each member is its distance from the s - **An item nothing links to** comes back as a one-member subgraph: itself, at depth `0`, with `max_depth_reached: 0`. That is a real answer ("this stands alone"), distinct from an unknown id, which has no members. - **Bounding the traversal:** Threads and hierarchies can be large. `depth` bounds how far the walk goes; `max_sources` bounds how many members it returns. When `max_sources` clips it, `is_truncated` is `true` and the members you have are the ones closest to the start item. `auxiliary_truncated` reports the same for the structural graph. -- **Knowledge vs memory:** If the `id` belongs to a memory, set `type=memory`; the two graphs never connect to each other. - **Completeness:** An item's links populate once its `indexing_status` reaches `completed`. Items still in `graph_creation` may appear with fewer connections than they will have. - **Cost:** One request fans out into a bounded series of graph reads, so it is rate-limited like a Query, not like a status poll. diff --git a/api-reference/v2/endpoint/submit-feedback.mdx b/api-reference/v2/endpoint/submit-feedback.mdx index a25cbb84..fc62da39 100644 --- a/api-reference/v2/endpoint/submit-feedback.mdx +++ b/api-reference/v2/endpoint/submit-feedback.mdx @@ -202,10 +202,6 @@ If you already know the right answer (you are running an evaluation set, or you Send either on its own or both together. If `ground_truth` is your only signal, at least one of the two has to carry something; values that are empty or all whitespace are treated as not sent. - -When you send `ground_truth`, the `feedback` comment becomes optional: an evaluation run with an answer key does not need prose for every row. A submission with neither is rejected. - - ```python Evaluation run for case in eval_set: result = client.query(database="acme_corp", query=case.question) diff --git a/api-reference/v2/endpoint/tenant-stats.mdx b/api-reference/v2/endpoint/tenant-stats.mdx index 51a4b499..f90dd915 100644 --- a/api-reference/v2/endpoint/tenant-stats.mdx +++ b/api-reference/v2/endpoint/tenant-stats.mdx @@ -5,9 +5,7 @@ description: "Retrieve usage statistics for a database." openapi: "api-reference/v2/openapi.json GET /databases/stats" --- -Get detailed metrics including object counts. - -The response splits stats into two database-wide collections: `data.knowledge_collection` (Knowledge) and `data.memory_collection` (User memories), each with its own row count. Counts aggregate across all collections in the database. +Returns row counts for a database: one for the knowledge store (`data.knowledge_collection`) and one for memories (`data.memory_collection`), each summed across every collection in the database. @@ -89,11 +87,11 @@ curl -X GET 'https://api.hydradb.com/databases/stats?database=my_first_database' ## Behavior notes - **`row_count` is chunks, not sources.** One ingested document typically becomes many chunks (e.g., a 30-page PDF can produce 100\+ rows). To count distinct sources, use [List Documents](/api-reference/v2/endpoint/list-documents) with `page_size=1` and read `total`. + **`row_count` is chunks, not sources.** One ingested document typically becomes many chunks (e.g., a 30-page PDF can produce 100\+ rows). To count distinct sources, use [List Documents](/api-reference/v2/endpoint/list-documents) with `page_size=1` and read `pagination.total`. -- **Both collections always exist:** Even if you've only ingested knowledge (or only memories), both collections are provisioned. Empty collections report `row_count: 0`. -- **Stats are eventually consistent:** Immediately after ingestion or deletion, counts may lag until background processing completes. Check [Ingestion Status](/api-reference/v2/endpoint/source-status) for recently ingested IDs before treating counts as final. +- **Both stores always report:** even if you've only ingested knowledge (or only memories), both counts are returned. An empty store reports `row_count: 0`. +- **Stats are eventually consistent:** immediately after ingestion or deletion, counts may lag until background processing completes. Check [Ingestion Status](/api-reference/v2/endpoint/source-status) for recently ingested IDs before treating counts as final.
    diff --git a/api-reference/v2/endpoint/tenant-status.mdx b/api-reference/v2/endpoint/tenant-status.mdx index e2556529..e13b13b7 100644 --- a/api-reference/v2/endpoint/tenant-status.mdx +++ b/api-reference/v2/endpoint/tenant-status.mdx @@ -89,10 +89,6 @@ curl -X GET 'https://api.hydradb.com/databases/status?database=my_first_database **Querying too early:** a `POST /query` sent before `ready_for_ingestion` is `true` returns `422 TENANT_INFRA_NOT_READY`. - - **Common mistake:** `row_count` from [Database Stats](/api-reference/v2/endpoint/tenant-stats) counts individual chunks, not documents. For distinct source or memory counts, use [List Documents](/api-reference/v2/endpoint/list-documents) and read `pagination.total` from the response. - -
    diff --git a/api-reference/v2/endpoint/tenants-overview.mdx b/api-reference/v2/endpoint/tenants-overview.mdx index b0e5a1b4..7020382d 100644 --- a/api-reference/v2/endpoint/tenants-overview.mdx +++ b/api-reference/v2/endpoint/tenants-overview.mdx @@ -1,6 +1,6 @@ --- title: "Databases: Overview" -description: "Quick reference for all databases endpoints, their lifecycle, and when to call each." +description: "Every database endpoint, the order to call them in, and what each is for." --- Databases are physically isolated spaces for storing context. In most integrations you create a database once, wait for provisioning, then ingest and query inside it. @@ -13,17 +13,18 @@ Databases are physically isolated spaces for storing context. In most integratio | [`/databases`](/api-reference/v2/endpoint/delete-tenant) | `DELETE` | `databases.delete` | Permanently remove a database | Yes | | [`/databases`](/api-reference/v2/endpoint/list-tenants) | `GET` | `databases.list` | List all databases for the organization | No | | [`/databases/status`](/api-reference/v2/endpoint/tenant-status) | `GET` | `databases.status` | Check provisioning readiness | No | -| [`/databases/stats`](/api-reference/v2/endpoint/tenant-stats) | `GET` | `databases.stats` | Monitor database load | No | +| [`/databases/stats`](/api-reference/v2/endpoint/tenant-stats) | `GET` | `databases.stats` | Row counts for a database | No | | [`/databases/collections`](/api-reference/v2/endpoint/list-sub-tenants) | `GET` | TypeScript: `databases.collections`
    Python: `databases.collections` | List active collections | No | | [`/databases/collections`](/api-reference/v2/endpoint/delete-collection) | `DELETE` | TypeScript: `databases.deleteCollection`
    Python: `databases.delete_collection` | Permanently remove one collection | Yes | | [`/databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema) | `PATCH` | TypeScript: `databases.updateMetadataSchema`
    Python: `databases.update_metadata_schema` | Add metadata schema fields | No | +| `/databases/{database}/metadata-schema` | `GET` | TypeScript: `databases.getMetadataSchema`
    Python: `databases.get_metadata_schema` | Read the metadata schema | No | ## Typical call sequence For a new database from scratch: 1. Create the database: `POST /databases` -2. Wait for provisioning: `GET /databases/status` until `scheduler_status`, `graph_status`, `vectorstore_status.knowledge`, and `vectorstore_status.memories` are all `true` +2. Wait for provisioning: `GET /databases/status` until `infra.ready_for_ingestion` is `true` 3. Ingest content: `POST /context/ingest` 4. Wait for indexing: `GET /context/status` until sources are searchable 5. Retrieve context: `POST /query` @@ -32,11 +33,9 @@ For a new database from scratch: For routine operations on an existing database: -```text -GET /databases -> list databases in the org -GET /databases/collections -> list active collections -GET /databases/stats -> check database health & growth -``` +- `GET /databases`: list databases in the org +- `GET /databases/collections`: list active collections +- `GET /databases/stats`: row counts for knowledge and memories ## Key concepts diff --git a/api-reference/v2/endpoint/update-metadata-schema.mdx b/api-reference/v2/endpoint/update-metadata-schema.mdx index 652a8c98..273775ec 100644 --- a/api-reference/v2/endpoint/update-metadata-schema.mdx +++ b/api-reference/v2/endpoint/update-metadata-schema.mdx @@ -97,16 +97,10 @@ Each `add_fields[]` item uses the same field shape as `database_metadata_schema` ## Rules -- Additions only. - Existing field names cannot be reused with a different definition, case-insensitively. A field that differs from the existing one in type, `max_length`, or flags returns `409`, as does one name declared twice in a request with different definitions. Resending a field with its identical definition is a no-op. -- Existing fields cannot be deleted or changed. - Total custom database metadata fields cannot exceed 32. Skipped fields do not count toward the new total. - Reserved names such as `source_id`, `chunk_id`, `metadata`, and `document_metadata` are rejected. -- Dense/sparse embedding flags are rejected on new fields. Resending an existing field that already has those flags is a no-op. - - - This endpoint persists the updated schema and MongoDB filter indexes. It does not alter existing Milvus collections, so it cannot add dense/sparse metadata fields. Declare the desired semantic metadata fields at database creation, or migrate/re-ingest into a database with the final schema if those fields must participate in semantic/BM25 metadata search. - +- Resending an existing field that already has dense or sparse flags is a no-op. ## Response @@ -146,7 +140,7 @@ A success returns `database`, the deprecated `tenant_id`, and `added_fields` dir | `400` | Invalid request body (`INVALID_INPUT`), empty `add_fields`, invalid field name/type, `ARRAY`, too many fields, or an embedding flag on a new field. | | `404` | Database not found (`DATABASE_NOT_FOUND`). | | `409` | A field conflicts with an existing field or with another entry in the same request, or the database changed during the request (retry this last case). | -| `500` | Backend persistence or index creation failed. Retry. | +| `500` | Saving the schema failed. Retry. | Apart from a malformed body and a missing database, this endpoint currently returns `error.code: "INTERNAL_ERROR"` for `400`, `409`, and `500` alike. Branch on the HTTP status and read `error.message`, which names the field and the rule it broke. diff --git a/api-reference/v2/endpoint/update-source-metadata.mdx b/api-reference/v2/endpoint/update-source-metadata.mdx index 79046c1f..e63310bb 100644 --- a/api-reference/v2/endpoint/update-source-metadata.mdx +++ b/api-reference/v2/endpoint/update-source-metadata.mdx @@ -1,6 +1,6 @@ --- title: "Update Source Metadata" -description: "Merge tenant metadata and additional metadata for one existing source without re-ingesting its content." +description: "Merge database metadata and additional metadata into one existing source, without re-ingesting it." openapi: "api-reference/v2/openapi.json PATCH /context/{id}/metadata" --- @@ -103,7 +103,7 @@ const response = await fetch("https://api.hydradb.com/context/policy_main/metada | --- | --- | | | Owning database. (deprecated alias: `tenant_id`) | | | Collection that contains the source. This endpoint does not default it. (deprecated alias: `sub_tenant_id`) | -| | Schema-backed metadata fields to merge into the source's `metadata`. Keys must satisfy the tenant metadata schema when one exists. (deprecated alias: `tenant_metadata`) | +| | Schema-backed metadata fields to merge into the source's `metadata`. Keys must satisfy the database metadata schema when one exists. (deprecated alias: `tenant_metadata`) | | | Free-form metadata fields to merge into the source's `additional_metadata`. | | | Replaces the source's access-control list, so send the complete new list. `[]` or `null` revokes access for everyone. Omit it to keep the stored list. See [Access Control](/essentials/v2/access-control). | @@ -122,7 +122,7 @@ At least one of `database_metadata`, `additional_metadata`, or `acl` is required - The source must already exist. This endpoint does not create sources. - The endpoint edits one source at a time. Bulk metadata edits are not supported. - Updated metadata is visible to [`/query`](/api-reference/v2/endpoint/query) metadata filters and [`/context/list`](/api-reference/v2/endpoint/list-documents) filters. -- If an edited tenant metadata field has `enable_dense_embedding` or `enable_sparse_embedding`, HydraDB synchronously refreshes the relevant vector store metadata search lane. +- If an edited database metadata field has `enable_dense_embedding` or `enable_sparse_embedding`, HydraDB synchronously refreshes the relevant vector store metadata search lane. - If no edited field has an embedding flag, the edit remains MongoDB-only and `vector_sync_required` is `false`. ## Response @@ -213,7 +213,7 @@ At least one of `database_metadata`, `additional_metadata`, or `acl` is required | | Database metadata keys included in the request. | | | Deprecated alias for `database_metadata_keys`; still emitted for backward compatibility. | | | Additional metadata keys included in the request. | -| | `true` when at least one changed tenant metadata field has dense/sparse embedding enabled. | +| | `true` when at least one changed database metadata field has dense/sparse embedding enabled. | | | Present (`true`) when a required sync completed. Absent when no sync was required or the sync failed. | | | Number of chunk rows synced to the vector store when sync was required. | | | Present when a required sync failed after the metadata was saved. Retry the same edit. | @@ -230,7 +230,7 @@ At least one of `database_metadata`, `additional_metadata`, or `acl` is required | Status | When it happens | | --- | --- | -| `400` | Missing `database`, missing `collection`, empty metadata payload, `document_metadata` supplied, unknown tenant metadata key when a schema exists, wrong type, reserved key, over-size payload, too-deep nesting, or `null` for a dense/sparse-enabled field. | +| `400` | Missing `database`, missing `collection`, empty metadata payload, `document_metadata` supplied, unknown database metadata key when a schema exists, wrong type, reserved key, over-size payload, too-deep nesting, or `null` for a dense/sparse-enabled field. | | `404` | Source does not exist for the `(database, collection, id)` scope. | | `500` | The edit failed. Retry the same idempotent edit to converge. | @@ -244,14 +244,6 @@ measured on its compact JSON encoding in UTF-8 bytes: keys, quotes and punctuation count toward the budget, so budget in bytes rather than in characters of content. - - `document_metadata` has no size limit here because it is **not accepted on this - endpoint at all**: any non-null value returns `400`, whatever its size. It is a - valid alias for `additional_metadata` on - [`/context/ingest`](/api-reference/v2/endpoint/ingest-context), but not on this - one. Send `additional_metadata`. - - The cap is checked against the payload in **this** request, before the merge, not against the stored map the merge produces. A small edit to an already-large map is therefore accepted, so treat the cap as a per-request budget rather than a diff --git a/api-reference/v2/error-responses.mdx b/api-reference/v2/error-responses.mdx index c189b546..beaa7b66 100644 --- a/api-reference/v2/error-responses.mdx +++ b/api-reference/v2/error-responses.mdx @@ -101,10 +101,6 @@ Many storage- and capacity-related ingestion errors are **transient**: the pipel `E1002` is the one ingestion code that does **not** follow the polling advice above. The file is rejected at upload, so it never enters the pipeline and never gets a status record. Polling [`/context/status`](/api-reference/v2/endpoint/source-status) for it returns `FILE_NOT_FOUND`, not `E1002`. Read `error_code` on each item in the upload response instead. See [Supported file formats](/api-reference/v2/endpoint/ingest-context#supported-file-formats) for what is accepted, and note that one rejected file does not affect the other files in the same request. - -`E6001` is **transient**, not terminal. If you observe it on an in-flight item, keep polling [`/context/status`](/api-reference/v2/endpoint/source-status): the item normally advances to `graph_creation` / `completed` on a subsequent retry with no action on your part. Only contact support if the item is still reported as `errored` after retries are exhausted. - - ## Retry pattern Retry only transient failures: `429`, `500`, and `503`. Use exponential backoff with jitter and keep retries bounded. Two other codes clear on their own, so wait instead of backing off blindly: `409 SOURCE_PROCESSING` (wait for `Retry-After`) and `422 TENANT_INFRA_NOT_READY` (wait until the database is ready). @@ -268,7 +264,7 @@ Common causes: - `document_metadata` length does not match the `documents` array length. - `app_knowledge`, `memories`, or `document_metadata` was sent as an object instead of a JSON-stringified multipart field. - A memory item has neither `text` nor `user_assistant_pairs`. -- A typed `tenant_metadata` value does not match the database metadata schema. +- A typed metadata value does not match the database metadata schema. ### Empty query results diff --git a/api-reference/v2/index.mdx b/api-reference/v2/index.mdx index 5b841ebc..9227743a 100644 --- a/api-reference/v2/index.mdx +++ b/api-reference/v2/index.mdx @@ -1,6 +1,6 @@ --- title: "API Reference" -description: "Single reference to all HydraDB endpoints" +description: "Every v2 endpoint, and the conventions they share." --- ## Quick links @@ -31,7 +31,7 @@ description: "Single reference to all HydraDB endpoints" | [Knowledge](/essentials/v2/knowledge) | Shared source material such as PDFs, docs, app pages, tickets, Slack threads, or webpages. | Use `type=knowledge` when many users or agents should query the same content. | | [Memory](/essentials/v2/memories) | User-specific context such as preferences, conversation history, notes, and inferred traits. | Use `type=memory` when the content should personalize answers for a specific user or collection. | | `database_metadata_schema` | Database-level fields you define up front so metadata can be filtered or queried consistently. | Use it for stable fields like department, customer, region, plan, category, or compliance label. | -| `document_metadata` | JSON-stringified per-document metadata array sent during file ingestion. | Use it to attach source IDs, titles, schema-backed `metadata`, free-form `additional_metadata`, or forceful relations to each uploaded document. | +| `document_metadata` | JSON-stringified per-document metadata array sent during file ingestion. | Use it to attach a source `id`, schema-backed `metadata`, free-form `additional_metadata`, and relations to other sources to each uploaded file, in the same order as `documents`. | | `ids` | IDs returned by ingestion or visible from `/context/list`. | Use them when polling processing status, inspecting content, listing a specific subset, deleting sources, or inspecting relations. | ## End-to-end lifecycle @@ -174,17 +174,15 @@ Errors use the same envelope on every endpoint, including the exceptions above, `meta` may also include a `deprecation` list when a request uses a legacy `/tenants` route or a deprecated field (`tenant_id`/`sub_tenant_id`, or `sub_tenant_ids` on `/query`); each entry carries `deprecated`, a `message`, and `deprecated_since` (field-level notices also add `deprecated_field` and `preferred_field`). It is a non-breaking migration nudge (the status code is unchanged) and is accompanied by a `Deprecation: true` response header. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). -- **Quick reference vs API details.** Each endpoint page starts with a short cheat sheet (what to send, what to save, common gotchas). Later on the page you will see a complete field reference with types, defaults, and examples that is kept in sync with the API. Use the cheat sheet to get moving quickly, and the API details when you need exact request/response shapes (especially for agents and strict validators). +- **Database scoping:** Most database-scoped endpoints require a `database` (formerly `tenant_id`). Many source and query endpoints also accept an optional `collection` (formerly `sub_tenant_id`) for finer-grained scoping. If omitted, the default collection is used. The old `tenant_id`/`sub_tenant_id` names (and the old `/tenants` routes) remain accepted as deprecated aliases; sending a canonical name and its alias with **different** values returns `400`. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). -- **Database scoping.** Most database-scoped endpoints require a `database` (formerly `tenant_id`). Many source and query endpoints also accept an optional `collection` (formerly `sub_tenant_id`) for finer-grained scoping. If omitted, the default collection is used. The old `tenant_id`/`sub_tenant_id` names (and the old `/tenants` routes) remain accepted as deprecated aliases; sending a canonical name and its alias with **different** values returns `400`. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). +- **Async operations:** Database creation, database and collection deletion, and content ingestion are asynchronous. They return immediately after queuing. Use the relevant status endpoint to confirm completion before downstream operations. -- **Async operations.** Database creation, database and collection deletion, and content ingestion are asynchronous. They return immediately after queuing. Use the relevant status endpoint to confirm completion before downstream operations. +- **Pagination:** Listing endpoints (`/context/list`) return pagination fields for browsing large result sets. -- **Pagination.** Listing endpoints (`/context/list`) return pagination fields for browsing large result sets. +- **Parameter casing:** The REST API uses snake_case (`database`, `max_results`). The TypeScript SDK uses camelCase keys and method names (`maxResults`, `deleteCollection`) and ignores request keys it does not recognize, including snake_case ones. The Python SDK uses snake_case throughout. See [SDKs](/api-reference/v2/sdks#naming-conventions). -- **Parameter casing.** The REST API uses snake_case (`database`, `max_results`). The TypeScript SDK uses camelCase keys and method names (`maxResults`, `deleteCollection`) and ignores request keys it does not recognize, including snake_case ones. The Python SDK uses snake_case throughout. See [SDKs](/api-reference/v2/sdks#naming-conventions). - -- **Query modes.** `POST /query` supports `query_by: "hybrid"` or `"text"` and `type: "knowledge"`, `"memory"`, or `"all"`. The same `type` enum is used across ingestion, listing, deletion, and query; query additionally accepts `"all"`. +- **Query modes:** `POST /query` supports `query_by: "hybrid"` or `"text"` and `type: "knowledge"`, `"memory"`, or `"all"`. The same `type` enum is used across ingestion, listing, deletion, and query; query additionally accepts `"all"`. **Status codes:** Successful responses return `200` (or `202` for async accepts). Errors follow standard HTTP semantics: @@ -216,6 +214,6 @@ Rate limits apply per API key. For production deployments, build retry logic wit Existing v1 endpoints remain available under the v1 API Reference. - **Build something:** [Quickstart](/get-started/v2/quickstart) walks through your first integration in five minutes -- **Understand the model:** [Core Concepts](/get-started/v2/core-concepts) explains databases, memories, Query, and metadata +- **Understand the model:** [Core Concepts](/get-started/v2/core-concepts) explains the five primitives - **Go deeper:** [Usage](/essentials/v2/query) covers each primitive in depth - **Install an SDK:** [Python](https://pypi.org/project/hydradb-sdk/) · [TypeScript](https://www.npmjs.com/package/@hydradb/sdk) diff --git a/api-reference/v2/sdks.mdx b/api-reference/v2/sdks.mdx index 5ef5850c..95b83e68 100644 --- a/api-reference/v2/sdks.mdx +++ b/api-reference/v2/sdks.mdx @@ -1,5 +1,5 @@ --- -title: "SDKs: Python and Node" +title: "SDKs: Python and TypeScript" description: "Official Python and TypeScript/Node.js SDKs for the HydraDB API." --- @@ -46,12 +46,6 @@ const client = new HydraDBClient({ **Python:** Both synchronous (`HydraDB`) and asynchronous (`AsyncHydraDB`) clients are available. They share an identical surface; choose based on your application's concurrency model. -## Versioning - -The SDKs automatically include `API-Version: 2` on every outbound request. Server-side routing resolves to the matching routes; the response includes an `X-API-Version: 2` header echoing the resolved version. - -You can verify which version your client is using by inspecting the response headers in any SDK call. - ## Naming conventions The REST API uses **snake_case** for all request and response fields, and the Python SDK preserves those field names for request/response objects. TypeScript camelCases multi-word **method names** and all request and response **fields**. @@ -70,8 +64,6 @@ The Python SDK also uses snake_case throughout (e.g., `client.context.list(datab `database` and `collection` are the current field names (formerly `tenant_id` and `sub_tenant_id`). The old names remain accepted as deprecated aliases for full backward compatibility. -Your IDE's autocomplete and type checking work directly off the API contract: if a field is optional in the API, it's optional in the SDK. - ## SDK method structure SDK methods are grouped under top-level namespaces, one per `/api-reference/v2` group: @@ -85,7 +77,6 @@ SDK methods are grouped under top-level namespaces, one per `/api-reference/v2` | `/connectors/*` and `/connectors` | `client.connectors` | Create, configure, sync, and manage app connectors. `GET /connectors/providers` is `client.list_providers()`. | | `/webhooks/indexing*` | `client.webhooks` | Register, inspect, test, and delete the indexing webhook; list and retry deliveries | - ## Method reference ### Context @@ -164,7 +155,7 @@ A single method covers all retrieval. Pick `type` (`"knowledge"`, `"memory"`, `" | `client.webhooks.get_delivery()` | `GET /webhooks/indexing/deliveries/{delivery_id}` | | `client.webhooks.retry_delivery()` | `POST /webhooks/indexing/deliveries/{delivery_id}/retry` | -The method-reference tables above use the Python (snake_case) method names. TypeScript keeps the same method names but camelCases any that are multi-word (for example, the connector method `list_resources()` in Python is `listResources()` in TypeScript). Request parameters are snake_case in Python and camelCase in TypeScript. +The tables above use the Python method names. TypeScript camelCases multi-word ones (`list_resources()` is `listResources()`). ## Migrating from v1 @@ -196,7 +187,7 @@ The v1 SDK methods remain available on the `<2.0.0` releases of `hydradb-sdk` / ### Create a database -A database is an isolated workspace. Most organizations create one database total, with collections for users or teams. See [Multi-Tenant](/essentials/v2/multi-tenant) for the full model. +A database is an isolated workspace. B2C apps usually use one database with a collection per user; B2B apps use one database per customer. See [Multi-tenancy](/essentials/v2/multi-tenant) for the full model. ```python Python SDK @@ -235,13 +226,7 @@ import time while True: status = client.databases.status(database="my_first_database") - infra = status.data.infra - if ( - infra.scheduler_status - and infra.graph_status - and infra.vectorstore_status.knowledge - and infra.vectorstore_status.memories - ): + if status.data.infra.ready_for_ingestion: break time.sleep(2) ``` @@ -251,13 +236,7 @@ while (true) { database: "my_first_database", }); // Response fields are optional in the TypeScript types. - const infra = status.data?.infra; - if ( - infra?.schedulerStatus && - infra.graphStatus && - infra.vectorstoreStatus?.knowledge && - infra.vectorstoreStatus?.memories - ) break; + if (status.data?.infra?.readyForIngestion) break; await new Promise((r) => setTimeout(r, 2000)); } ``` @@ -401,7 +380,7 @@ knowledge = client.query( graph_context=True, ) -# Personalized — searches knowledge AND user memories together +# Personalized: knowledge and this user's memories, from the same collection personalized = client.query( database="my_first_database", collection="user_alex", @@ -431,7 +410,7 @@ const knowledge = await client.query({ graphContext: true, }); -// Personalized — searches knowledge AND user memories together +// Personalized: knowledge and this user's memories, from the same collection const personalized = await client.query({ database: "my_first_database", collection: "user_alex", @@ -537,25 +516,6 @@ All responses, except connector endpoints and `PATCH /databases/{database}/metad The SDKs return the full envelope. Read the payload from the `data` field (e.g. `response.data`); on failure the SDK raises a typed exception instead of returning an envelope with `error` populated. -## Type safety - - -Both SDKs are fully typed: - -- **Autocomplete** for all method names and parameters -- **Type checking** for request and response objects -- **Inline documentation** for each parameter, sourced from the OpenAPI spec -- **Compile-time validation** for required vs optional fields -- **Enum-typed values** for `type`, `query_by`, `operator` and `mode` - - -The SDKs provide exact type parity with the API specification: - -- **Request parameters:** every field documented in the API reference is reflected in method signatures -- **Response objects:** return types match the JSON schema for each endpoint -- **Error types:** exception structures mirror error response formats -- **Nested objects:** complex parameters and responses keep their full structure - ## Error handling Both SDKs throw exceptions for non-2xx responses. Error objects expose the envelope's `error` payload: @@ -576,7 +536,7 @@ except NotFoundError: pass except ApiError as exc: if exc.status_code == 429: - # Rate limit — retry with backoff + # Rate limit: retry with backoff pass else: raise @@ -599,7 +559,7 @@ try { if (code === "DATABASE_NOT_FOUND") { // Handle missing database } else if (error.statusCode === 429) { - // Rate limit — retry with backoff + // Rate limit: retry with backoff } else { throw error; } @@ -618,16 +578,6 @@ Not every status has its own class: `401` and `429` do not. Catch the base `ApiE For the full list of error codes and retry patterns, see [Error Responses](/api-reference/v2/error-responses). -## IDE-driven discovery - -Whether you're using TypeScript, Python, VS Code, PyCharm, or any modern IDE, the workflow is the same: - -1. Type the method name to see all available methods -2. Open the parentheses to see all required and optional parameters -3. Press `Cmd+Space` (macOS) or `Ctrl+Space` (Windows/Linux) to get inline documentation - -This works because the SDKs are fully typed with comprehensive parameter docs sourced from the OpenAPI spec. - ## Related sections - [API Reference](/api-reference/v2): complete endpoint documentation diff --git a/essentials/v2/api-results.mdx b/essentials/v2/api-results.mdx index 6ebce124..33e4a5e4 100644 --- a/essentials/v2/api-results.mdx +++ b/essentials/v2/api-results.mdx @@ -466,11 +466,11 @@ If the memory query fails or times out, fall back to the knowledge-only prompt. ## 5. Practical guidance -- **Preserve server order.** Don't re-sort chunks client-side. -- **Start small on chunks.** `max_results: 10` is a reasonable default. Drop to 5 if you hit token limits, raise to 20 if you rerank downstream. -- **Use graph context selectively.** It improves relational queries and bloats simple lookups. See [Context Graphs](/essentials/v2/context-graphs). +- **Preserve server order:** Don't re-sort chunks client-side. +- **Start small on chunks:** `max_results: 10` is a reasonable default. Drop to 5 if you hit token limits, raise to 20 if you rerank downstream. +- **Use graph context selectively:** It improves relational queries and bloats simple lookups. See [Context Graphs](/essentials/v2/context-graphs). - **Give the model a grounding instruction:** Use a system prompt like "answer only from the provided context" to reduce unsupported answers when retrieval is thin. -- **Format consistently.** Whatever section delimiters you choose (`=== CONTEXT ===`, `Chunk N`, `Source:`), keep them stable across calls so the model learns the structure. +- **Format consistently:** Whatever section delimiters you choose (`=== CONTEXT ===`, `Chunk N`, `Source:`), keep them stable across calls so the model learns the structure. --- From 98e41701855844b9214a62391ea88692727a1936 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Tue, 6 Oct 2026 15:27:24 +0530 Subject: [PATCH 03/11] docs(v2): align API examples and specification with behavior Signed-off-by: SohamRatnaparkhi --- api-reference/v2/endpoint/ingest-context.mdx | 22 +++-- api-reference/v2/endpoint/query.mdx | 2 +- api-reference/v2/endpoint/submit-feedback.mdx | 4 +- api-reference/v2/openapi.json | 93 ++++++++++++------- 4 files changed, 77 insertions(+), 44 deletions(-) diff --git a/api-reference/v2/endpoint/ingest-context.mdx b/api-reference/v2/endpoint/ingest-context.mdx index 4a3b5faa..dfb7c27d 100644 --- a/api-reference/v2/endpoint/ingest-context.mdx +++ b/api-reference/v2/endpoint/ingest-context.mdx @@ -45,8 +45,10 @@ with open("/path/to/policy.pdf", "rb") as policy: "database": "acme_corp", "collection": "team_docs", "title": "Pricing discussion", - "type": "slack", - "content": {"text": "We agreed on three tiers..."}, + "kind": "message", + "provider": "slack", + "external_id": "1716213600.000100", + "fields": {"kind": "message", "body": "We agreed on three tiers..."}, "metadata": {"department": "product"}, "additional_metadata": {"channel": "pricing"}, } @@ -111,8 +113,10 @@ const knowledgeResult = await client.context.ingest({ database: "acme_corp", collection: "team_docs", title: "Pricing discussion", - type: "slack", - content: { text: "We agreed on three tiers..." }, + kind: "message", + provider: "slack", + external_id: "1716213600.000100", + fields: { kind: "message", body: "We agreed on three tiers..." }, metadata: { department: "product" }, additional_metadata: { channel: "pricing" }, }, @@ -172,8 +176,10 @@ curl -X POST 'https://api.hydradb.com/context/ingest' \ "database": "acme_corp", "collection": "team_docs", "title": "Pricing discussion", - "type": "slack", - "content": { "text": "We agreed on three tiers..." }, + "kind": "message", + "provider": "slack", + "external_id": "1716213600.000100", + "fields": { "kind": "message", "body": "We agreed on three tiers..." }, "metadata": { "department": "product" }, "additional_metadata": { "channel": "pricing" } } @@ -504,7 +510,6 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in { "id": "slack_thread_001", "title": "Pricing discussion", - "type": "slack", "kind": "message", "provider": "slack", "external_id": "1716213600.000100", @@ -524,7 +529,6 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in { id: "slack_thread_001", title: "Pricing discussion", - type: "slack", kind: "message", provider: "slack", external_id: "1716213600.000100", @@ -546,7 +550,6 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in { "id": "slack_thread_001", "title": "Pricing discussion", - "type": "slack", "kind": "message", "provider": "slack", "external_id": "1716213600.000100", @@ -564,7 +567,6 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in | | Optional. If sent, must match the form-level `database`. Formerly `tenant_id` (deprecated alias). | | | Optional. If sent, must match the form-level `collection`. Formerly `sub_tenant_id` (deprecated alias). | | | Short title or subject shown in search results. | - | | Source category (`slack`, `notion`, `gmail`, `webpage`, etc.). Used for filtering and display. | | | Optional long-form description. | | | Canonical URL or reference link. | | | ISO-8601 timestamp (creation or last-updated). | diff --git a/api-reference/v2/endpoint/query.mdx b/api-reference/v2/endpoint/query.mdx index 87245b4c..985662e7 100644 --- a/api-reference/v2/endpoint/query.mdx +++ b/api-reference/v2/endpoint/query.mdx @@ -508,7 +508,7 @@ result = client.query( | `contains` | a single value | sources whose field **holds** that value. Multi-value fields are stored comma-joined, so this matches one member of that list | | `contains_any` | an array | sources holding **any one** of the listed values (OR/IN) | - ```json + ```text "metadata_filters": { "department": { "equals": "legal" }, "attendee_emails": { "contains": "b@company.com" }, diff --git a/api-reference/v2/endpoint/submit-feedback.mdx b/api-reference/v2/endpoint/submit-feedback.mdx index fc62da39..3488b7f0 100644 --- a/api-reference/v2/endpoint/submit-feedback.mdx +++ b/api-reference/v2/endpoint/submit-feedback.mdx @@ -14,7 +14,7 @@ Both people and agents can submit. An agent that can tell a retrieval was unhelp Every HydraDB response carries a `request_id` in `meta`, and the same value in the `X-Request-ID` header. Send that id back and we can line your comment up with the exact query it is about: the text queried, what came back, how long it took. -```json Query response {6} +```text Query response {6} { "success": true, "data": { "chunks": [ /* ... */ ] }, @@ -190,7 +190,7 @@ curl -X POST 'https://api.hydradb.com/feedback' \ If you already know the right answer (you are running an evaluation set, or you know which document the user needed), send it. It is a much stronger signal than a comment, because we can score it without a human reading it. -```json +```text "ground_truth": { "answer": "Refunds are processed within 14 days.", "source_ids": ["policy_2024", "handbook_q3"] diff --git a/api-reference/v2/openapi.json b/api-reference/v2/openapi.json index 0bd72b6e..db4be8c3 100644 --- a/api-reference/v2/openapi.json +++ b/api-reference/v2/openapi.json @@ -412,7 +412,7 @@ "fetch.V2SourceFetchResponse": { "properties": { "content": { - "description": "Extracted text content of the source document.", + "description": "Stored original bytes as text when valid UTF-8. Binary originals, such as PDF and DOCX, are returned in content_base64 instead. Null in url mode.", "example": "# Q4 Report\n\nRevenue grew 23% quarter over quarter.", "type": "string" }, @@ -436,7 +436,7 @@ "type": "string" }, "inferred_content": { - "description": "LLM-generated summary of the source content.", + "description": "Model-derived content when available in content or both mode. Null in url mode.", "example": "Summary: Q4 revenue rose 23% QoQ, driven by enterprise expansion.", "type": "string" }, @@ -4494,9 +4494,16 @@ "type": "string" }, "indexing_status": { - "description": "Current processing state: `queued`, `processing`, `completed`, or `failed`.", + "description": "Current indexing state. graph_creation is searchable; completed has finished; errored is terminal.", "example": "completed", - "type": "string" + "type": "string", + "enum": [ + "queued", + "processing", + "graph_creation", + "completed", + "errored" + ] }, "message": { "description": "Human-readable status description.", @@ -4505,7 +4512,7 @@ }, "success": { "deprecated": true, - "description": "Deprecated for API clients: this reads like a per-source outcome but is\na constant echo of the envelope's `success` — it is true even for a\nsource that failed indexing. For the state of THIS source read\nindexing_status (and error_code/error_message when it is errored); for\nwhether the request itself succeeded read the HTTP status code or the\nenvelope's top-level `success`. Still emitted unchanged for existing\nclients (PRO-1208).", + "description": "Deprecated per-item success flag. Use indexing_status and error_code to determine whether indexing completed.", "example": true, "type": "boolean", "x-deprecated": "true" @@ -5170,7 +5177,8 @@ "OperatorOr", "OperatorAnd", "OperatorPhrase" - ] + ], + "default": "or" }, "search.PathTriplet": { "properties": { @@ -5215,7 +5223,8 @@ "x-enum-varnames": [ "QueryByHybrid", "QueryByText" - ] + ], + "default": "hybrid" }, "search.QueryRequest": { "properties": { @@ -5233,7 +5242,21 @@ "type": "string" }, "alpha": { - "description": "Weighting balance between dense and sparse retrieval in hybrid mode. `\"auto\"` lets HydraDB choose; a number from 0 (full BM25) to 1 (full dense) sets it explicitly." + "description": "Semantic weight in hybrid search: 0 is BM25 only and 1 is semantic only. The default is 0.8; \"auto\" also resolves to 0.8.", + "default": 0.8, + "oneOf": [ + { + "type": "number", + "minimum": 0, + "maximum": 1 + }, + { + "type": "string", + "enum": [ + "auto" + ] + } + ] }, "collection": { "description": "Collection scope. Defaults to the default collection when omitted. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated).", @@ -5241,7 +5264,7 @@ "type": "string" }, "collections": { - "description": "Preferred /query scope selector. Send either a list of collection IDs for equal normalized weighting, or an object mapping collection ID to a positive relative ranking weight with at most one decimal place. Do not send together with the deprecated sub_tenant_ids or sub_tenant_id.", + "description": "A list of existing collections (equal weights), or an object mapping collection IDs to positive ranking weights with at most one decimal place. Maximum 100 collections. A missing collection returns 400. Use instead of a single collection selector.", "example": [ "team_docs", "engineering" @@ -5277,14 +5300,15 @@ "x-preferred": true }, "database": { - "description": "Database is the canonical v2 name for the tenant scope. TenantID is its\ndeprecated alias and remains fully accepted. The TenantAliases middleware\nreconciles the two before binding, so TenantID is always populated and the\nhandler reads it; Database/Collection are carried only for docs/OpenAPI.", + "description": "Database to query. The deprecated tenant_id alias is still accepted.", "example": "acme_corp", "type": "string" }, "graph_context": { - "description": "Whether to include graph context in the response. Defaults to true for /query when omitted.", + "description": "Include the graph slice. False takes effect only in fast mode; thinking always includes graph context.", "example": true, - "type": "boolean" + "type": "boolean", + "default": true }, "graph_vector_prune": { "description": "GraphVectorPrune switches the graph-connected-chunks lane from \"fetch\ngraph-selected chunks and let the fusion reranker sort them out\" to \"fetch\na wider graph-selected candidate pool, then rank that pool by Milvus vector\nsimilarity, fully replacing the final chunk list.\" Works in either fast or\nthinking mode. Default false preserves existing behavior. Also gated\nserver-side by a repo-level config flag (SearchService's\ngraphVectorPruneEnabled) — if that flag is off, this is forced to false\nregardless of what the request sets, so a deployment can disable the\nmechanism without any client-side change.", @@ -5311,7 +5335,10 @@ "max_results": { "description": "Maximum number of chunks to return.", "example": 10, - "type": "integer" + "type": "integer", + "default": 10, + "minimum": 1, + "maximum": 250 }, "metadata_filters": { "$ref": "#/components/schemas/search.MetadataFilters" @@ -5327,7 +5354,8 @@ }, "operator": { "$ref": "#/components/schemas/search.Operator", - "example": "and" + "example": "and", + "description": "BM25 term operator. \"and\" and \"phrase\" require query_by: \"text\"; otherwise the request returns 400." }, "query": { "description": "Natural-language search query.", @@ -5335,9 +5363,10 @@ "type": "string" }, "query_apps": { - "description": "Whether to include app-aware knowledge retrieval. Applies to knowledge hybrid queries.", + "description": "Adds app-aware retrieval to knowledge hybrid queries without excluding other knowledge. Send false to disable it.", "example": true, - "type": "boolean" + "type": "boolean", + "default": true }, "query_by": { "$ref": "#/components/schemas/search.QueryBy", @@ -5345,14 +5374,18 @@ "example": "hybrid" }, "query_forceful_relations": { - "description": "Whether to force relation expansion for graph-aware query retrieval. Defaults to true when omitted.", + "description": "Include declared related knowledge sources in additional_context when mode resolves to thinking. Ignored in fast mode.", "example": true, - "type": "boolean" + "type": "boolean", + "default": true }, "recency_bias": { "description": "Recency boost applied to ranking. 0 disables it; higher values favour more recent sources.", "example": 0.2, - "type": "number" + "type": "number", + "default": 0.4, + "minimum": 0, + "maximum": 1 }, "sub_tenant_id": { "deprecated": true, @@ -5424,7 +5457,7 @@ }, "type": { "$ref": "#/components/schemas/search.SourceType", - "description": "Corpus to query: knowledge, memory, or all." + "description": "Store to query: knowledge (default), memory, or all. Both stores use the same collection scope." } }, "type": "object" @@ -5440,7 +5473,9 @@ "RecallModeFast", "RecallModeThinking", "RecallModeAuto" - ] + ], + "default": "auto", + "description": "fast: one retrieval pass; thinking: expanded retrieval and reranking; auto (default): select fast or thinking per query." }, "search.ScoredPathResponse": { "properties": { @@ -5684,7 +5719,8 @@ "SourceKnowledge", "SourceMemory", "SourceAll" - ] + ], + "default": "knowledge" }, "search.TemporalDuration": { "description": "TemporalDuration is the computed event-duration answer, when resolved.", @@ -6395,11 +6431,6 @@ "example": true, "type": "boolean" }, - "enable_match": { - "description": "Whether to enable exact-match filtering on this field.", - "example": true, - "type": "boolean" - }, "enable_sparse_embedding": { "description": "Whether to enable BM25 (sparse) embedding search on this field.", "example": false, @@ -8679,7 +8710,7 @@ "schema": { "properties": { "app_knowledge": { - "description": "App-knowledge items as a JSON array (type=knowledge). Per item, `metadata` is capped at 16 KiB and `additional_metadata` at 1 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes (keys and punctuation count). The deprecated `tenant_metadata` / `document_metadata` spellings are accepted here and held to the same caps. Over-cap returns 400 with the actual byte count. Each item may also carry `acl`, a list of principals (`user_email:`, a bare email, `group::`, `domain:`, or `__public__`) restricting who may retrieve it; omit it to leave the document unrestricted, and send an empty list to restrict it to nobody. A malformed principal rejects the whole request with 400.", + "description": "JSON-stringified array of app sources (type=knowledge). Typed items use kind, provider, external_id, and fields. Metadata limits: 16 KiB metadata and 1 KiB additional_metadata. Optional acl restricts retrieval when the query supplies principals; an empty item ACL restricts it to nobody.", "title": "app_knowledge", "type": "string" }, @@ -8692,7 +8723,7 @@ "type": "string" }, "document_metadata": { - "description": "Per-document metadata as a JSON array (type=knowledge). Per item, `metadata` is capped at 16 KiB and `additional_metadata` at 1 KiB. Both caps are measured on the compact JSON encoding of the whole map in UTF-8 bytes, so keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", + "description": "JSON-stringified array, one item per uploaded file in the same order. Items accept id, metadata, additional_metadata, relations, infer, evidence_kind, and evidence_subject. Metadata limits: 16 KiB schema fields and 1 KiB free-form fields, measured as compact JSON in UTF-8 bytes.", "title": "document_metadata", "type": "string" }, @@ -8706,7 +8737,7 @@ "type": "string" }, "memories": { - "description": "Memory items as a JSON array (type=memory). Per item, `metadata` is capped at 16 KiB and `additional_metadata` at 1 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes (keys and punctuation count). Over-cap returns 400 with the actual byte count.", + "description": "JSON-stringified array of memory items (type=memory). Each item has text or user_assistant_pairs, and optional infer (default false), id, user_name, metadata, and additional_metadata. Nested metadata values are objects, not JSON strings. Limits: 16 KiB metadata and 1 KiB additional_metadata.", "title": "memories", "type": "string" }, @@ -8786,7 +8817,7 @@ } } }, - "description": "Body is not multipart/form-data (e.g. a JSON body)" + "description": "Unsupported Content-Type: the body is neither multipart/form-data nor JSON." }, "422": { "content": { From e134a1958db16346c00a58283774a5effb04ec7f Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Tue, 6 Oct 2026 15:35:11 +0530 Subject: [PATCH 04/11] docs(v2): correct generated response and metadata schema examples Signed-off-by: SohamRatnaparkhi --- api-reference/v2/endpoint/subgraph.mdx | 2 +- api-reference/v2/openapi.json | 638 ++++++++++++++----------- 2 files changed, 373 insertions(+), 267 deletions(-) diff --git a/api-reference/v2/endpoint/subgraph.mdx b/api-reference/v2/endpoint/subgraph.mdx index d5f4f67c..a93d159c 100644 --- a/api-reference/v2/endpoint/subgraph.mdx +++ b/api-reference/v2/endpoint/subgraph.mdx @@ -54,7 +54,7 @@ hydradb --output json subgraph slack_C0BE77_1788320073 | jq '.sources[].source_i | Name | Description | | --- | --- | -| | The item to start from. Any `id` returned by Query, List Documents or Ingest. URL-encode it if it contains reserved characters. An id containing a literal `/` cannot be written as one path segment; pass those as `GET /context/subgraph?id=...` instead (the form the SDKs use for every id). | +| | The item to start from. Any `id` returned by Query, List Documents or Ingest. URL-encode it if it contains reserved characters. An id containing a literal `/` cannot be written as one path segment; pass those as `GET /context/subgraph?id=...` instead (Python and TypeScript SDK 2.1.7 use this form for every id). | ## Query parameters diff --git a/api-reference/v2/openapi.json b/api-reference/v2/openapi.json index db4be8c3..3b896fc3 100644 --- a/api-reference/v2/openapi.json +++ b/api-reference/v2/openapi.json @@ -174,7 +174,7 @@ "feedback.GroundTruth": { "properties": { "answer": { - "description": "Answer is the response the caller expected — the text a correct system\nwould have produced from the retrieved context.", + "description": "Answer is the response the caller expected \u2014 the text a correct system\nwould have produced from the retrieved context.", "maxLength": 8000, "type": "string" }, @@ -318,7 +318,7 @@ }, "ground_truth": { "$ref": "#/components/schemas/feedback.GroundTruth", - "description": "What you already know the right answer to be, when you know it. Supply an expected `answer`, the `source_ids` that contain it, or both — at least one is required if the field is present. Machine-checkable, so it is a stronger signal than a comment: submit it alone and `feedback` becomes optional.", + "description": "What you already know the right answer to be, when you know it. Supply an expected `answer`, the `source_ids` that contain it, or both \u2014 at least one is required if the field is present. Machine-checkable, so it is a stronger signal than a comment: submit it alone and `feedback` becomes optional.", "example": { "source_ids": [ "HydraDoc1234", @@ -347,7 +347,7 @@ "description": "Optional overall judgement: `positive`, `negative`, or `neutral`. Omit to send a comment with no rating." }, "request_id": { - "description": "The `request_id` from `response.meta` of the query this feedback is about. Required — it is what links the feedback to the query that ran.", + "description": "The `request_id` from `response.meta` of the query this feedback is about. Required \u2014 it is what links the feedback to the query that ran.", "example": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "format": "uuid", "type": "string" @@ -457,7 +457,7 @@ }, "success": { "deprecated": true, - "description": "Deprecated for API clients: to decide whether the request succeeded,\ncheck the HTTP status code — 2xx is success — or equivalently the\nenvelope's top-level `success`. This nested copy always carries the same\nvalue as that flag and never carries independent information. Still\nemitted unchanged for existing clients (PRO-1208).", + "description": "Deprecated for API clients: to decide whether the request succeeded,\ncheck the HTTP status code \u2014 2xx is success \u2014 or equivalently the\nenvelope's top-level `success`. This nested copy always carries the same\nvalue as that flag and never carries independent information. Still\nemitted unchanged for existing clients (PRO-1208).", "example": true, "type": "boolean", "x-deprecated": "true" @@ -611,11 +611,11 @@ "type": "string" }, "hydration": { - "description": "Hydration is set on Source nodes in AuxiliaryRelations only, and omitted\neverywhere else. RELATES_TO MERGEs its target by source_id, so a target\nthat has not been ingested yet still exists as a node — callers must be\nable to tell a real document from a forward reference to one.\n\n\tresolved — ingested; source_id and app_provider both present\n\tstub — MERGE-created target; source_id present, no app_provider\n\tplaceholder — source_id IS NULL, keyed by app_external_id, awaiting\n\t builder.py's reconciliation pass", + "description": "Hydration is set on Source nodes in AuxiliaryRelations only, and omitted\neverywhere else. RELATES_TO MERGEs its target by source_id, so a target\nthat has not been ingested yet still exists as a node \u2014 callers must be\nable to tell a real document from a forward reference to one.\n\n\tresolved \u2014 ingested; source_id and app_provider both present\n\tstub \u2014 MERGE-created target; source_id present, no app_provider\n\tplaceholder \u2014 source_id IS NULL, keyed by app_external_id, awaiting\n\t builder.py's reconciliation pass", "type": "string" }, "identifier": { - "description": "NO omitempty — serialize as null", + "description": "NO omitempty \u2014 serialize as null", "example": "Acme Corp", "type": "string" }, @@ -645,7 +645,7 @@ "graph.GraphRelationsResponse": { "properties": { "auxiliary_relations": { - "description": "AuxiliaryRelations carries the structural graph around the entity\nrelations: Entity->Source presence, Source->Comment/Attachment,\nActor->Source/Comment, and Source->Source links. Same item shape as\nRelations, so a caller wanting one graph concatenates the two.\n\nDeliberately a SEPARATE array rather than merged into Relations:\ncapPreservingTies counts triplets against the caller's limit, and\ncomputeNextCursor keys on relation timestamps. Auxiliary edges carry\ncreated_at — a different clock — so merging them would both shrink the\nentity relations returned for a given limit and corrupt the cursor.", + "description": "AuxiliaryRelations carries the structural graph around the entity\nrelations: Entity->Source presence, Source->Comment/Attachment,\nActor->Source/Comment, and Source->Source links. Same item shape as\nRelations, so a caller wanting one graph concatenates the two.\n\nDeliberately a SEPARATE array rather than merged into Relations:\ncapPreservingTies counts triplets against the caller's limit, and\ncomputeNextCursor keys on relation timestamps. Auxiliary edges carry\ncreated_at \u2014 a different clock \u2014 so merging them would both shrink the\nentity relations returned for a given limit and corrupt the cursor.", "example": [ { "chunk_id": "HydraEmbeddings123_0", @@ -756,7 +756,7 @@ }, "success": { "deprecated": true, - "description": "Deprecated for API clients: to decide whether the request succeeded,\ncheck the HTTP status code — 2xx is success — or equivalently the\nenvelope's top-level `success`. This nested copy always carries the same\nvalue and never carries independent information. Still emitted unchanged\nfor existing clients (PRO-1208).", + "description": "Deprecated for API clients: to decide whether the request succeeded,\ncheck the HTTP status code \u2014 2xx is success \u2014 or equivalently the\nenvelope's top-level `success`. This nested copy always carries the same\nvalue and never carries independent information. Still emitted unchanged\nfor existing clients (PRO-1208).", "example": true, "type": "boolean", "x-deprecated": "true" @@ -991,7 +991,7 @@ "type": "string" }, "discovered_via": { - "description": "DiscoveredVia and DiscoveredRelation record the traversal provenance:\nwhich already-admitted source this member was first reached from, and\nthrough which mechanism — a RELATES_TO relation_type (reply_to,\nchild_of, ...), same_thread, parent or child. Empty on the seed.", + "description": "DiscoveredVia and DiscoveredRelation record the traversal provenance:\nwhich already-admitted source this member was first reached from, and\nthrough which mechanism \u2014 a RELATES_TO relation_type (reply_to,\nchild_of, ...), same_thread, parent or child. Empty on the seed.", "type": "string" }, "hydration": { @@ -1085,12 +1085,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1119,12 +1123,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1134,8 +1142,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1152,7 +1160,7 @@ "$ref": "#/components/schemas/fetch.V2SourceFetchResponse", "example": { "content": "# Q4 Report\n\nRevenue grew 23% quarter over quarter.", - "content_type": "application/pdf", + "content_type": "text/markdown", "error": "", "id": "HydraDoc1234", "inferred_content": "Summary: Q4 revenue rose 23% QoQ, driven by enterprise expansion.", @@ -1163,12 +1171,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1178,8 +1190,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1219,12 +1231,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1234,8 +1250,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1333,12 +1349,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1348,8 +1368,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1457,12 +1477,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1472,8 +1496,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1494,12 +1518,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1509,8 +1537,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1539,12 +1567,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1554,8 +1586,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1585,16 +1617,20 @@ } ], "success": true, - "success_count": 2 + "success_count": 1 } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1604,8 +1640,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1687,12 +1723,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1702,8 +1742,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1751,12 +1791,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1766,8 +1810,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -1961,12 +2005,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -1976,8 +2024,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2007,12 +2055,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2022,8 +2074,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2055,12 +2107,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2070,8 +2126,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2099,12 +2155,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2114,8 +2174,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2138,12 +2198,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2153,8 +2217,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2177,12 +2241,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2192,8 +2260,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2235,12 +2303,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2250,8 +2322,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2272,7 +2344,6 @@ { "data_type": "VARCHAR", "enable_dense_embedding": true, - "enable_match": true, "enable_sparse_embedding": false, "max_length": 256, "name": "category" @@ -2282,12 +2353,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2297,8 +2372,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2326,12 +2401,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2341,8 +2420,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2371,12 +2450,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2386,8 +2469,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2422,12 +2505,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2437,8 +2524,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2460,12 +2547,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2475,8 +2566,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2498,12 +2589,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2513,8 +2608,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2535,12 +2630,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2550,8 +2649,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2576,12 +2675,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2591,8 +2694,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2619,12 +2722,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2634,8 +2741,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2657,12 +2764,16 @@ } }, "error": { - "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", - "example": { - "code": "DATABASE_NOT_FOUND", - "message": "Database not found" - } + "anyOf": [ + { + "$ref": "#/components/schemas/handler.apiError" + }, + { + "type": "null" + } + ], + "description": "Null on success; an object with code and message on failure.", + "example": null }, "meta": { "$ref": "#/components/schemas/handler.responseMeta", @@ -2672,8 +2783,8 @@ "latency_ms": 12.3, "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d", "source_type": "file", - "sub_tenant_id": "sub_tenant_4567", - "tenant_id": "tenant_1234" + "sub_tenant_id": "team_docs", + "tenant_id": "acme_corp" } }, "success": { @@ -2751,7 +2862,7 @@ }, "error": { "$ref": "#/components/schemas/handler.apiError", - "description": "Error message, empty string on success.", + "description": "Error code and message for the failed request.", "example": { "code": "DATABASE_NOT_FOUND", "message": "Database not found" @@ -3807,7 +3918,7 @@ }, "additional_metadata": { "additionalProperties": {}, - "description": "Free-form key-value pairs to merge into the source's `additional_metadata`. The only accepted spelling for document metadata on this endpoint. Capped at 1 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes — keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", + "description": "Free-form key-value pairs to merge into the source's `additional_metadata`. The only accepted spelling for document metadata on this endpoint. Capped at 1 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes \u2014 keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", "example": { "author": "ada", "doc_version": 3 @@ -3826,7 +3937,7 @@ }, "database_metadata": { "additionalProperties": {}, - "description": "Schema-backed metadata fields to merge into the source's `metadata` (database metadata). Canonical name; `tenant_metadata` is a deprecated alias. Capped at 16 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes — keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", + "description": "Schema-backed metadata fields to merge into the source's `metadata` (database metadata). Canonical name; `tenant_metadata` is a deprecated alias. Capped at 16 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes \u2014 keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", "example": { "department": "legal", "priority": 7 @@ -3857,7 +3968,7 @@ "tenant_metadata": { "additionalProperties": {}, "deprecated": true, - "description": "Deprecated alias for `database_metadata`, still accepted here; `database_metadata` wins when both are sent. Capped at 16 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes — keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", + "description": "Deprecated alias for `database_metadata`, still accepted here; `database_metadata` wins when both are sent. Capped at 16 KiB, measured on the compact JSON encoding of the whole map in UTF-8 bytes \u2014 keys, quotes, commas and braces count toward the budget. Over-cap returns 400 with the actual byte count.", "example": { "department": "legal", "priority": 7 @@ -4258,7 +4369,7 @@ "handler.responseMeta": { "properties": { "api_version": { - "description": "APIVersion echoes the version of the API that served the request (PRO-1209),\nsourced from reqmeta.APIVersion — the same value carried by OpenAPI\ninfo.version and /health — so a client always knows which API version\nproduced a response. Always present (no omitempty).", + "description": "APIVersion echoes the version of the API that served the request (PRO-1209),\nsourced from reqmeta.APIVersion \u2014 the same value carried by OpenAPI\ninfo.version and /health \u2014 so a client always knows which API version\nproduced a response. Always present (no omitempty).", "type": "string" }, "collection": { @@ -4272,7 +4383,7 @@ "type": "string" }, "deprecation": { - "description": "Deprecation lists any migration nudges that apply to this request — the\ncaller used a legacy /tenants route, a legacy tenant_id/sub_tenant_id field,\nor the deprecated sub_tenant_ids selector. It is a non-breaking signal (the\nstatus code is unchanged); omitempty keeps it absent for fully-migrated\nrequests. A list so independent deprecations coexist without clobbering.", + "description": "Deprecation lists any migration nudges that apply to this request \u2014 the\ncaller used a legacy /tenants route, a legacy tenant_id/sub_tenant_id field,\nor the deprecated sub_tenant_ids selector. It is a non-breaking signal (the\nstatus code is unchanged); omitempty keeps it absent for fully-migrated\nrequests. A list so independent deprecations coexist without clobbering.", "items": { "$ref": "#/components/schemas/handler.deprecationNotice" }, @@ -4345,7 +4456,7 @@ "type": "object" }, "ingestion.SourceStatus": { - "description": "Status is the item's initial lifecycle state. Both modes share this\nvocabulary — memory mode reuses the same values.", + "description": "Status is the item's initial lifecycle state. Both modes share this\nvocabulary \u2014 memory mode reuses the same values.", "enum": [ "queued", "processing", @@ -4416,7 +4527,7 @@ }, "success": { "deprecated": true, - "description": "Deprecated for API clients: whether the REQUEST was accepted is the HTTP\nstatus code (202) or equivalently the envelope's top-level `success`.\nWhether each SOURCE ingested is per-item — read results[].status and\nresults[].error, then poll GET /context/status, since a 202 only means\nqueued. This flag answers neither question independently: it always\nmirrors the envelope. Still emitted unchanged for existing clients\n(PRO-1208).", + "description": "Deprecated for API clients: whether the REQUEST was accepted is the HTTP\nstatus code (202) or equivalently the envelope's top-level `success`.\nWhether each SOURCE ingested is per-item \u2014 read results[].status and\nresults[].error, then poll GET /context/status, since a 202 only means\nqueued. This flag answers neither question independently: it always\nmirrors the envelope. Still emitted unchanged for existing clients\n(PRO-1208).", "example": true, "type": "boolean", "x-deprecated": "true" @@ -4595,7 +4706,7 @@ "uniqueItems": false }, "include_fields": { - "description": "Field projection — only the listed fields plus id, database, collection are returned. Only applies to type=knowledge.", + "description": "Field projection \u2014 only the listed fields plus id, database, collection are returned. Only applies to type=knowledge.", "example": [ "id", "title", @@ -4697,7 +4808,7 @@ }, "success": { "deprecated": true, - "description": "Deprecated for API clients: to decide whether the request succeeded,\ncheck the HTTP status code — 2xx is success — or equivalently the\nenvelope's top-level `success`. This nested copy always carries the same\nvalue and never carries independent information. Still emitted unchanged\nfor existing clients (PRO-1208).", + "description": "Deprecated for API clients: to decide whether the request succeeded,\ncheck the HTTP status code \u2014 2xx is success \u2014 or equivalently the\nenvelope's top-level `success`. This nested copy always carries the same\nvalue and never carries independent information. Still emitted unchanged\nfor existing clients (PRO-1208).", "example": true, "type": "boolean", "x-deprecated": "true" @@ -4807,7 +4918,7 @@ "type": "string" }, "memory_id": { - "description": "MemoryID is the memory identifier — the type=memory spelling of ID, and\npresent on exactly the same terms.", + "description": "MemoryID is the memory identifier \u2014 the type=memory spelling of ID, and\npresent on exactly the same terms.", "example": "memory_1234", "type": "string" }, @@ -5311,7 +5422,7 @@ "default": true }, "graph_vector_prune": { - "description": "GraphVectorPrune switches the graph-connected-chunks lane from \"fetch\ngraph-selected chunks and let the fusion reranker sort them out\" to \"fetch\na wider graph-selected candidate pool, then rank that pool by Milvus vector\nsimilarity, fully replacing the final chunk list.\" Works in either fast or\nthinking mode. Default false preserves existing behavior. Also gated\nserver-side by a repo-level config flag (SearchService's\ngraphVectorPruneEnabled) — if that flag is off, this is forced to false\nregardless of what the request sets, so a deployment can disable the\nmechanism without any client-side change.", + "description": "GraphVectorPrune switches the graph-connected-chunks lane from \"fetch\ngraph-selected chunks and let the fusion reranker sort them out\" to \"fetch\na wider graph-selected candidate pool, then rank that pool by Milvus vector\nsimilarity, fully replacing the final chunk list.\" Works in either fast or\nthinking mode. Default false preserves existing behavior. Also gated\nserver-side by a repo-level config flag (SearchService's\ngraphVectorPruneEnabled) \u2014 if that flag is off, this is forced to false\nregardless of what the request sets, so a deployment can disable the\nmechanism without any client-side change.", "example": true, "type": "boolean" }, @@ -5444,7 +5555,7 @@ "type": "string" }, "temporal_reasoning": { - "description": "TemporalReasoning activates the temporal read path: the query is classified\ninto a temporal mode (current/as-of/range/upcoming...), matching edge-level\ntemporal facts are resolved from the edge_temporal store and ride back on\nthe response (temporal_facts / temporal_duration / temporal_filter).\nCONTRACT: chunk ranking is NEVER altered — ON returns the same chunks as\nOFF; the layer is additive payload + computed answers only (rank shaping\nmeasured net-negative on BEAM/LongMemEval/TEMPO; see temporal_filters.go).\nOptional; ON by default — pass temporal_reasoning:false to disable.\nResolved by GetTemporalReasoningOrDefault (ownership rule).", + "description": "TemporalReasoning activates the temporal read path: the query is classified\ninto a temporal mode (current/as-of/range/upcoming...), matching edge-level\ntemporal facts are resolved from the edge_temporal store and ride back on\nthe response (temporal_facts / temporal_duration / temporal_filter).\nCONTRACT: chunk ranking is NEVER altered \u2014 ON returns the same chunks as\nOFF; the layer is additive payload + computed answers only (rank shaping\nmeasured net-negative on BEAM/LongMemEval/TEMPO; see temporal_filters.go).\nOptional; ON by default \u2014 pass temporal_reasoning:false to disable.\nResolved by GetTemporalReasoningOrDefault (ownership rule).", "example": true, "type": "boolean" }, @@ -5726,7 +5837,7 @@ "description": "TemporalDuration is the computed event-duration answer, when resolved.", "properties": { "approximate": { - "description": "Approximate is set when either endpoint's granularity is coarser than a\nday (month/year brackets) — the day count is then a floor-to-floor\nestimate, not an exact span; consumers should not present it as exact.", + "description": "Approximate is set when either endpoint's granularity is coarser than a\nday (month/year brackets) \u2014 the day count is then a floor-to-floor\nestimate, not an exact span; consumers should not present it as exact.", "example": true, "type": "boolean" }, @@ -5826,7 +5937,7 @@ "description": "TemporalFilter reports what the temporal layer did for this request.", "properties": { "applied": { - "description": "Applied is true when the temporal layer engaged for a classified temporal\nquery — including when it matched zero dated facts; MatchedFacts carries the\nactual count. It is false only when the fact lookup degraded (Degraded).", + "description": "Applied is true when the temporal layer engaged for a classified temporal\nquery \u2014 including when it matched zero dated facts; MatchedFacts carries the\nactual count. It is false only when the fact lookup degraded (Degraded).", "example": true, "type": "boolean" }, @@ -5835,7 +5946,7 @@ "type": "integer" }, "degraded": { - "description": "Degraded is true when the fact lookup FAILED (as opposed to matching\nnothing) — callers must not read an empty payload as \"no temporal facts\nexist\" when this is set.", + "description": "Degraded is true when the fact lookup FAILED (as opposed to matching\nnothing) \u2014 callers must not read an empty payload as \"no temporal facts\nexist\" when this is set.", "example": true, "type": "boolean" }, @@ -5863,7 +5974,7 @@ "type": "object" }, "search.TemporalIntentOverride": { - "description": "TemporalIntent (EXPERIMENTAL) lets the caller supply the classification\n(mode/window/phrases) directly, bypassing the regex classifier — for\nagents whose own LLM already understands the query, and for non-English\nqueries. Invalid overrides fall back to the classifier.", + "description": "TemporalIntent (EXPERIMENTAL) lets the caller supply the classification\n(mode/window/phrases) directly, bypassing the regex classifier \u2014 for\nagents whose own LLM already understands the query, and for non-English\nqueries. Invalid overrides fall back to the classifier.", "properties": { "cutoff": { "type": "string" @@ -5896,7 +6007,7 @@ "properties": { "additional_metadata": { "additionalProperties": {}, - "description": "Pydantic aliases (see VectorStoreChunk): document_metadata→additional_metadata,\ntenant_metadata→metadata. FastAPI serializes by_alias, so the wire uses the aliases.", + "description": "Pydantic aliases (see VectorStoreChunk): document_metadata\u2192additional_metadata,\ntenant_metadata\u2192metadata. FastAPI serializes by_alias, so the wire uses the aliases.", "example": { "author": "ada", "doc_version": 3 @@ -6219,7 +6330,7 @@ "properties": { "additional_metadata": { "additionalProperties": {}, - "description": "Pydantic aliases: document_metadata→additional_metadata, tenant_metadata→metadata.\nFastAPI serializes responses by_alias, so the wire uses the alias names.", + "description": "Pydantic aliases: document_metadata\u2192additional_metadata, tenant_metadata\u2192metadata.\nFastAPI serializes responses by_alias, so the wire uses the alias names.", "example": { "author": "ada", "doc_version": 3 @@ -6516,7 +6627,7 @@ "type": "boolean" }, "ready_for_ingestion": { - "description": "Derived readiness flag: true only when scheduler_status (lifecycle provisioning finished), graph_status, and both vectorstore_status.knowledge and vectorstore_status.memories are true — i.e. the database is fully provisioned and ready to accept ingestion and serve queries. Database creation is asynchronous: collections may appear before provisioning completes, so poll GET /databases/status until this is true before ingesting or querying.", + "description": "Derived readiness flag: true only when scheduler_status (lifecycle provisioning finished), graph_status, and both vectorstore_status.knowledge and vectorstore_status.memories are true \u2014 i.e. the database is fully provisioned and ready to accept ingestion and serve queries. Database creation is asynchronous: collections may appear before provisioning completes, so poll GET /databases/status until this is true before ingesting or querying.", "example": true, "type": "boolean" }, @@ -6637,7 +6748,6 @@ { "data_type": "VARCHAR", "enable_dense_embedding": true, - "enable_match": true, "enable_sparse_embedding": false, "max_length": 256, "name": "category" @@ -6822,7 +6932,6 @@ { "data_type": "VARCHAR", "enable_dense_embedding": true, - "enable_match": true, "enable_sparse_embedding": false, "max_length": 256, "name": "category" @@ -6846,15 +6955,12 @@ "tenants.TenantMetadataSchemaUpdateRequest": { "properties": { "add_fields": { - "description": "New metadata schema fields to add to the database. Additive only — no deletes, renames, or type changes.", + "description": "New metadata schema fields to add to the database. Additive only \u2014 no deletes, renames, or type changes.", "example": [ { + "name": "region", "data_type": "VARCHAR", - "enable_dense_embedding": true, - "enable_match": true, - "enable_sparse_embedding": false, - "max_length": 256, - "name": "category" + "max_length": 256 } ], "items": { @@ -7329,7 +7435,7 @@ "email": "support@hydradb.com", "name": "HydraDB Support" }, - "description": "HydraDB Application API — knowledge ingestion, search, and memory management.", + "description": "HydraDB Application API \u2014 knowledge ingestion, search, and memory management.", "license": { "name": "Proprietary" }, @@ -8471,10 +8577,10 @@ }, "/context": { "delete": { - "description": "Delete one or more knowledge sources or memories by ID.\n\nBy default this endpoint answers 200 for every outcome, including a delete\nthat removed nothing — check `data.deleted_count` and `data.results` rather\nthan the status code.\n\nSend `X-HydraDB-Delete-Status: strict` to opt in to honest status codes: a\ndelete that did not happen then answers 404/409/500 and never 200. This is\nthe recommended mode for new integrations. On those failures the response\n`data` still carries the same `results` / `deleted_count` payload a 200\ncarries, so per-id outcomes stay readable either way.\n\nThe default is expected to become strict in a future release, at which\npoint `X-HydraDB-Delete-Status: legacy` keeps the unconditional 200 for a\ncaller that is not ready.", + "description": "Delete one or more knowledge sources or memories by ID.\n\nBy default this endpoint answers 200 for every outcome, including a delete\nthat removed nothing \u2014 check `data.deleted_count` and `data.results` rather\nthan the status code.\n\nSend `X-HydraDB-Delete-Status: strict` to opt in to honest status codes: a\ndelete that did not happen then answers 404/409/500 and never 200. This is\nthe recommended mode for new integrations. On those failures the response\n`data` still carries the same `results` / `deleted_count` payload a 200\ncarries, so per-id outcomes stay readable either way.\n\nThe default is expected to become strict in a future release, at which\npoint `X-HydraDB-Delete-Status: legacy` keeps the unconditional 200 for a\ncaller that is not ready.", "parameters": [ { - "description": "Selects the status behaviour for this request. `strict` opts in to honest 404/409/500 codes when the delete did not happen; `legacy` forces the unconditional 200. Omitted, the server default applies — currently `legacy`.", + "description": "Selects the status behaviour for this request. `strict` opts in to honest 404/409/500 codes when the delete did not happen; `legacy` forces the unconditional 200. Omitted, the server default applies \u2014 currently `legacy`.", "in": "header", "name": "X-HydraDB-Delete-Status", "schema": { @@ -10018,7 +10124,7 @@ "x-fern-sdk-method-name": "get_metadata_schema" }, "patch": { - "description": "Add new metadata schema fields to an existing database. Additive only — existing fields cannot be deleted or retyped.", + "description": "Add new metadata schema fields to an existing database. Additive only \u2014 existing fields cannot be deleted or retyped.", "parameters": [ { "description": "Database identifier", From f2f357fa3011fed97617650f14526bbbacaa001c Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Tue, 6 Oct 2026 15:46:19 +0530 Subject: [PATCH 05/11] docs(v2): model nullable fetch errors in the API schema Signed-off-by: SohamRatnaparkhi --- api-reference/v2/openapi.json | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/api-reference/v2/openapi.json b/api-reference/v2/openapi.json index 3b896fc3..de0d7784 100644 --- a/api-reference/v2/openapi.json +++ b/api-reference/v2/openapi.json @@ -426,9 +426,12 @@ "type": "string" }, "error": { - "description": "Error message, empty string on success.", - "example": "", - "type": "string" + "description": "Legacy per-source error string; null when the fetch succeeds.", + "example": null, + "type": [ + "string", + "null" + ] }, "id": { "description": "Unique identifier for this resource.", @@ -1161,7 +1164,7 @@ "example": { "content": "# Q4 Report\n\nRevenue grew 23% quarter over quarter.", "content_type": "text/markdown", - "error": "", + "error": null, "id": "HydraDoc1234", "inferred_content": "Summary: Q4 revenue rose 23% QoQ, driven by enterprise expansion.", "message": "Success", From 3dc9455e6939efdc471fc9ef8e12ee9efc449b09 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Tue, 6 Oct 2026 15:46:47 +0530 Subject: [PATCH 06/11] docs(v2): allow null fetch fields in content and URL modes Signed-off-by: SohamRatnaparkhi --- api-reference/v2/openapi.json | 30 ++++++++++++++++++++++++------ 1 file changed, 24 insertions(+), 6 deletions(-) diff --git a/api-reference/v2/openapi.json b/api-reference/v2/openapi.json index de0d7784..332ab76d 100644 --- a/api-reference/v2/openapi.json +++ b/api-reference/v2/openapi.json @@ -414,16 +414,25 @@ "content": { "description": "Stored original bytes as text when valid UTF-8. Binary originals, such as PDF and DOCX, are returned in content_base64 instead. Null in url mode.", "example": "# Q4 Report\n\nRevenue grew 23% quarter over quarter.", - "type": "string" + "type": [ + "string", + "null" + ] }, "content_base64": { "description": "Base64-encoded binary content, for binary file types.", - "type": "string" + "type": [ + "string", + "null" + ] }, "content_type": { "description": "MIME type of the source (e.g. `application/pdf`, `text/plain`).", "example": "application/pdf", - "type": "string" + "type": [ + "string", + "null" + ] }, "error": { "description": "Legacy per-source error string; null when the fetch succeeds.", @@ -441,7 +450,10 @@ "inferred_content": { "description": "Model-derived content when available in content or both mode. Null in url mode.", "example": "Summary: Q4 revenue rose 23% QoQ, driven by enterprise expansion.", - "type": "string" + "type": [ + "string", + "null" + ] }, "message": { "description": "Human-readable result message.", @@ -451,12 +463,18 @@ "presigned_url": { "description": "Time-limited download URL for the original file.", "example": "https://storage.hydradb.com/sources/HydraDoc1234?sig=...", - "type": "string" + "type": [ + "string", + "null" + ] }, "size_bytes": { "description": "File size in bytes.", "example": 20480, - "type": "integer" + "type": [ + "integer", + "null" + ] }, "success": { "deprecated": true, From 4d42afea79f9adee61194ea6c0df0f662a8d9e5b Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Wed, 7 Oct 2026 15:59:41 +0530 Subject: [PATCH 07/11] docs(v2): explain API results before advanced graph fields Signed-off-by: SohamRatnaparkhi --- api-reference/v2/endpoint/query-overview.mdx | 14 +++++++------- api-reference/v2/endpoint/source-relations.mdx | 2 +- api-reference/v2/endpoint/subgraph.mdx | 2 +- api-reference/v2/index.mdx | 14 ++++++++------ essentials/v2/api-results.mdx | 14 +++++++++----- 5 files changed, 26 insertions(+), 20 deletions(-) diff --git a/api-reference/v2/endpoint/query-overview.mdx b/api-reference/v2/endpoint/query-overview.mdx index cc067607..016b9339 100644 --- a/api-reference/v2/endpoint/query-overview.mdx +++ b/api-reference/v2/endpoint/query-overview.mdx @@ -5,7 +5,7 @@ description: "Choose type, query_by, and mode for a query, with recipes for comm import { Field } from "/snippets/field.jsx"; -Use this page to choose the right query shape before opening the full [Query](/api-reference/v2/endpoint/query) endpoint reference. Query has three main decisions: what to query (`type`), how to match (`query_by`), and how much retrieval work to spend (`mode`). +A query searches stored content and returns ranked passages (**chunks**) with their source details. It does not produce a model-written answer. Use this page to choose the right query shape before opening the full [Query](/api-reference/v2/endpoint/query) endpoint reference. Query has three main decisions: what to query (`type`), how to match (`query_by`), and how much retrieval work to spend (`mode`). ```mermaid flowchart LR @@ -37,9 +37,9 @@ linkStyle default stroke:#64748b,stroke-width:2px; |---|---|---| | | `"knowledge"`, `"memory"`, `"all"` | Choose what to query. Use `"knowledge"` for shared docs/app sources, `"memory"` for user context, and `"all"` when an answer should use both. | | | `"hybrid"`, `"text"` | Choose the matching method. Use `"hybrid"` by default and `"text"` for exact terms or phrases. | -| | `"fast"`, `"thinking"`, `"auto"` | Choose latency vs quality, or let HydraDB decide. Use `"fast"` for low-latency paths, `"thinking"` for multi-query retrieval, reranking, and forceful-relation context, and `"auto"` to score the query and route to one of the two automatically (defaults to `"thinking"` when the signal is inconclusive; **the default if `mode` is omitted**). | +| | `"fast"`, `"thinking"`, `"auto"` | Choose latency vs quality, or let HydraDB decide. Use `"fast"` for low-latency paths, `"thinking"` for multi-query retrieval, reordering results for relevance, and context from explicit links between sources, and `"auto"` to score the query and route to one of the two automatically (defaults to `"thinking"` when the signal is inconclusive; **the default if `mode` is omitted**). | | | integer | Control prompt size. Default `10`, maximum `250`. Start with `10`, reduce for tight context windows, increase only when you rerank or summarize downstream. | -| | `0.0` to `1.0`, or `"auto"` | Tune hybrid query. Lower values favor BM25 keywords; higher values favor semantic similarity. Default `0.8`, which `"auto"` also resolves to. | +| | `0.0` to `1.0`, or `"auto"` | Tune hybrid query. Lower values favor keyword matching; higher values favor semantic similarity. Default `0.8`, which `"auto"` also resolves to. | | | object | Narrow candidates before ranking. Top-level keys match `metadata`; nested `additional_metadata` filters free-form per-source fields. | | | `string[]` or weighted object | Query one or more user/workspace/team scopes. A list uses equal normalized weights; an object like `{ "workspace_42": 2, "user_alex": 1 }` applies relative ranking weights with at most one decimal place. Max 100 collections, and each must exist. | | | boolean | Include entity/relation context with the chunks. On by default; set `false` with `mode: "fast"` for chunk-only responses (`thinking` always includes it). | @@ -49,8 +49,8 @@ linkStyle default stroke:#64748b,stroke-width:2px; | User intent | Recommended config | |---|---| -| Fast document RAG | `type="knowledge"`, `query_by="hybrid"`, `mode="fast"`, `max_results=5-10`, `graph_context=false` | -| Highest-quality document RAG | `type="knowledge"`, `query_by="hybrid"`, `mode="thinking"` (thinking always includes graph context) | +| Fast answers from documents | `type="knowledge"`, `query_by="hybrid"`, `mode="fast"`, `max_results=5-10`, `graph_context=false` | +| Broader document search | `type="knowledge"`, `query_by="hybrid"`, `mode="thinking"` (thinking always includes graph context) | | Personalized answer | `type="all"`, the user's `collection`, `query_by="hybrid"`, `mode="thinking"`. If shared docs live in another collection, list both in `collections` | | User preferences only | `type="memory"`, include `collection`, `query_by="hybrid"` | | Exact keyword or phrase | `type="knowledge"`, `query_by="text"`, `operator="phrase"` | @@ -62,7 +62,7 @@ linkStyle default stroke:#64748b,stroke-width:2px; -Use this for standard RAG over docs, PDFs, tickets, pages, or app sources. +Use this for answering questions from docs, PDFs, tickets, pages, or app sources. ```json { @@ -117,7 +117,7 @@ Use this when the same question should search several collection scopes and retu -Use `metadata_filters` when you already know the slice you want. Top-level keys match schema-backed `metadata` fields; declare hot filters in `database_metadata_schema`. Free-form per-source fields go under `additional_metadata` (`document_metadata` is a legacy alias). Multiple filters are ANDed constraints. +Use `metadata_filters` when you already know the slice you want. Top-level keys match schema-backed `metadata` fields; declare frequently used fields in `database_metadata_schema`. Free-form per-source fields go under `additional_metadata` (`document_metadata` is a legacy alias). Multiple filters are ANDed constraints. ```json { diff --git a/api-reference/v2/endpoint/source-relations.mdx b/api-reference/v2/endpoint/source-relations.mdx index a15e6771..7a943b7f 100644 --- a/api-reference/v2/endpoint/source-relations.mdx +++ b/api-reference/v2/endpoint/source-relations.mdx @@ -6,7 +6,7 @@ description: "Read the entity-relation triplets extracted from your content, for import { Field } from "/snippets/field.jsx"; -This endpoint queries entity-and-relationship triplets extracted from your ingested content. +Inspect relationships extracted from your content, such as `PaymentsWorker → depends_on → OrdersDB`. Each three-part relationship is a **triplet**: a starting entity (a named person, service, or topic), a relationship, and a target entity. Pass `id` to scope to a single ingested item, or omit it to return all relations in the collection. Set `type=memory` to inspect a memory's relations. Pagination handles large result sets. diff --git a/api-reference/v2/endpoint/subgraph.mdx b/api-reference/v2/endpoint/subgraph.mdx index a93d159c..5a2c61cd 100644 --- a/api-reference/v2/endpoint/subgraph.mdx +++ b/api-reference/v2/endpoint/subgraph.mdx @@ -6,7 +6,7 @@ description: "Everything connected to one item: its thread, its replies, its par import { Field } from "/snippets/field.jsx"; -This endpoint returns the **connected subgraph** of one ingested item: every item reachable from it through item-level relations, traversed breadth-first up to `depth` hops, together with the relations among those members and the structural graph around them (entities, comments, attachments, people). +Get the content connected to one stored item: for example, the replies around a Slack message or the pages linked from a wiki page. This set is the item's **connected subgraph**. The endpoint follows links up to `depth` connections (**hops**), visits nearby items first, and returns connected items and their relationships. `max_sources` limits how many items it returns. It answers a different question from [Inspecting Context Relations](/api-reference/v2/endpoint/source-relations). Relations are the entity-and-predicate triplets *extracted from text* (`PaymentsWorker` `depends_on` `OrdersDB`). The subgraph is about *items*: which Slack message replies to which, which page links to which, which ticket a comment belongs to. Use it after [Query](/api-reference/v2/endpoint/query) or [List Documents](/api-reference/v2/endpoint/list-documents) when a single result is not enough and you need what surrounds it. diff --git a/api-reference/v2/index.mdx b/api-reference/v2/index.mdx index 9227743a..563b68be 100644 --- a/api-reference/v2/index.mdx +++ b/api-reference/v2/index.mdx @@ -3,6 +3,8 @@ title: "API Reference" description: "Every v2 endpoint, and the conventions they share." --- +Use this reference to look up request fields and response formats. If you are making your first call, follow the [Quickstart](/get-started/v2/quickstart). **Context** means the documents, passages, and memories your model uses to answer a question; HydraDB stores and retrieves them, while your model writes the answer. + ## Quick links - **New to HydraDB?** Start with the [Quickstart](/get-started/v2/quickstart) @@ -18,7 +20,7 @@ description: "Every v2 endpoint, and the conventions they share." |---|---|---| | [Databases](/api-reference/v2/endpoint/tenants-overview) | Create, monitor, and manage isolated workspaces | First step in any integration, and any time you need usage stats, provisioning status, or to tear down a workspace | | [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB: document uploads, app sources, user memories, and lifecycle ops | -| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | At query time: the only endpoint you call to feed an LLM | +| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | Find relevant passages to include in your model prompt | | [Connectors](/api-reference/v2/endpoint/connectors-overview) | Connect, configure, sync, and manage app connectors such as Slack, GitHub, Google Drive, and Supabase | When you want app data synced into a database without writing ingest code | | [Webhooks](/essentials/v2/webhooks) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | @@ -26,10 +28,10 @@ description: "Every v2 endpoint, and the conventions they share." | Concept | What it means | When you use it | |---|---|---| -| `database` | Your isolated workspace for data, metadata schema, and query. | Send it on every API call so HydraDB knows which workspace to read or write. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | -| `collection` | Optional partition inside a database, often a user, team, account, or customer. | Use it when one database contains data for multiple users or customers. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). Read more about our [multi-tenant architecture](/essentials/v2/multi-tenant) | -| [Knowledge](/essentials/v2/knowledge) | Shared source material such as PDFs, docs, app pages, tickets, Slack threads, or webpages. | Use `type=knowledge` when many users or agents should query the same content. | -| [Memory](/essentials/v2/memories) | User-specific context such as preferences, conversation history, notes, and inferred traits. | Use `type=memory` when the content should personalize answers for a specific user or collection. | +| `database` | Your isolated workspace for data, metadata schema, and query. | Send it on database-scoped calls such as ingestion and query. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | +| `collection` | Optional partition inside a database, often a user, team, account, or customer. | Send the same collection on writes and reads. Omitting it uses the default collection, not all collections. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). Read more about our [multi-tenant architecture](/essentials/v2/multi-tenant) | +| [Knowledge](/essentials/v2/knowledge) | Source material such as PDFs, docs, app pages, tickets, Slack threads, or webpages. | Use `type=knowledge` to search documents and app content. | +| [Memory](/essentials/v2/memories) | Preferences, conversation history, notes, and saved facts. | Use `type=memory` when the content should personalize answers for a specific user or collection. | | `database_metadata_schema` | Database-level fields you define up front so metadata can be filtered or queried consistently. | Use it for stable fields like department, customer, region, plan, category, or compliance label. | | `document_metadata` | JSON-stringified per-document metadata array sent during file ingestion. | Use it to attach a source `id`, schema-backed `metadata`, free-form `additional_metadata`, and relations to other sources to each uploaded file, in the same order as `documents`. | | `ids` | IDs returned by ingestion or visible from `/context/list`. | Use them when polling processing status, inspecting content, listing a specific subset, deleting sources, or inspecting relations. | @@ -214,6 +216,6 @@ Rate limits apply per API key. For production deployments, build retry logic wit Existing v1 endpoints remain available under the v1 API Reference. - **Build something:** [Quickstart](/get-started/v2/quickstart) walks through your first integration in five minutes -- **Understand the model:** [Core Concepts](/get-started/v2/core-concepts) explains the five primitives +- **Understand the model:** [Core Concepts](/get-started/v2/core-concepts) explains databases, collections, content categories, and query results - **Go deeper:** [Usage](/essentials/v2/query) covers each primitive in depth - **Install an SDK:** [Python](https://pypi.org/project/hydradb-sdk/) · [TypeScript](https://www.npmjs.com/package/@hydradb/sdk) diff --git a/essentials/v2/api-results.mdx b/essentials/v2/api-results.mdx index 33e4a5e4..2c14f479 100644 --- a/essentials/v2/api-results.mdx +++ b/essentials/v2/api-results.mdx @@ -3,7 +3,7 @@ title: "How to Use API Results" description: "Turn a query response into an LLM prompt." --- -`POST /query` returns structured JSON. Before you can pass it to an LLM, you need to convert the retrieval payload into a plain string. This page shows how. +`POST /query` returns matching passages (**chunks**) and source details, not a generated answer. Format these results as **context** (information supplied alongside the question), then pass it to your language model. The SDK helper below handles the formatting; you do not need to interpret the graph fields to get started. --- @@ -11,7 +11,9 @@ description: "Turn a query response into an LLM prompt." The SDKs return the full response envelope from `POST /query`; the retrieval payload is on its `data` field (e.g. `result.data`). Raw HTTP responses use the same envelope, with the payload under `data`. The `build_string` / `buildString` helper accepts either the full envelope or just the `data` payload, so you can pass the SDK return value to it directly. -The retrieval payload has the same core shape regardless of `type` or `query_by`: +The `data` payload has the same core shape regardless of `type` or `query_by`. Expand this reference when you need individual fields; continue to Section 2 to format a result. + + ```json { @@ -73,12 +75,14 @@ The retrieval payload has the same core shape regardless of `type` or `query_by` Four things matter for prompt construction: -- **`chunks`**: the primary retrieval output. Ranked by relevance; preserve the order HydraDB returns. -- **`graph_context.query_paths`**: entity traversal paths derived from your query. Useful for relational reasoning. See [Context Graphs](/essentials/v2/context-graphs). +- **`chunks`**: passages from matching documents or memories. Ranked by relevance; preserve the order HydraDB returns. +- **`graph_context.query_paths`**: chains of relationships connecting named people, services, or topics relevant to the query. See [Context Graphs](/essentials/v2/context-graphs). - **`graph_context.chunk_relations`** + **`chunk_id_to_group_ids`**: per-chunk graph relations grouped by `group_id`, so you can attach the right triplets to each chunk. - **`additional_context`**: a map keyed by `chunk_uuid`. When a chunk includes `extra_context_ids`, use those IDs to look up related chunks here. -The raw object has too much noise for an LLM: IDs, timestamps, metadata. Section 2 shows how to convert it into a clean string. +The helper in Section 2 formats passages, source details, and any relationships together. + + --- From 1afa03f17aa414fce6ffa0b91c5c1ef43b99a533 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Thu, 8 Oct 2026 16:09:16 +0530 Subject: [PATCH 08/11] docs(v2): put the four core calls first in the API reference Most integrations are create database, ingest, check status, query. The sidebar now opens with a "Start here" group holding exactly those four endpoints and the overview, and the overview page leads with the same four calls before the endpoint groups and the full inventory. The resource groups still list every endpoint, so browsing by resource works as before. SDKs and Error Responses move to a Reference group. Refs PRO-2457 Co-Authored-By: Claude Fable 5.1 Signed-off-by: SohamRatnaparkhi --- api-reference/v2/index.mdx | 62 ++++++++++++++++++-------------------- docs.json | 56 ++++++++++++++++++++++------------ 2 files changed, 66 insertions(+), 52 deletions(-) diff --git a/api-reference/v2/index.mdx b/api-reference/v2/index.mdx index 563b68be..64c13b17 100644 --- a/api-reference/v2/index.mdx +++ b/api-reference/v2/index.mdx @@ -1,42 +1,18 @@ --- title: "API Reference" -description: "Every v2 endpoint, and the conventions they share." +description: "The four calls most integrations are built on, then every v2 endpoint." --- -Use this reference to look up request fields and response formats. If you are making your first call, follow the [Quickstart](/get-started/v2/quickstart). **Context** means the documents, passages, and memories your model uses to answer a question; HydraDB stores and retrieves them, while your model writes the answer. +Use this reference to look up request fields and response formats. Every request carries `Authorization: Bearer ` and `API-Version: 2` against `https://api.hydradb.com`; the [SDKs](/api-reference/v2/sdks) set both for you. If you are making your first call, follow the [Quickstart](/get-started/v2/quickstart). AI agents can start from the [Agent Integration Guide](/AGENTS) and the [v2 OpenAPI spec](/api-reference/v2/openapi.json). **Context** means the documents, passages, and memories your model uses to answer a question; HydraDB stores and retrieves them, while your model writes the answer. -## Quick links +## The four calls -- **New to HydraDB?** Start with the [Quickstart](/get-started/v2/quickstart) -- **Prefer SDKs?** See [SDKs (Node and Python)](/api-reference/v2/sdks) -- **Authentication:** Every endpoint requires `Authorization: Bearer ` -- **Base URL:** `https://api.hydradb.com` -- **Errors:** See [Error Responses](/api-reference/v2/error-responses) -- **For AI agents:** See the [Agent Integration Guide](/AGENTS) and [v2 OpenAPI spec](/api-reference/v2/openapi.json) +Most integrations are these four requests, in this order. Everything else in this reference manages what they create. -## Endpoint groups - -| Group | Purpose | When to reach for it | -|---|---|---| -| [Databases](/api-reference/v2/endpoint/tenants-overview) | Create, monitor, and manage isolated workspaces | First step in any integration, and any time you need usage stats, provisioning status, or to tear down a workspace | -| [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB: document uploads, app sources, user memories, and lifecycle ops | -| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | Find relevant passages to include in your model prompt | -| [Connectors](/api-reference/v2/endpoint/connectors-overview) | Connect, configure, sync, and manage app connectors such as Slack, GitHub, Google Drive, and Supabase | When you want app data synced into a database without writing ingest code | -| [Webhooks](/essentials/v2/webhooks) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | - -## Core concepts - -| Concept | What it means | When you use it | -|---|---|---| -| `database` | Your isolated workspace for data, metadata schema, and query. | Send it on database-scoped calls such as ingestion and query. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | -| `collection` | Optional partition inside a database, often a user, team, account, or customer. | Send the same collection on writes and reads. Omitting it uses the default collection, not all collections. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). Read more about our [multi-tenant architecture](/essentials/v2/multi-tenant) | -| [Knowledge](/essentials/v2/knowledge) | Source material such as PDFs, docs, app pages, tickets, Slack threads, or webpages. | Use `type=knowledge` to search documents and app content. | -| [Memory](/essentials/v2/memories) | Preferences, conversation history, notes, and saved facts. | Use `type=memory` when the content should personalize answers for a specific user or collection. | -| `database_metadata_schema` | Database-level fields you define up front so metadata can be filtered or queried consistently. | Use it for stable fields like department, customer, region, plan, category, or compliance label. | -| `document_metadata` | JSON-stringified per-document metadata array sent during file ingestion. | Use it to attach a source `id`, schema-backed `metadata`, free-form `additional_metadata`, and relations to other sources to each uploaded file, in the same order as `documents`. | -| `ids` | IDs returned by ingestion or visible from `/context/list`. | Use them when polling processing status, inspecting content, listing a specific subset, deleting sources, or inspecting relations. | - -## End-to-end lifecycle +1. [Create Database](/api-reference/v2/endpoint/create-tenant): `POST /databases`. One isolated workspace per customer, environment, or product. Creation runs in the background, so poll [Database Status](/api-reference/v2/endpoint/tenant-status) until `ready_for_ingestion` is `true`. +2. [Ingest Context](/api-reference/v2/endpoint/ingest-context): `POST /context/ingest`. Files, app sources, or memories, in one multipart request. +3. [Ingestion Status](/api-reference/v2/endpoint/source-status): `GET /context/status`. Poll the returned ids until `indexing_status` is `completed`, or [register a webhook](/essentials/v2/webhooks) instead. +4. [Query](/api-reference/v2/endpoint/query): `POST /query`. Retrieve the context your model needs, from knowledge, memories, or both. ```mermaid flowchart LR @@ -60,6 +36,28 @@ flowchart LR style G fill:#0f172a,stroke:#334155,stroke-width:2px,color:#f8fafc ``` +## Endpoint groups + +| Group | Purpose | When to reach for it | +|---|---|---| +| [Databases](/api-reference/v2/endpoint/tenants-overview) | Create, monitor, and manage isolated workspaces | First step in any integration, and any time you need usage stats, provisioning status, or to tear down a workspace | +| [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB: document uploads, app sources, user memories, and lifecycle ops | +| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | Find relevant passages to include in your model prompt | +| [Connectors](/api-reference/v2/endpoint/connectors-overview) | Connect, configure, sync, and manage app connectors such as Slack, GitHub, Google Drive, and Supabase | When you want app data synced into a database without writing ingest code | +| [Webhooks](/essentials/v2/webhooks) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | + +## Core concepts + +| Concept | What it means | When you use it | +|---|---|---| +| `database` | Your isolated workspace for data, metadata schema, and query. | Send it on database-scoped calls such as ingestion and query. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | +| `collection` | Optional partition inside a database, often a user, team, account, or customer. | Send the same collection on writes and reads. Omitting it uses the default collection, not all collections. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). Read more about our [multi-tenant architecture](/essentials/v2/multi-tenant) | +| [Knowledge](/essentials/v2/knowledge) | Source material such as PDFs, docs, app pages, tickets, Slack threads, or webpages. | Use `type=knowledge` to search documents and app content. | +| [Memory](/essentials/v2/memories) | Preferences, conversation history, notes, and saved facts. | Use `type=memory` when the content should personalize answers for a specific user or collection. | +| `database_metadata_schema` | Database-level fields you define up front so metadata can be filtered or queried consistently. | Use it for stable fields like department, customer, region, plan, category, or compliance label. | +| `document_metadata` | JSON-stringified per-document metadata array sent during file ingestion. | Use it to attach a source `id`, schema-backed `metadata`, free-form `additional_metadata`, and relations to other sources to each uploaded file, in the same order as `documents`. | +| `ids` | IDs returned by ingestion or visible from `/context/list`. | Use them when polling processing status, inspecting content, listing a specific subset, deleting sources, or inspecting relations. | + ## SDKs HydraDB publishes official SDKs for Python and TypeScript/Node. They wrap every endpoint in this reference with typed methods and IDE autocomplete. diff --git a/docs.json b/docs.json index ef8780a2..2dd8f792 100644 --- a/docs.json +++ b/docs.json @@ -22,7 +22,11 @@ }, "favicon": "/favicon.png", "contextual": { - "options": ["copy", "chatgpt", "claude"] + "options": [ + "copy", + "chatgpt", + "claude" + ] }, "integrations": { "posthog": { @@ -33,7 +37,11 @@ }, "api": { "examples": { - "languages": ["python", "javascript", "curl"] + "languages": [ + "python", + "javascript", + "curl" + ] } }, "navigation": { @@ -101,7 +109,9 @@ { "group": "For Agents", "public": true, - "pages": ["AGENTS"] + "pages": [ + "AGENTS" + ] } ] }, @@ -133,43 +143,45 @@ "tab": "API Reference", "groups": [ { - "group": "API Documentation", - "public": true, - "pages": ["api-reference/v2/index", "api-reference/v2/sdks"] + "group": "Start here", + "pages": [ + "api-reference/v2/index", + "api-reference/v2/endpoint/create-tenant", + "api-reference/v2/endpoint/ingest-context", + "api-reference/v2/endpoint/source-status", + "api-reference/v2/endpoint/query" + ] }, { "group": "Databases", - "public": true, "pages": [ "api-reference/v2/endpoint/tenants-overview", "api-reference/v2/endpoint/create-tenant", - "api-reference/v2/endpoint/update-metadata-schema", - "api-reference/v2/endpoint/list-tenants", - "api-reference/v2/endpoint/delete-tenant", "api-reference/v2/endpoint/tenant-status", + "api-reference/v2/endpoint/list-tenants", + "api-reference/v2/endpoint/tenant-stats", + "api-reference/v2/endpoint/update-metadata-schema", "api-reference/v2/endpoint/list-sub-tenants", "api-reference/v2/endpoint/delete-collection", - "api-reference/v2/endpoint/tenant-stats" + "api-reference/v2/endpoint/delete-tenant" ] }, { "group": "Context", - "public": true, "pages": [ "api-reference/v2/endpoint/sources-overview", "api-reference/v2/endpoint/ingest-context", "api-reference/v2/endpoint/source-status", - "api-reference/v2/endpoint/fetch-content", "api-reference/v2/endpoint/list-documents", + "api-reference/v2/endpoint/fetch-content", "api-reference/v2/endpoint/update-source-metadata", - "api-reference/v2/endpoint/delete-source", "api-reference/v2/endpoint/source-relations", - "api-reference/v2/endpoint/subgraph" + "api-reference/v2/endpoint/subgraph", + "api-reference/v2/endpoint/delete-source" ] }, { "group": "Query", - "public": true, "pages": [ "api-reference/v2/endpoint/query-overview", "api-reference/v2/endpoint/query", @@ -205,9 +217,11 @@ ] }, { - "group": "Miscellaneous", - "public": true, - "pages": ["api-reference/v2/error-responses"] + "group": "Reference", + "pages": [ + "api-reference/v2/sdks", + "api-reference/v2/error-responses" + ] } ] } @@ -259,7 +273,9 @@ { "group": "For Agents", "public": true, - "pages": ["AGENTS"] + "pages": [ + "AGENTS" + ] } ] }, From 05a422c69779534357b561754766f4dac4dc3bc9 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Thu, 8 Oct 2026 20:43:59 +0530 Subject: [PATCH 09/11] docs(v2): link BYOG references to the merged page's non-Cypher section The two Bring Your Own Graph pages are being merged into one (PR #304, carried into #309), so the ingest and query references point at the section that covers graph_payload. Refs PRO-2457 Co-Authored-By: Claude Fable 5.1 Signed-off-by: SohamRatnaparkhi --- api-reference/v2/endpoint/ingest-context.mdx | 4 ++-- api-reference/v2/endpoint/query.mdx | 2 +- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/api-reference/v2/endpoint/ingest-context.mdx b/api-reference/v2/endpoint/ingest-context.mdx index dfb7c27d..de42aef5 100644 --- a/api-reference/v2/endpoint/ingest-context.mdx +++ b/api-reference/v2/endpoint/ingest-context.mdx @@ -320,7 +320,7 @@ response.raise_for_status() | | Binary uploads, **knowledge only**. Required when `type=knowledge` and you want HydraDB to parse documents. Omit when ingesting memories. See [Supported file formats](#supported-file-formats). (default=`[]`) | | | One entry per file in `documents`, in the same order. The counts must match, or the request returns `400`. If omitted, documents index with inferred defaults such as filename/title. See the item shape below. | | | Pre-extracted source objects (Slack, Notion, web pages, etc.), **knowledge only**. See the `app_knowledge` item shape below. | -| | Map of source id to your own entities + relations. Replaces LLM graph extraction for each keyed source. Works for `type=knowledge` (key = a `document_metadata` id or `app_knowledge` item id) and `type=memory` (key = a memory `id`). See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) and the shape below. | +| | Map of source id to your own entities + relations. Replaces LLM graph extraction for each keyed source. Works for `type=knowledge` (key = a `document_metadata` id or `app_knowledge` item id) and `type=memory` (key = a memory `id`). See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph#using-byog-without-cypher) and the shape below. | | | Memory items, **memory only**. Required and non-empty when `type=memory`. Use plural `memories` for the form field, even though `type` is singular `memory`. See the `memories` item shape below. | @@ -710,7 +710,7 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in - `graph_payload` is a **map of source id to graph** that **replaces LLM graph extraction** for each keyed source. For `type=knowledge`, the key is a `document_metadata` id or an `app_knowledge` item id; for `type=memory`, the key is a memory `id`. Keyed sources are still chunked and embedded, so they stay searchable. See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) for the full guide. + `graph_payload` is a **map of source id to graph** that **replaces LLM graph extraction** for each keyed source. For `type=knowledge`, the key is a `document_metadata` id or an `app_knowledge` item id; for `type=memory`, the key is a memory `id`. Keyed sources are still chunked and embedded, so they stay searchable. See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph#using-byog-without-cypher) for the full guide. diff --git a/api-reference/v2/endpoint/query.mdx b/api-reference/v2/endpoint/query.mdx index 985662e7..d428fdfd 100644 --- a/api-reference/v2/endpoint/query.mdx +++ b/api-reference/v2/endpoint/query.mdx @@ -435,7 +435,7 @@ result = client.query( | | Maximum chunks to return. Default `10`; maximum `250`. Start with `10`, use `5` for tight prompts, and increase only when reranking downstream. | | | Hybrid weight (`1.0` = pure semantic, `0.0` = pure BM25). Applies to `query_by: "hybrid"` only. `"auto"` currently resolves to `0.8`. (default=`0.8`) | | | Boost newer content. Omitted applies a mild `0.4` tilt; send `0` to turn it off. (default=`0.4`) | -| | When `true`, includes the entity/relation graph slice in the response under `graph_context`. Set to `false` when you only need ranked chunks. Relations you supplied via [Bring Your Own Graph](/essentials/v2/bring-your-own-graph) appear here identically to extracted ones. (default=`true`). **`false` only takes effect in `fast` mode: `thinking`, and `auto` when it resolves to thinking, always include the graph slice.** | +| | When `true`, includes the entity/relation graph slice in the response under `graph_context`. Set to `false` when you only need ranked chunks. Relations you supplied via [Bring Your Own Graph](/essentials/v2/bring-your-own-graph#using-byog-without-cypher) appear here identically to extracted ones. (default=`true`). **`false` only takes effect in `fast` mode: `thinking`, and `auto` when it resolves to thinking, always include the graph slice.** | | | Pull author-declared related sources into `additional_context`. **Only takes effect when `mode` resolves to `"thinking"`;** silently skipped in `fast` mode, and under `mode: "auto"` whether it takes effect depends on the automatic routing decision. (default=`true`) | | | Request-time hint to guide retrieval (e.g., "user is on the billing page"). This is different from the response `additional_context` map. (default=`null`) | | | Deterministic narrowing before ranking. See [Filters](#decision-matrix). Each list holds at most 500 values, and the whole object is capped at 64 KiB of compact JSON, measured after operator objects are reduced to their values; over either returns `400`. (default=`null`) | From afb5c62c7fb6e49734cd260d015a2dd22dd1793f Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Fri, 9 Oct 2026 00:31:14 +0530 Subject: [PATCH 10/11] docs(v2): add a Webhooks overview and link connector instructions to Connectors - Webhooks: Overview, first in the Webhooks group, lists every webhook endpoint including the signing-secret calls, the typical sequence, and how deliveries are signed and retried (exponential backoff, 16 attempts or 48 hours on Cloud, per domain/webhooks/retry.go). - The connector API pages link to the custom instructions section of the Connectors guide, where that content now lives. Refs PRO-2457 Co-Authored-By: Claude Opus 5.5 (1M context) Signed-off-by: SohamRatnaparkhi --- .../v2/endpoint/create-connector.mdx | 2 +- .../v2/endpoint/update-connector-resource.mdx | 4 +- .../v2/endpoint/update-connector.mdx | 4 +- .../v2/endpoint/webhooks-overview.mdx | 45 +++++++++++++++++++ api-reference/v2/index.mdx | 2 +- docs.json | 1 + 6 files changed, 52 insertions(+), 6 deletions(-) create mode 100644 api-reference/v2/endpoint/webhooks-overview.mdx diff --git a/api-reference/v2/endpoint/create-connector.mdx b/api-reference/v2/endpoint/create-connector.mdx index 18f63ba4..a0c503dd 100644 --- a/api-reference/v2/endpoint/create-connector.mdx +++ b/api-reference/v2/endpoint/create-connector.mdx @@ -10,7 +10,7 @@ Creates a connector for one provider account. Nothing syncs yet: next, call [Dis Optional settings you can also pass here, and change later with [Update Connector](/api-reference/v2/endpoint/update-connector): -- `custom_instructions`: guidance for how this connector's documents are interpreted and indexed. See [Custom Ingestion Instructions](/essentials/v2/connector-instructions). +- `custom_instructions`: guidance for how this connector's documents are interpreted and indexed. See [Custom Ingestion Instructions](/essentials/v2/connectors#custom-ingestion-instructions). - `sync_interval_seconds`: how often scheduled syncs run. Omit it for the provider default (one hour for most providers). diff --git a/api-reference/v2/endpoint/update-connector-resource.mdx b/api-reference/v2/endpoint/update-connector-resource.mdx index 13a7a050..bb3c34af 100644 --- a/api-reference/v2/endpoint/update-connector-resource.mdx +++ b/api-reference/v2/endpoint/update-connector-resource.mdx @@ -8,7 +8,7 @@ Updates one configured resource in place. Send `custom_instructions`, `acl`, or `resource_id` in the path is the id from [List Connector Resources](/api-reference/v2/endpoint/connector-resources), for example a Slack channel id or a Supabase `schema.table` name. URL-encode it if it contains spaces or other reserved characters. -- **`custom_instructions`** sets instructions for documents from this resource only. They **replace** the connector-level instructions for this resource rather than adding to them. Send `""` to clear them so the resource uses the connector's instructions again. Up to 4,000 characters. They apply from the next sync; documents already synced are not processed again. See [Custom Ingestion Instructions](/essentials/v2/connector-instructions). +- **`custom_instructions`** sets instructions for documents from this resource only. They **replace** the connector-level instructions for this resource rather than adding to them. Send `""` to clear them so the resource uses the connector's instructions again. Up to 4,000 characters. They apply from the next sync; documents already synced are not processed again. See [Custom Ingestion Instructions](/essentials/v2/connectors#custom-ingestion-instructions). - **`acl`** sets who can read this resource's objects, and applies to already-synced objects on the next query, with no re-sync. Use `["__public__"]` to open the resource to everyone and `[]` to allow nobody. For providers whose own permissions HydraDB reads, those permissions take over again at the next sync; your rule applies whenever the provider reports the resource as public or its permissions cannot be read. See [Access Control](/essentials/v2/access-control). @@ -50,5 +50,5 @@ The response includes `custom_instructions` and `acl` only when the request set ## Related Resources - [Update Connector](/api-reference/v2/endpoint/update-connector): set the connector-level instructions these override -- [Custom Ingestion Instructions](/essentials/v2/connector-instructions): how connector and resource instructions combine +- [Custom Ingestion Instructions](/essentials/v2/connectors#custom-ingestion-instructions): how connector and resource instructions combine - [List Connector Resources](/api-reference/v2/endpoint/connector-resources): read the stored values back diff --git a/api-reference/v2/endpoint/update-connector.mdx b/api-reference/v2/endpoint/update-connector.mdx index f1055051..42cf77e6 100644 --- a/api-reference/v2/endpoint/update-connector.mdx +++ b/api-reference/v2/endpoint/update-connector.mdx @@ -6,7 +6,7 @@ openapi: "api-reference/v2/openapi.json PATCH /connectors/{id}" Updates a connector in place. Every field is optional and only the fields you send change, so it is safe to call without re-sending anything else. Every field is checked before anything is saved, so if one is invalid, nothing changes. Credentials are saved before the other settings, so in the rare case of a server error after that, only the credentials may have changed; retrying the same request is safe. -- **`custom_instructions`** sets the connector-level ingestion instructions, used by every resource that has no instructions of its own. Send `""` to clear them. Up to 4,000 characters. They apply from the next sync; documents already synced are not processed again. See [Custom Ingestion Instructions](/essentials/v2/connector-instructions). +- **`custom_instructions`** sets the connector-level ingestion instructions, used by every resource that has no instructions of its own. Send `""` to clear them. Up to 4,000 characters. They apply from the next sync; documents already synced are not processed again. See [Custom Ingestion Instructions](/essentials/v2/connectors#custom-ingestion-instructions). - **`sync_interval_seconds`** sets how often scheduled syncs run. The allowed range depends on the provider and comes back in the response as `min_sync_interval_seconds` and `max_sync_interval_seconds`. Values outside it are rejected, and `0` resets to the provider default. The next sync is rescheduled right away, so a shorter interval takes effect immediately. - **`credentials`** reconnects the connector, for example after a token expired or was revoked. Send the provider's full credential set, the same as on create. The connector keeps its id, resources and sync positions, and a reconnect-required state is cleared. @@ -55,5 +55,5 @@ curl -X PATCH 'https://api.hydradb.com/connectors/{id}' \ ## Related Resources - [Update Connector Resource](/api-reference/v2/endpoint/update-connector-resource): instructions for one resource -- [Custom Ingestion Instructions](/essentials/v2/connector-instructions): what instructions do and how they combine +- [Custom Ingestion Instructions](/essentials/v2/connectors#custom-ingestion-instructions): what instructions do and how they combine - [Get Connector](/api-reference/v2/endpoint/get-connector): read the stored settings back diff --git a/api-reference/v2/endpoint/webhooks-overview.mdx b/api-reference/v2/endpoint/webhooks-overview.mdx new file mode 100644 index 00000000..71b600ce --- /dev/null +++ b/api-reference/v2/endpoint/webhooks-overview.mdx @@ -0,0 +1,45 @@ +--- +title: "Webhooks: Overview" +description: "Every webhook endpoint, the order to call them in, and how deliveries are signed and retried." +--- + +A webhook is an HTTP `POST` that HydraDB sends to your endpoint when ingested content reaches a terminal indexing state, `completed` or `errored`, so you do not have to poll [Ingestion Status](/api-reference/v2/endpoint/source-status). One webhook is registered per workspace, for the `indexing.status_changed` event. The [Webhooks guide](/essentials/v2/webhooks) covers the payload, signature verification, and receiver examples. + +## Endpoints + +Set up the webhook: + +- [Register Webhook](/api-reference/v2/endpoint/register-webhook): `POST /webhooks/indexing` registers your endpoint, or replaces the current registration. Send `generate_signing_secret: true` to get a signing secret in the same call. +- [Send Test Delivery](/api-reference/v2/endpoint/test-webhook): `POST /webhooks/indexing/test` sends a synthetic, signed payload to confirm your endpoint and your signature check work. +- [Get Webhook](/api-reference/v2/endpoint/get-webhook): `GET /webhooks/indexing` reads the current registration and whether a signing secret is set. + +Track deliveries: + +- [List Deliveries](/api-reference/v2/endpoint/list-webhook-deliveries): `GET /webhooks/indexing/deliveries` lists delivery attempts, most recent first. Filter by `status` to find failures. +- [Get Delivery](/api-reference/v2/endpoint/get-webhook-delivery): `GET /webhooks/indexing/deliveries/{delivery_id}` reads one delivery, using the id from the `X-HydraDB-Delivery-ID` header. +- [Retry Delivery](/api-reference/v2/endpoint/retry-webhook-delivery): `POST /webhooks/indexing/deliveries/{delivery_id}/retry` queues a `failed` or `permanently_failed` delivery for another attempt. + +Change or remove it: + +- `POST /webhooks/indexing/signing-secret` generates a new signing secret, or stores one you supply. The secret is returned once and takes effect immediately, so rotate only once your receiver accepts the new one. +- `DELETE /webhooks/indexing/signing-secret` turns signing off. It is the only way to; editing the registration never clears the secret. +- [Delete Webhook](/api-reference/v2/endpoint/delete-webhook): `DELETE /webhooks/indexing` removes the registration and its signing secret. + +## Typical call sequence + +1. `POST /webhooks/indexing` with your HTTPS endpoint and `generate_signing_secret: true`. Store the secret it returns. +2. `POST /webhooks/indexing/test` to check your receiver accepts the delivery and verifies the signature. +3. Ingest content as usual. Each item that finishes indexing produces one delivery. +4. `GET /webhooks/indexing/deliveries?status=failed` to find deliveries your endpoint rejected, then retry them once the problem is fixed. + +## Delivery and retries + +Each delivery carries the event in its body and an `X-HydraDB-Delivery-ID` header; with signing on, it also carries an `X-HydraDB-Signature` header. A failed delivery is retried with exponential backoff until it succeeds. On HydraDB Cloud it is marked `permanently_failed` after 16 attempts, or 48 hours after the first failure, whichever comes first. A retry, automatic or manual, is signed with your current secret. Test deliveries do not appear in the delivery history. + +Your endpoint must be reachable over public HTTPS; localhost and private network addresses are rejected. + +## Related Resources + +- [Webhooks guide](/essentials/v2/webhooks): payload fields, signature verification, receiver examples, and secret rotation +- [Ingestion Status](/api-reference/v2/endpoint/source-status): polling instead of a webhook +- [Ingest Context](/api-reference/v2/endpoint/ingest-context): the ingestion that produces deliveries diff --git a/api-reference/v2/index.mdx b/api-reference/v2/index.mdx index 64c13b17..8ed50985 100644 --- a/api-reference/v2/index.mdx +++ b/api-reference/v2/index.mdx @@ -44,7 +44,7 @@ flowchart LR | [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB: document uploads, app sources, user memories, and lifecycle ops | | [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | Find relevant passages to include in your model prompt | | [Connectors](/api-reference/v2/endpoint/connectors-overview) | Connect, configure, sync, and manage app connectors such as Slack, GitHub, Google Drive, and Supabase | When you want app data synced into a database without writing ingest code | -| [Webhooks](/essentials/v2/webhooks) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | +| [Webhooks](/api-reference/v2/endpoint/webhooks-overview) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | ## Core concepts diff --git a/docs.json b/docs.json index 2dd8f792..e1e9f8c4 100644 --- a/docs.json +++ b/docs.json @@ -207,6 +207,7 @@ "group": "Webhooks", "public": true, "pages": [ + "api-reference/v2/endpoint/webhooks-overview", "api-reference/v2/endpoint/register-webhook", "api-reference/v2/endpoint/get-webhook", "api-reference/v2/endpoint/delete-webhook", From c6de5607a777c9cfde2506779fcf9ccce39d4e45 Mon Sep 17 00:00:00 2001 From: SohamRatnaparkhi Date: Fri, 9 Oct 2026 00:40:53 +0530 Subject: [PATCH 11/11] docs(v2): apply the review principles across the API reference - Code tabs: every connector and webhook endpoint now shows Python SDK, TypeScript SDK, and cURL, like the rest of the reference. Update Metadata Schema, Update Source Metadata, and an Ingest Context example used raw HTTP in their SDK tabs; they now call the SDK. Every new SDK call was run against a local mock server and type-checked against SDK 2.1.7, and sends the same request as its cURL. - Callouts: 40 routine Note, Info, Tip, and Warning boxes become plain text. Warnings remain only where data or access is at risk: replacing a connector resource, deleting a database, deleting a source that is still indexing, and an acl_warning on a connector resource. - Overview: drops the core concepts table, the endpoint groups table, the SDK setup code, and the status code table, each covered on its own page; the endpoint inventory is grouped by resource, each group linking to its overview. 2,202 words to 1,655. - Query: the tuning list that repeated the Query guide becomes a link; Behavior notes becomes Common mistakes. - Error Responses: notes that both SDKs already retry 408, 429, and 5xx twice with exponential backoff (Python also 409), and how to raise it. - Wording: no "material", no per-user or one-user framing of collections and memories; connector examples write to one company_docs collection; no arrows in prose. Refs PRO-2457 Co-Authored-By: Claude Opus 5.5 (1M context) Signed-off-by: SohamRatnaparkhi --- .../v2/endpoint/add-connector-resource.mdx | 25 +++- .../v2/endpoint/configure-connector.mdx | 50 +++++++- .../v2/endpoint/connector-resources.mdx | 8 ++ .../v2/endpoint/connectors-overview.mdx | 4 +- .../v2/endpoint/create-connector.mdx | 30 ++++- api-reference/v2/endpoint/create-tenant.mdx | 10 +- .../v2/endpoint/delete-connector-resource.mdx | 8 ++ .../v2/endpoint/delete-connector.mdx | 8 ++ api-reference/v2/endpoint/delete-source.mdx | 22 ++-- api-reference/v2/endpoint/delete-tenant.mdx | 2 +- api-reference/v2/endpoint/delete-webhook.mdx | 18 +++ .../endpoint/discover-connector-resources.mdx | 13 ++ .../v2/endpoint/get-connector-status.mdx | 10 ++ api-reference/v2/endpoint/get-connector.mdx | 8 ++ .../v2/endpoint/get-webhook-delivery.mdx | 18 +++ api-reference/v2/endpoint/get-webhook.mdx | 22 +++- api-reference/v2/endpoint/ingest-context.mdx | 61 ++-------- .../v2/endpoint/list-connector-providers.mdx | 16 ++- api-reference/v2/endpoint/list-connectors.mdx | 8 ++ api-reference/v2/endpoint/list-documents.mdx | 4 +- api-reference/v2/endpoint/list-tenants.mdx | 2 +- .../v2/endpoint/list-webhook-deliveries.mdx | 18 +++ api-reference/v2/endpoint/pause-connector.mdx | 8 ++ api-reference/v2/endpoint/query.mdx | 26 ++-- .../v2/endpoint/register-webhook.mdx | 38 +++++- .../v2/endpoint/resume-connector.mdx | 8 ++ .../v2/endpoint/retry-webhook-delivery.mdx | 18 +++ .../v2/endpoint/source-relations.mdx | 6 +- api-reference/v2/endpoint/source-status.mdx | 8 +- .../v2/endpoint/sources-overview.mdx | 10 +- api-reference/v2/endpoint/subgraph.mdx | 4 +- api-reference/v2/endpoint/submit-feedback.mdx | 2 - api-reference/v2/endpoint/sync-connector.mdx | 8 ++ api-reference/v2/endpoint/tenant-stats.mdx | 6 +- .../v2/endpoint/tenants-overview.mdx | 2 +- api-reference/v2/endpoint/test-webhook.mdx | 24 +++- .../v2/endpoint/update-connector-resource.mdx | 20 +++ .../v2/endpoint/update-connector.mdx | 21 ++++ .../v2/endpoint/update-metadata-schema.mdx | 64 ++++------ .../v2/endpoint/update-source-metadata.mdx | 78 ++++-------- api-reference/v2/error-responses.mdx | 14 +-- api-reference/v2/index.mdx | 115 ++++++------------ api-reference/v2/sdks.mdx | 8 -- 43 files changed, 514 insertions(+), 339 deletions(-) diff --git a/api-reference/v2/endpoint/add-connector-resource.mdx b/api-reference/v2/endpoint/add-connector-resource.mdx index dbafa6a4..f6c80fdd 100644 --- a/api-reference/v2/endpoint/add-connector-resource.mdx +++ b/api-reference/v2/endpoint/add-connector-resource.mdx @@ -14,6 +14,26 @@ Adds a single resource to a connector without going through [Configure Connector +```python Python SDK +client.connectors.create_resource( + "{connector_id}", + resource_id="C0123456789", + resource_type="channel", + display_name="general", + filters={"lookback_days": 30}, +) +``` + +```typescript TypeScript SDK +await client.connectors.createResource({ + id: "{connector_id}", + resourceId: "C0123456789", + resourceType: "channel", + displayName: "general", + filters: { lookback_days: 30 }, +}); +``` + ```bash cURL curl -X POST 'https://api.hydradb.com/connectors/{id}/resources' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ @@ -23,7 +43,6 @@ curl -X POST 'https://api.hydradb.com/connectors/{id}/resources' \ "resource_id": "C0123456789", "resource_type": "channel", "display_name": "general", - "collection_override": "all-hands", "filters": { "lookback_days": 30 } }' ``` @@ -41,9 +60,9 @@ curl -X POST 'https://api.hydradb.com/connectors/{id}/resources' \ "status": "active", "provider_cursor": "", "database_override": "", - "collection_override": "all-hands", + "collection_override": "", "tenant_id_override": "", - "sub_tenant_id_override": "all-hands", + "sub_tenant_id_override": "", "provider_metadata": null, "filters": { "lookback_days": 30 diff --git a/api-reference/v2/endpoint/configure-connector.mdx b/api-reference/v2/endpoint/configure-connector.mdx index bde80070..8cd24614 100644 --- a/api-reference/v2/endpoint/configure-connector.mdx +++ b/api-reference/v2/endpoint/configure-connector.mdx @@ -18,12 +18,54 @@ You can call configure again at any time to add resources or change their settin - The resource's sync position is kept, so it does not sync again from the beginning. - `custom_instructions` and `acl` are kept when you omit them. To clear instructions, send `""` with [Update Connector Resource](/api-reference/v2/endpoint/update-connector-resource); configure cannot clear them. - - To change only a resource's instructions or access rule, use [Update Connector Resource](/api-reference/v2/endpoint/update-connector-resource) instead. It changes just the fields you send. - +To change only a resource's instructions or access rule, use [Update Connector Resource](/api-reference/v2/endpoint/update-connector-resource) instead. It changes just the fields you send. +```python Python SDK +result = client.connectors.configure( + "{connector_id}", + lookback_days=30, + resources=[ + { + "resource_id": "C0123456789", + "resource_type": "channel", + "name": "general", + "metadata": {"department": "all-hands"}, + "additional_metadata": {"internal_label": "general-slack"}, + }, + { + "resource_id": "C0987654321", + "resource_type": "channel", + "name": "incidents", + "custom_instructions": "These threads are incident retros. Extract root cause, impact and owner.", + }, + ], +) +``` + +```typescript TypeScript SDK +const result = await client.connectors.configure({ + id: "{connector_id}", + lookbackDays: 30, + resources: [ + { + resourceId: "C0123456789", + resourceType: "channel", + name: "general", + metadata: { department: "all-hands" }, + additionalMetadata: { internal_label: "general-slack" }, + }, + { + resourceId: "C0987654321", + resourceType: "channel", + name: "incidents", + customInstructions: "These threads are incident retros. Extract root cause, impact and owner.", + }, + ], +}); +``` + ```bash cURL curl -X POST 'https://api.hydradb.com/connectors/{id}/configure' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ @@ -36,7 +78,6 @@ curl -X POST 'https://api.hydradb.com/connectors/{id}/configure' \ "resource_id": "C0123456789", "resource_type": "channel", "name": "general", - "collection": "all-hands", "metadata": { "department": "all-hands" }, "additional_metadata": { "internal_label": "general-slack" } }, @@ -44,7 +85,6 @@ curl -X POST 'https://api.hydradb.com/connectors/{id}/configure' \ "resource_id": "C0987654321", "resource_type": "channel", "name": "incidents", - "collection": "engineering", "custom_instructions": "These threads are incident retros. Extract root cause, impact and owner." } ] diff --git a/api-reference/v2/endpoint/connector-resources.mdx b/api-reference/v2/endpoint/connector-resources.mdx index 21db3c57..84d6e198 100644 --- a/api-reference/v2/endpoint/connector-resources.mdx +++ b/api-reference/v2/endpoint/connector-resources.mdx @@ -12,6 +12,14 @@ The `resource_id` values returned here are the ones to use in `PATCH` and `DELET +```python Python SDK +resources = client.connectors.list_resources("{connector_id}") +``` + +```typescript TypeScript SDK +const resources = await client.connectors.listResources({ id: "{connector_id}" }); +``` + ```bash cURL curl 'https://api.hydradb.com/connectors/{id}/resources' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/connectors-overview.mdx b/api-reference/v2/endpoint/connectors-overview.mdx index 8f86a3f6..763e6d01 100644 --- a/api-reference/v2/endpoint/connectors-overview.mdx +++ b/api-reference/v2/endpoint/connectors-overview.mdx @@ -45,9 +45,7 @@ After that, syncs run every `sync_interval_seconds` (one hour by default). Call ## Which endpoint changes what - - To change a resource that is already configured, use `PATCH /connectors/{id}/resources/{resource_id}`. `POST /connectors/{id}/resources` on an existing `resource_id` replaces the whole resource: every field you leave out is cleared, and the resource syncs again from the beginning. - +To change a resource that is already configured, use `PATCH /connectors/{id}/resources/{resource_id}`. `POST /connectors/{id}/resources` on an existing `resource_id` replaces the whole resource: every field you leave out is cleared, and the resource syncs again from the beginning. - Instructions or access rule on one resource: `PATCH /connectors/{id}/resources/{resource_id}`. Only the fields you send change. - Instructions, sync interval or credentials for the whole connector: `PATCH /connectors/{id}`. Only the fields you send change. diff --git a/api-reference/v2/endpoint/create-connector.mdx b/api-reference/v2/endpoint/create-connector.mdx index a0c503dd..dd7002c2 100644 --- a/api-reference/v2/endpoint/create-connector.mdx +++ b/api-reference/v2/endpoint/create-connector.mdx @@ -15,6 +15,30 @@ Optional settings you can also pass here, and change later with [Update Connecto +```python Python SDK +connector = client.connectors.create( + provider="slack", + name="acme-engineering", + database="acme_corp", + collection="company_docs", + provider_account_scope="T12345ACME", + credentials={"access_token": "xoxp-..."}, +) +connector_id = connector.connector_id +``` + +```typescript TypeScript SDK +const connector = await client.connectors.create({ + provider: "slack", + name: "acme-engineering", + database: "acme_corp", + collection: "company_docs", + providerAccountScope: "T12345ACME", + credentials: { access_token: "xoxp-..." }, +}); +const connectorId = connector.connectorId; +``` + ```bash cURL curl -X POST 'https://api.hydradb.com/connectors' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ @@ -24,7 +48,7 @@ curl -X POST 'https://api.hydradb.com/connectors' \ "provider": "slack", "name": "acme-engineering", "database": "acme_corp", - "collection": "engineering", + "collection": "company_docs", "provider_account_scope": "T12345ACME", "credentials": { "access_token": "xoxp-..." @@ -42,9 +66,9 @@ curl -X POST 'https://api.hydradb.com/connectors' \ "provider": "slack", "name": "acme-engineering", "database": "acme_corp", - "collection": "engineering", + "collection": "company_docs", "tenant_id": "acme_corp", - "sub_tenant_id": "engineering", + "sub_tenant_id": "company_docs", "provider_account_scope": "T12345ACME", "lifecycle": "pending_setup", "sync_status": "idle", diff --git a/api-reference/v2/endpoint/create-tenant.mdx b/api-reference/v2/endpoint/create-tenant.mdx index 84fae23d..c2c952bc 100644 --- a/api-reference/v2/endpoint/create-tenant.mdx +++ b/api-reference/v2/endpoint/create-tenant.mdx @@ -76,10 +76,6 @@ curl -X POST 'https://api.hydradb.com/databases' \ ## Request body - -`database` and `collection` are the current field names (formerly `tenant_id` and `sub_tenant_id`). The old names remain accepted as deprecated aliases for full backward compatibility. - - | Name | Description | | --- | --- | | | Account-scoped database identifier. Use a stable ID up to 255 characters of lowercase letters, digits, `-`, and `_`; anything else, including uppercase or spaces, returns `400`. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | @@ -136,9 +132,7 @@ Always check if a database is ready before using it. Use [Database Status](/api- ## Defining metadata schema - - Schema field names are **immutable** after database creation. You can add per-document free-form metadata fields at ingestion time, and add new database-level fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but updates are additive only: no delete, rename, or type change. Dense and sparse metadata lanes (`enable_dense_embedding`, `enable_sparse_embedding`) can only be declared here, at creation. Plan your schema before you create the database. - +Schema field names are **immutable** after database creation. You can add per-document free-form metadata fields at ingestion time, and add new database-level fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema), but updates are additive only: no delete, rename, or type change. Dense and sparse metadata lanes (`enable_dense_embedding`, `enable_sparse_embedding`) can only be declared here, at creation. Plan your schema before you create the database. You can define a custom schema at database creation to declare the `metadata` fields you filter on, and to enable semantic/BM25 search over metadata text fields (`enable_dense_embedding` / `enable_sparse_embedding`). Each dense or sparse flag adds one vector field, so a field with both uses two; a database can have at most 6. Going over returns `400`, as does declaring an `ARRAY` field. @@ -160,5 +154,5 @@ Common codes: `400 INVALID_INPUT` (missing or invalid `database`, or an invalid - **Next:** [Ingest Context](/api-reference/v2/endpoint/ingest-context): start ingesting data once status is ready - **Related:** [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema): add metadata schema fields later - **Related:** [Delete Database](/api-reference/v2/endpoint/delete-tenant): teardown -- **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) +- **Read more:** [Multi-tenancy](/essentials/v2/multi-tenant) - **Read more:** [Usage: Metadata](/essentials/v2/metadata) diff --git a/api-reference/v2/endpoint/delete-connector-resource.mdx b/api-reference/v2/endpoint/delete-connector-resource.mdx index 9ea23415..31acacac 100644 --- a/api-reference/v2/endpoint/delete-connector-resource.mdx +++ b/api-reference/v2/endpoint/delete-connector-resource.mdx @@ -10,6 +10,14 @@ Removes a resource from the connector so it is no longer synced. Objects already +```python Python SDK +client.connectors.delete_resource("{connector_id}", "{resource_id}") +``` + +```typescript TypeScript SDK +await client.connectors.deleteResource({ id: "{connector_id}", resourceId: "{resource_id}" }); +``` + ```bash cURL curl -X DELETE 'https://api.hydradb.com/connectors/{id}/resources/{resource_id}' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/delete-connector.mdx b/api-reference/v2/endpoint/delete-connector.mdx index 834e88ed..7002bd60 100644 --- a/api-reference/v2/endpoint/delete-connector.mdx +++ b/api-reference/v2/endpoint/delete-connector.mdx @@ -10,6 +10,14 @@ To stop syncing for a while without losing the configuration, [pause](/api-refer +```python Python SDK +client.connectors.delete("{connector_id}") +``` + +```typescript TypeScript SDK +await client.connectors.delete({ id: "{connector_id}" }); +``` + ```bash cURL curl -X DELETE 'https://api.hydradb.com/connectors/{id}' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/delete-source.mdx b/api-reference/v2/endpoint/delete-source.mdx index c604d1f8..b96fdadf 100644 --- a/api-reference/v2/endpoint/delete-source.mdx +++ b/api-reference/v2/endpoint/delete-source.mdx @@ -257,18 +257,16 @@ The header always wins. Without it, the server default applies. | `X-HydraDB-Delete-Status: strict` | Honest `404` / `409` / `500`. | | `X-HydraDB-Delete-Status: legacy` | `200` for every outcome, whatever the server default. | - - **The default is expected to become `strict` in a future release.** Adopting - `strict` now means that change is a no-op for you. - - When it happens, `X-HydraDB-Delete-Status: legacy` keeps the unconditional - `200` for any integration that is not ready. Both header values are supported - and neither has a removal date. If that ever changes, we will announce it. - - If your integration checks `response.ok` or `status == 200` today, it is - treating blocked deletes as successful. That is the failure strict mode - surfaces. - +**The default is expected to become `strict` in a future release.** Adopting +`strict` now means that change is a no-op for you. + +When it happens, `X-HydraDB-Delete-Status: legacy` keeps the unconditional +`200` for any integration that is not ready. Both header values are supported +and neither has a removal date. If that ever changes, we will announce it. + +If your integration checks `response.ok` or `status == 200` today, it is +treating blocked deletes as successful. That is the failure strict mode +surfaces. ## Notes diff --git a/api-reference/v2/endpoint/delete-tenant.mdx b/api-reference/v2/endpoint/delete-tenant.mdx index b318f70b..7fa06c48 100644 --- a/api-reference/v2/endpoint/delete-tenant.mdx +++ b/api-reference/v2/endpoint/delete-tenant.mdx @@ -91,5 +91,5 @@ Common codes: `400 INVALID_INPUT` (missing `database`), `401 UNAUTHORIZED`, `404 - **Before this:** [List Databases](/api-reference/v2/endpoint/list-tenants): find the database ID - **Alternative:** [Delete Collection](/api-reference/v2/endpoint/delete-collection): remove one collection without deleting the whole database - **Alternative:** [Delete Context](/api-reference/v2/endpoint/delete-source): remove specific knowledge or memories without deleting the whole database -- **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) +- **Read more:** [Multi-tenancy](/essentials/v2/multi-tenant)
    diff --git a/api-reference/v2/endpoint/delete-webhook.mdx b/api-reference/v2/endpoint/delete-webhook.mdx index 2efcfae9..b204081c 100644 --- a/api-reference/v2/endpoint/delete-webhook.mdx +++ b/api-reference/v2/endpoint/delete-webhook.mdx @@ -7,3 +7,21 @@ openapi: "api-reference/v2/openapi.json DELETE /webhooks/indexing" Removes the webhook registration. HydraDB stops sending deliveries, and the stored signing secret is discarded along with the registration. To stop signing deliveries while keeping the webhook itself, call `DELETE /webhooks/indexing/signing-secret` instead. See [Manage the signing secret](/essentials/v2/webhooks#manage-the-signing-secret). + + + +```python Python SDK +client.webhooks.delete() +``` + +```typescript TypeScript SDK +await client.webhooks.delete(); +``` + +```bash cURL +curl -X DELETE 'https://api.hydradb.com/webhooks/indexing' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" +``` + + diff --git a/api-reference/v2/endpoint/discover-connector-resources.mdx b/api-reference/v2/endpoint/discover-connector-resources.mdx index ab259934..6f619ab0 100644 --- a/api-reference/v2/endpoint/discover-connector-resources.mdx +++ b/api-reference/v2/endpoint/discover-connector-resources.mdx @@ -14,6 +14,19 @@ For providers with many resources, page through the list with `limit` (1 to 100) +```python Python SDK +found = client.connectors.discover("{connector_id}") +for resource in found.resources: + print(resource.id, resource.name, resource.resource_type) +``` + +```typescript TypeScript SDK +const found = await client.connectors.discover({ id: "{connector_id}" }); +for (const resource of found.resources ?? []) { + console.log(resource.id, resource.name, resource.resourceType); +} +``` + ```bash cURL curl 'https://api.hydradb.com/connectors/{id}/discover' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/get-connector-status.mdx b/api-reference/v2/endpoint/get-connector-status.mdx index af1fd2d0..07e8d4f7 100644 --- a/api-reference/v2/endpoint/get-connector-status.mdx +++ b/api-reference/v2/endpoint/get-connector-status.mdx @@ -8,6 +8,16 @@ Answers "is this connector working?" in one call. It returns an overall `status` +```python Python SDK +status = client.connectors.status("{connector_id}") +print(status.status, status.lifecycle) +``` + +```typescript TypeScript SDK +const status = await client.connectors.status({ id: "{connector_id}" }); +console.log(status.status, status.lifecycle); +``` + ```bash cURL curl 'https://api.hydradb.com/connectors/{id}/status' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/get-connector.mdx b/api-reference/v2/endpoint/get-connector.mdx index 20f107cf..6a7386de 100644 --- a/api-reference/v2/endpoint/get-connector.mdx +++ b/api-reference/v2/endpoint/get-connector.mdx @@ -10,6 +10,14 @@ Returns one connector: where it syncs to, its connector-level `custom_instructio +```python Python SDK +connector = client.connectors.get("{connector_id}") +``` + +```typescript TypeScript SDK +const connector = await client.connectors.get({ id: "{connector_id}" }); +``` + ```bash cURL curl 'https://api.hydradb.com/connectors/{id}' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/get-webhook-delivery.mdx b/api-reference/v2/endpoint/get-webhook-delivery.mdx index 16f39248..cd854363 100644 --- a/api-reference/v2/endpoint/get-webhook-delivery.mdx +++ b/api-reference/v2/endpoint/get-webhook-delivery.mdx @@ -7,3 +7,21 @@ openapi: "api-reference/v2/openapi.json GET /webhooks/indexing/deliveries/{deliv Returns one delivery record, including its current status, attempt count, and the last error code and message when an attempt failed. The `delivery_id` is the value sent in the `X-HydraDB-Delivery-ID` header of the delivery itself, so you can look up any request your endpoint received. See [Webhooks](/essentials/v2/webhooks). + + + +```python Python SDK +delivery = client.webhooks.get_delivery("{delivery_id}") +``` + +```typescript TypeScript SDK +const delivery = await client.webhooks.getDelivery({ deliveryId: "{delivery_id}" }); +``` + +```bash cURL +curl 'https://api.hydradb.com/webhooks/indexing/deliveries/{delivery_id}' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" +``` + + diff --git a/api-reference/v2/endpoint/get-webhook.mdx b/api-reference/v2/endpoint/get-webhook.mdx index 2e5cb94b..6745bd0e 100644 --- a/api-reference/v2/endpoint/get-webhook.mdx +++ b/api-reference/v2/endpoint/get-webhook.mdx @@ -6,8 +6,24 @@ openapi: "api-reference/v2/openapi.json GET /webhooks/indexing" Returns the registered endpoint, the subscribed event types, and whether a signing secret is configured. - - The signing secret itself is never returned. `signing_secret_configured` tells you only whether one is set. If you have lost your secret, rotate it with `POST /webhooks/indexing/signing-secret` rather than trying to read it back. - +The signing secret itself is never returned. `signing_secret_configured` tells you only whether one is set. If you have lost your secret, rotate it with `POST /webhooks/indexing/signing-secret` rather than trying to read it back. When no webhook is registered, `registered` is `false` and `url` is `null`. See [Webhooks](/essentials/v2/webhooks). + + + +```python Python SDK +webhook = client.webhooks.get() +``` + +```typescript TypeScript SDK +const webhook = await client.webhooks.get(); +``` + +```bash cURL +curl 'https://api.hydradb.com/webhooks/indexing' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" +``` + + diff --git a/api-reference/v2/endpoint/ingest-context.mdx b/api-reference/v2/endpoint/ingest-context.mdx index de42aef5..a98badf9 100644 --- a/api-reference/v2/endpoint/ingest-context.mdx +++ b/api-reference/v2/endpoint/ingest-context.mdx @@ -13,7 +13,7 @@ When context is of `type=knowledge`: When context is of `type=memory`: -Use `memories` for per-user content, scoped with `collection`. Set `infer: true` to let HydraDB extract preferences from raw signals, or `infer: false` to store the text verbatim. Read more about [ingesting memories](/essentials/v2/memories). +Use `memories` for preferences, decisions, and past conversations, saved in the `collection` you will query. Set `infer: true` to let HydraDB extract preferences from raw signals, or `infer: false` to store the text verbatim. Read more about [ingesting memories](/essentials/v2/memories). @@ -273,39 +273,16 @@ const response = await client.context.ingest({ }); ``` -```python API -import io -import json -import requests - -text = """Q4 planning notes - -- Launch checklist is owned by Priya. -- Legal review is due by Friday. -""" - -# Create a file-like object in memory. No local .txt file is required. -txt_file = io.BytesIO(text.encode("utf-8")) - -response = requests.post( - "https://api.hydradb.com/context/ingest", - headers={ - "Authorization": "Bearer ", - "API-Version": "2", - }, - data={ - "type": "knowledge", - "database": "acme_corp", - "collection": "team_docs", - "document_metadata": json.dumps([ - {"id": "meeting_notes_q4", "metadata": {"department": "product"}} - ]), - }, - files={ - "documents": ("meeting-notes.txt", txt_file, "text/plain"), - }, -) -response.raise_for_status() +```bash cURL +printf 'Q4 planning notes\n\n- Launch checklist is owned by Priya.\n- Legal review is due by Friday.\n' | \ +curl -X POST 'https://api.hydradb.com/context/ingest' \ + -H "Authorization: Bearer " \ + -H "API-Version: 2" \ + -F "type=knowledge" \ + -F "database=acme_corp" \ + -F "collection=team_docs" \ + -F "documents=@-;filename=meeting-notes.txt;type=text/plain" \ + -F 'document_metadata=[{"id":"meeting_notes_q4","metadata":{"department":"product"}}]' ``` @@ -323,11 +300,9 @@ response.raise_for_status() | | Map of source id to your own entities + relations. Replaces LLM graph extraction for each keyed source. Works for `type=knowledge` (key = a `document_metadata` id or `app_knowledge` item id) and `type=memory` (key = a memory `id`). See [Bring Your Own Graph](/essentials/v2/bring-your-own-graph#using-byog-without-cypher) and the shape below. | | | Memory items, **memory only**. Required and non-empty when `type=memory`. Use plural `memories` for the form field, even though `type` is singular `memory`. See the `memories` item shape below. | - 1. **`id` must not contain a comma (`,`):** The comma is reserved as the id separator on [Ingestion Status](/api-reference/v2/endpoint/source-status) (`GET /context/status?ids=a,b`), so an `id` containing a comma cannot be looked up unambiguously. This applies to every `id` you supply: `document_metadata`, `app_knowledge`, and `memories` items. Ingesting an item whose `id` contains a comma is rejected with a `400`. 2. **`202 Accepted` means queued, not indexed:** Ingestion is asynchronous. A successful response only confirms your sources were accepted, not that they are ready to query. Before querying, poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with the returned IDs until each source reaches `graph_creation` (searchable), `completed`, or `errored`. Alternatively, register a webhook for `indexing.status_changed` events (see [Webhooks](/essentials/v2/webhooks)). - ## Supported file formats @@ -367,9 +342,7 @@ These are rejected the moment you upload them, before anything is queued. You ge | `.epub` `.wpd` | Convert to PDF or DOCX | | `.zip` | Upload the files individually | - `.heic` is the format iPhones use for photos by default, so it is the one people hit most often without realizing. On iOS you can change this under **Settings > Camera > Formats > Most Compatible**, which makes the camera save JPEGs instead. Or export the photo as JPEG or PDF before uploading. - ### How we decide the format @@ -377,9 +350,7 @@ The file extension is what counts. `report.pdf` is treated as a PDF because of t If the filename has no extension at all, the `Content-Type` you send with that part is used instead, so a file with no extension in its name, sent as `application/pdf`, is accepted. - Renaming a file does not convert it. A `.heic` photo renamed to `photo.pdf` passes this check, because the check reads the name, then fails later during parsing and comes back as `errored` on [`GET /context/status`](/api-reference/v2/endpoint/source-status). Upload files under their real extension. - ### Size limit @@ -410,13 +381,9 @@ The request returns `202` as usual. The rejected file comes back with `status: " } ``` - `results`, `success_count` and `failed_count` sit inside `data`, not at the top level, like every other v2 response. The outer `success: true` means the request was accepted; `data.success` is what tells you whether every file in it was queued. - - **Do not poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) for a file rejected with `E1002`.** The file never entered the pipeline, so it has no status record, and looking it up returns `FILE_NOT_FOUND` rather than the format error you were given. The upload response is the only place `E1002` appears. Read `error_code` on each item in `data.results` and act on it there. - ## Common use-cases and their configurations @@ -489,11 +456,10 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in Size limits are measured on the compact JSON of the whole object, in UTF-8 bytes, and apply the same way to `app_knowledge` and `memories` items. - Those seven, plus the legacy aliases `source_id` and `file_id` (for `id`) and `document_metadata` (for `additional_metadata`), are the only keys accepted. Anything else (including `title`, `type`, `url` and `timestamp`) is rejected with a `400` naming the unsupported key, rather than being silently dropped. In particular, a document's **title is derived, not settable**. It defaults to the uploaded filename and is returned as `source_title` on query results and `title` on `/context/list`. It cannot be overridden at ingest, and [`PATCH /context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata) only merges `additional_metadata` and `database_metadata`. Use `additional_metadata` for your own display fields, or ingest through `app_knowledge`, whose items carry an explicit `title`. - + @@ -798,9 +764,8 @@ Per-document metadata (`id`, `metadata`, `additional_metadata`, `relations`, `in | | Optional sentence supporting the edge. At most 2,000 chars. | | | Optional timing info (e.g. "since 2021", "in Q3"). | - **Caps per graph:** at most 5,000 entities, 10,000 relations, and 500 relations per entity; over-cap returns `400`. Graphs survive re-ingest: a re-upload or connector re-sync re-applies the stored graph. - + ## Notes diff --git a/api-reference/v2/endpoint/list-connector-providers.mdx b/api-reference/v2/endpoint/list-connector-providers.mdx index bf5ee85c..a563a680 100644 --- a/api-reference/v2/endpoint/list-connector-providers.mdx +++ b/api-reference/v2/endpoint/list-connector-providers.mdx @@ -13,13 +13,23 @@ Use a catalog entry's `provider` value as `provider` in [Create Connector](/api- -```bash List every provider +```python Python SDK +catalog = client.list_providers() + +affinity = client.list_providers(id="affinity") +``` + +```typescript TypeScript SDK +const catalog = await client.listProviders(); + +const affinity = await client.listProviders({ id: "affinity" }); +``` + +```bash cURL curl 'https://api.hydradb.com/connectors/providers' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ -H "API-Version: 2" -``` -```bash Describe one provider curl 'https://api.hydradb.com/connectors/providers?id=affinity' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ -H "API-Version: 2" diff --git a/api-reference/v2/endpoint/list-connectors.mdx b/api-reference/v2/endpoint/list-connectors.mdx index 6dac5d13..fdb12e49 100644 --- a/api-reference/v2/endpoint/list-connectors.mdx +++ b/api-reference/v2/endpoint/list-connectors.mdx @@ -10,6 +10,14 @@ Each entry has the same fields as [Get Connector](/api-reference/v2/endpoint/get +```python Python SDK +connectors = client.connectors.list() +``` + +```typescript TypeScript SDK +const connectors = await client.connectors.list(); +``` + ```bash cURL curl 'https://api.hydradb.com/connectors' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/list-documents.mdx b/api-reference/v2/endpoint/list-documents.mdx index 7d65063d..d3db3ad5 100644 --- a/api-reference/v2/endpoint/list-documents.mdx +++ b/api-reference/v2/endpoint/list-documents.mdx @@ -105,9 +105,7 @@ When you don't need every field on every row, pass `include_fields` to keep resp Allowed values are `title`, `type`, `description`, `note`, `timestamp`, `metadata`, `additional_metadata`, and `relations`, plus the legacy names `tenant_metadata` and `document_metadata`. Omit or pass `null` to return everything. - - **Projectable vs. fetchable fields:** `content`, `url`, and `attachments` are **not** valid `include_fields` values: they are stripped from list responses, and requesting one returns `400`. Fetch them per-source via [Inspect Context](/api-reference/v2/endpoint/fetch-content). - +**Projectable vs. fetchable fields:** `content`, `url`, and `attachments` are **not** valid `include_fields` values: they are stripped from list responses, and requesting one returns `400`. Fetch them per-source via [Inspect Context](/api-reference/v2/endpoint/fetch-content). diff --git a/api-reference/v2/endpoint/list-tenants.mdx b/api-reference/v2/endpoint/list-tenants.mdx index dce7f5de..d6637aac 100644 --- a/api-reference/v2/endpoint/list-tenants.mdx +++ b/api-reference/v2/endpoint/list-tenants.mdx @@ -93,5 +93,5 @@ A database in `data.failed_databases` carries diagnostic entries (see the **Prov - **Inspect:** [Database Status](/api-reference/v2/endpoint/tenant-status) - **Inspect:** [Database Stats](/api-reference/v2/endpoint/tenant-stats) - **Delete:** [Delete Database](/api-reference/v2/endpoint/delete-tenant) - - **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) + - **Read more:** [Multi-tenancy](/essentials/v2/multi-tenant) \ No newline at end of file diff --git a/api-reference/v2/endpoint/list-webhook-deliveries.mdx b/api-reference/v2/endpoint/list-webhook-deliveries.mdx index a4acb509..114f9201 100644 --- a/api-reference/v2/endpoint/list-webhook-deliveries.mdx +++ b/api-reference/v2/endpoint/list-webhook-deliveries.mdx @@ -19,3 +19,21 @@ Filter by `status` to isolate failures, and page through results with `limit` an `limit` is 1 to 100 (default 20). To get the next page, pass the response's `next_cursor` as `cursor`. `next_cursor` is `null` on the last page. A `limit` outside that range returns `422`, and an unknown `status` returns `400`. See [Webhooks](/essentials/v2/webhooks) for retry behaviour. + + + +```python Python SDK +deliveries = client.webhooks.list_deliveries(status="failed", limit=20) +``` + +```typescript TypeScript SDK +const deliveries = await client.webhooks.listDeliveries({ status: "failed", limit: 20 }); +``` + +```bash cURL +curl 'https://api.hydradb.com/webhooks/indexing/deliveries?status=failed&limit=20' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" +``` + + diff --git a/api-reference/v2/endpoint/pause-connector.mdx b/api-reference/v2/endpoint/pause-connector.mdx index 620b9e32..d83f9d00 100644 --- a/api-reference/v2/endpoint/pause-connector.mdx +++ b/api-reference/v2/endpoint/pause-connector.mdx @@ -10,6 +10,14 @@ While paused, `lifecycle` reads `paused` and [Sync Connector](/api-reference/v2/ +```python Python SDK +client.connectors.pause("{connector_id}") +``` + +```typescript TypeScript SDK +await client.connectors.pause({ id: "{connector_id}" }); +``` + ```bash cURL curl -X POST 'https://api.hydradb.com/connectors/{id}/pause' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/query.mdx b/api-reference/v2/endpoint/query.mdx index d428fdfd..20dddc20 100644 --- a/api-reference/v2/endpoint/query.mdx +++ b/api-reference/v2/endpoint/query.mdx @@ -424,7 +424,7 @@ result = client.query( | Name | Description | | --- | --- | | | Owning database. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | -| | Single collection scope. Required for per-user memory queries. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=default collection) | +| | Single collection scope. Send the collection the content was saved in. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). (default=default collection) | | | Multi-collection scope. Send a list of collection IDs for equal weighting, or an object mapping collection ID to a positive relative weight (at most one decimal place, e.g. `{"finance": 1.5, "legal": 0.8}`) to bias ranking. Up to 100 collections. Do not combine with `collection`/`sub_tenant_id`. Formerly `sub_tenant_ids`; the `sub_tenant_ids` alias is still accepted (deprecated since 2.0.1). | | | Query terms or natural-language question. Cannot be empty. | | | What to query. `"all"` runs knowledge and memory in parallel over the same scope and merges by `relevancy_score`. (default=`"knowledge"`) | @@ -440,14 +440,7 @@ result = client.query( | | Request-time hint to guide retrieval (e.g., "user is on the billing page"). This is different from the response `additional_context` map. (default=`null`) | | | Deterministic narrowing before ranking. See [Filters](#decision-matrix). Each list holds at most 500 values, and the whole object is capped at 64 KiB of compact JSON, measured after operator objects are reduced to their values; over either returns `400`. (default=`null`) | - -**Tuning heuristics:** -
      -
    • alpha: start at 0.8. Lower toward 0.3 to 0.5 when the query contains literal tokens (error codes, SKUs, product names). Raise toward 0.9 for conceptual questions.
    • -
    • recency_bias: set to 0 for static reference material. Set 0.2 to 0.4 for mixed content, 0.6 to 0.8 for changelogs, news, or status updates.
    • -
    • max_results: start at 10. Drop to 5 for tight context windows; raise to 20 if you rerank downstream.
    • -
    -
    +To tune `alpha`, `recency_bias`, and `max_results` for your content, see the [Query guide](/essentials/v2/query). ### Decision matrix @@ -520,13 +513,9 @@ result = client.query( Operators apply to `metadata` (top-level keys) only. Inside `additional_metadata`, use a bare scalar for an exact match or a bare array to match any listed value. - - An operator used inside `additional_metadata` is **not** rejected. It is read as an exact-match filter against a stored object, so on a normal field it matches nothing and the request returns `200` with an empty result rather than an error. - + An operator used inside `additional_metadata` is **not** rejected. It is read as an exact-match filter against a stored object, so on a normal field it matches nothing and the request returns `200` with an empty result rather than an error. - - The bare forms still work and are unchanged, but are **deprecated** in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as `equals`, a bare array as `contains_any`, and a bare single-element array as `contains`, so `{"emails": "a@x"}` and `{"emails": ["a@x"]}` differ by one character and return different results. - + The bare forms still work and are unchanged, but are **deprecated** in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as `equals`, a bare array as `contains_any`, and a bare single-element array as `contains`, so `{"emails": "a@x"}` and `{"emails": ["a@x"]}` differ by one character and return different results. `contains`, `contains_any` and lists are supported on `VARCHAR` fields only: any of them passed for a declared field of another type is rejected with `400 VALIDATION_ERROR`. `equals` works on every declared type, so `{"priority": {"equals": 7}}` is valid on an `INT64` field. @@ -714,16 +703,15 @@ result = client.query( A zero-result query returns empty arrays/maps rather than an error, as shown in the **Zero results** tab. -## Behavior notes + + +## Common mistakes - -**Important Considerations & Common Mistakes** - **`query_forceful_relations` requires `mode` to resolve to `"thinking"`:** In `fast` mode the flag is silently skipped. The server does not error or warn; your `additional_context` will simply be empty. Under `mode: "auto"` this depends on that request's routing decision, not on what you asked for. - **Relation `timestamp` is a Unix epoch float here:** In the `graph_context` slice returned by `/query` (and in the passthrough relations returned by [List Documents](/api-reference/v2/endpoint/list-documents) with `include_fields: ["relations"]`), each relation's `timestamp` is a Unix epoch value in seconds (a float, e.g. `1778573640.0`). The dedicated [Context Relations](/api-reference/v2/endpoint/source-relations) endpoint returns the same field as an ISO-8601 string instead. Normalize before comparing relation timestamps across endpoints. - **Use the right metadata namespace:** Top-level `metadata_filters` keys match `metadata`; free-form per-document fields must be nested under `additional_metadata` (`document_metadata` is only a legacy alias). Declare hot top-level filter fields in `database_metadata_schema`. - **Check indexing first:** New documents are not searchable until [Ingestion Status](/api-reference/v2/endpoint/source-status) reports them indexed. - **Collection scope:** If you omit `collection`, HydraDB queries the default collection, and a query that names a collection does not read the default one. Use [List Collections](/api-reference/v2/endpoint/list-sub-tenants) to discover available IDs. - ## Errors diff --git a/api-reference/v2/endpoint/register-webhook.mdx b/api-reference/v2/endpoint/register-webhook.mdx index 7a48f0e4..d914d90e 100644 --- a/api-reference/v2/endpoint/register-webhook.mdx +++ b/api-reference/v2/endpoint/register-webhook.mdx @@ -6,10 +6,42 @@ openapi: "api-reference/v2/openapi.json POST /webhooks/indexing" Registers the endpoint HydraDB calls when ingested content reaches a terminal indexing state. One webhook is registered per workspace, so calling this again replaces the current registration. - - Omitting `signing_secret` **preserves** any secret you already have. Editing the URL or the event list never changes your signing configuration. To disable signing, call `DELETE /webhooks/indexing/signing-secret` explicitly. See [Manage the signing secret](/essentials/v2/webhooks#manage-the-signing-secret). - +Omitting `signing_secret` **preserves** any secret you already have. Editing the URL or the event list never changes your signing configuration. To disable signing, call `DELETE /webhooks/indexing/signing-secret` explicitly. See [Manage the signing secret](/essentials/v2/webhooks#manage-the-signing-secret). To have HydraDB create a signing secret, send `generate_signing_secret: true`. The secret comes back once, in this response, and cannot be read again. Send `generate_signing_secret` or `signing_secret`, not both. Your endpoint must be reachable over public HTTPS. Localhost and private network addresses are rejected. See [Webhooks](/essentials/v2/webhooks) for the payload format, retry behaviour, and receiver examples. + + + +```python Python SDK +webhook = client.webhooks.register( + url="https://example.com/hydradb/webhooks", + event_types=["indexing.status_changed"], + generate_signing_secret=True, +) +secret = webhook.data.signing_secret # returned once; store it now +``` + +```typescript TypeScript SDK +const webhook = await client.webhooks.register({ + url: "https://example.com/hydradb/webhooks", + eventTypes: ["indexing.status_changed"], + generateSigningSecret: true, +}); +const secret = webhook.data?.signingSecret; // returned once; store it now +``` + +```bash cURL +curl -X POST 'https://api.hydradb.com/webhooks/indexing' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" \ + -H "Content-Type: application/json" \ + -d '{ + "url": "https://example.com/hydradb/webhooks", + "event_types": ["indexing.status_changed"], + "generate_signing_secret": true + }' +``` + + diff --git a/api-reference/v2/endpoint/resume-connector.mdx b/api-reference/v2/endpoint/resume-connector.mdx index 5a14e26b..a12fcb8b 100644 --- a/api-reference/v2/endpoint/resume-connector.mdx +++ b/api-reference/v2/endpoint/resume-connector.mdx @@ -8,6 +8,14 @@ Puts a paused connector back on its schedule and makes it due for a sync right a +```python Python SDK +client.connectors.resume("{connector_id}") +``` + +```typescript TypeScript SDK +await client.connectors.resume({ id: "{connector_id}" }); +``` + ```bash cURL curl -X POST 'https://api.hydradb.com/connectors/{id}/resume' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/retry-webhook-delivery.mdx b/api-reference/v2/endpoint/retry-webhook-delivery.mdx index 467ca14a..c7c78833 100644 --- a/api-reference/v2/endpoint/retry-webhook-delivery.mdx +++ b/api-reference/v2/endpoint/retry-webhook-delivery.mdx @@ -9,3 +9,21 @@ Queues a `failed` or `permanently_failed` delivery for another attempt. Use it a The retry is signed with your **current** signing secret, not the one in force when the delivery was first attempted. If you have rotated since, your receiver must know the new secret. See [Webhooks](/essentials/v2/webhooks) for retry and backoff behaviour. + + + +```python Python SDK +result = client.webhooks.retry_delivery("{delivery_id}") +``` + +```typescript TypeScript SDK +const result = await client.webhooks.retryDelivery({ deliveryId: "{delivery_id}" }); +``` + +```bash cURL +curl -X POST 'https://api.hydradb.com/webhooks/indexing/deliveries/{delivery_id}/retry' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" +``` + + diff --git a/api-reference/v2/endpoint/source-relations.mdx b/api-reference/v2/endpoint/source-relations.mdx index 7a943b7f..0d8409f1 100644 --- a/api-reference/v2/endpoint/source-relations.mdx +++ b/api-reference/v2/endpoint/source-relations.mdx @@ -6,7 +6,7 @@ description: "Read the entity-relation triplets extracted from your content, for import { Field } from "/snippets/field.jsx"; -Inspect relationships extracted from your content, such as `PaymentsWorker → depends_on → OrdersDB`. Each three-part relationship is a **triplet**: a starting entity (a named person, service, or topic), a relationship, and a target entity. +Inspect relationships extracted from your content, such as `PaymentsWorker` depends on `OrdersDB`. Each three-part relationship is a **triplet**: a starting entity (a named person, service, or topic), a relationship, and a target entity. Pass `id` to scope to a single ingested item, or omit it to return all relations in the collection. Set `type=memory` to inspect a memory's relations. Pagination handles large result sets. @@ -199,9 +199,7 @@ while True: ## Notes - - **Cursor opacity:** `next_cursor` is opaque (currently the timestamp, in Unix seconds, of the last group returned). Don't construct it client-side or assume meaning: pass back exactly what the server returned. - +**Cursor opacity:** `next_cursor` is opaque (currently the timestamp, in Unix seconds, of the last group returned). Don't construct it client-side or assume meaning: pass back exactly what the server returned. - **Full-graph exports:** omit `id`, use a small `limit`, and paginate. - **Ordering:** `data.relations[]` comes back newest first, by each group's latest relation `timestamp`, not by relevance. diff --git a/api-reference/v2/endpoint/source-status.mdx b/api-reference/v2/endpoint/source-status.mdx index dcf4e91e..9708395e 100644 --- a/api-reference/v2/endpoint/source-status.mdx +++ b/api-reference/v2/endpoint/source-status.mdx @@ -10,9 +10,7 @@ Since ingestion is asynchronous, use this endpoint to determine when context is Pass the IDs you got back from ingestion in `ids` (or a single `id`). It works for documents, app sources, and memories. Whitespace is trimmed, empty entries are dropped, and duplicates are removed. - - **Prefer webhooks over polling?** Register a webhook for `indexing.status_changed` events and HydraDB will `POST` to your endpoint when content reaches a terminal state (`completed` or `errored`). See [Webhooks](/essentials/v2/webhooks) for setup and receiver examples. - +**Prefer webhooks over polling?** Register a webhook for `indexing.status_changed` events and HydraDB will `POST` to your endpoint when content reaches a terminal state (`completed` or `errored`). See [Webhooks](/essentials/v2/webhooks) for setup and receiver examples. @@ -132,9 +130,7 @@ Each entry in `data.statuses` describes one requested `id`: Status records can expire. A source that completed still reports `completed` after its record expires. A source that never finished before its record expired reports `errored` with `E9004`; ingest it again. - - Branch on `error_code`, not on the text in `message` or `error_message`. `message` describes the lookup, not the ingestion result, and human-readable text may change. The full list of codes an `errored` entry can carry is in the [Error Responses reference](/api-reference/v2/error-responses#ingestion-error-codes). - +Branch on `error_code`, not on the text in `message` or `error_message`. `message` describes the lookup, not the ingestion result, and human-readable text may change. The full list of codes an `errored` entry can carry is in the [Error Responses reference](/api-reference/v2/error-responses#ingestion-error-codes). ## Status values diff --git a/api-reference/v2/endpoint/sources-overview.mdx b/api-reference/v2/endpoint/sources-overview.mdx index 58d40326..c01a5382 100644 --- a/api-reference/v2/endpoint/sources-overview.mdx +++ b/api-reference/v2/endpoint/sources-overview.mdx @@ -55,16 +55,14 @@ flowchart LR style E fill:#0f172a,stroke:#ef4444,stroke-width:2px,color:#f8fafc ``` - - **Why both `type=knowledge` and `app_knowledge`?** They answer different questions. +**Why both `type=knowledge` and `app_knowledge`?** They answer different questions. - - `type` picks the **store**: `knowledge` (shared documents) or `memory` (per-user context). It routes the ingest to the right store. - - Within `type=knowledge`, you pick the **payload shape**: `documents` (binary documents HydraDB will parse: PDFs, DOCX, CSV) or `app_knowledge` (a JSON array of already-extracted content from your app: Slack messages, Notion pages, web pages). You can send both in the same request. - +- `type` picks the **store**: `knowledge` (documents, app sources, and facts) or `memory` (preferences, decisions, and past conversations). It routes the ingest to the right store. +- Within `type=knowledge`, you pick the **payload shape**: `documents` (binary documents HydraDB will parse: PDFs, DOCX, CSV) or `app_knowledge` (a JSON array of already-extracted content from your app: Slack messages, Notion pages, web pages). You can send both in the same request. ## Core ingestion concepts -- **Knowledge vs. memories:** [Knowledge](/essentials/v2/knowledge) is shared content (documents, app pages, Slack messages). [Memories](/essentials/v2/memories) are one user's preferences and conversation history. Both live in a `collection`, and `type: "all"` on `POST /query` searches both from the same scope. +- **Knowledge vs. memories:** [Knowledge](/essentials/v2/knowledge) is the documents, app sources, and facts your AI reasons over. [Memories](/essentials/v2/memories) are preferences, decisions, and past conversations. Both live in a `collection`, and `type: "all"` on `POST /query` searches both from the same scope. - **IDs:** Unique identifiers returned by `/context/ingest`. You can assign custom IDs using `id` in `document_metadata`, `app_knowledge`, or `memories` items. Use them for polling status, inspecting content, and deleting context. - **Metadata filtering:** You can scope queries using `metadata` (structured fields defined in your database schema) or `additional_metadata` (free-form per-document JSON). For detailed guidelines on structuring metadata, see the [Scoping using metadata](/essentials/v2/metadata) guide. diff --git a/api-reference/v2/endpoint/subgraph.mdx b/api-reference/v2/endpoint/subgraph.mdx index 5a2c61cd..44744ecd 100644 --- a/api-reference/v2/endpoint/subgraph.mdx +++ b/api-reference/v2/endpoint/subgraph.mdx @@ -189,9 +189,7 @@ Traversal is breadth-first, so `depth` on each member is its distance from the s ## Notes - - **An unknown `id` is an empty subgraph, not an error.** The endpoint does not confirm or deny that an item exists; the same answer comes back for an id that was never ingested and for one the `acl` principals may not see. - +**An unknown `id` is an empty subgraph, not an error.** The endpoint does not confirm or deny that an item exists; the same answer comes back for an id that was never ingested and for one the `acl` principals may not see. - **An item nothing links to** comes back as a one-member subgraph: itself, at depth `0`, with `max_depth_reached: 0`. That is a real answer ("this stands alone"), distinct from an unknown id, which has no members. - **Bounding the traversal:** Threads and hierarchies can be large. `depth` bounds how far the walk goes; `max_sources` bounds how many members it returns. When `max_sources` clips it, `is_truncated` is `true` and the members you have are the ones closest to the start item. `auxiliary_truncated` reports the same for the structural graph. diff --git a/api-reference/v2/endpoint/submit-feedback.mdx b/api-reference/v2/endpoint/submit-feedback.mdx index 3488b7f0..02e71887 100644 --- a/api-reference/v2/endpoint/submit-feedback.mdx +++ b/api-reference/v2/endpoint/submit-feedback.mdx @@ -27,9 +27,7 @@ Every HydraDB response carries a `request_id` in `meta`, and the same value in t } ``` - Send `request_id` back **exactly as you received it**. It must be the UUID from `meta.request_id` (or the `X-Request-ID` header). Any other value is rejected with `400`. - Submit feedback for queries that **returned**. If the query itself failed, handle the error instead: there is no retrieval to judge, and the fix is in the request rather than in the index. diff --git a/api-reference/v2/endpoint/sync-connector.mdx b/api-reference/v2/endpoint/sync-connector.mdx index 350da283..989df827 100644 --- a/api-reference/v2/endpoint/sync-connector.mdx +++ b/api-reference/v2/endpoint/sync-connector.mdx @@ -10,6 +10,14 @@ A `202` means the sync was started, not that it finished. If a sync is already r +```python Python SDK +run = client.connectors.sync("{connector_id}") +``` + +```typescript TypeScript SDK +const run = await client.connectors.sync({ id: "{connector_id}" }); +``` + ```bash cURL curl -X POST 'https://api.hydradb.com/connectors/{id}/sync' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/tenant-stats.mdx b/api-reference/v2/endpoint/tenant-stats.mdx index f90dd915..59e2ca61 100644 --- a/api-reference/v2/endpoint/tenant-stats.mdx +++ b/api-reference/v2/endpoint/tenant-stats.mdx @@ -86,9 +86,7 @@ curl -X GET 'https://api.hydradb.com/databases/stats?database=my_first_database' ## Behavior notes - - **`row_count` is chunks, not sources.** One ingested document typically becomes many chunks (e.g., a 30-page PDF can produce 100\+ rows). To count distinct sources, use [List Documents](/api-reference/v2/endpoint/list-documents) with `page_size=1` and read `pagination.total`. - +**`row_count` is chunks, not sources.** One ingested document typically becomes many chunks (e.g., a 30-page PDF can produce 100\+ rows). To count distinct sources, use [List Documents](/api-reference/v2/endpoint/list-documents) with `page_size=1` and read `pagination.total`. - **Both stores always report:** even if you've only ingested knowledge (or only memories), both counts are returned. An empty store reports `row_count: 0`. - **Stats are eventually consistent:** immediately after ingestion or deletion, counts may lag until background processing completes. Check [Ingestion Status](/api-reference/v2/endpoint/source-status) for recently ingested IDs before treating counts as final. @@ -101,5 +99,5 @@ curl -X GET 'https://api.hydradb.com/databases/stats?database=my_first_database' - **List sources/memories:** [List Documents](/api-reference/v2/endpoint/list-documents) - **List collections:** [List Collections](/api-reference/v2/endpoint/list-sub-tenants) - **Check provisioning:** [Database Status](/api-reference/v2/endpoint/tenant-status) - - **Read more:** [Concepts: Multi-Tenant Support](/essentials/v2/multi-tenant) + - **Read more:** [Multi-tenancy](/essentials/v2/multi-tenant) diff --git a/api-reference/v2/endpoint/tenants-overview.mdx b/api-reference/v2/endpoint/tenants-overview.mdx index 7020382d..33f5ab17 100644 --- a/api-reference/v2/endpoint/tenants-overview.mdx +++ b/api-reference/v2/endpoint/tenants-overview.mdx @@ -40,5 +40,5 @@ For routine operations on an existing database: ## Key concepts - **Database:** A top-level isolated space. For example, you can dedicate one database to one enterprise customer. -- **Collection:** Partitions within a database for per-user separation. The first collection is created implicitly at ingestion. Collections are useful when you need to scope data per user, team, or customer within a single database. +- **Collection:** A group of content inside a database, created on the first write. For a company brain, one collection holds everything, and metadata filters and access control keep results apart. Use separate collections only for data that must never meet in one search, such as each person's memories. - **Database Metadata & Schema:** Structured fields defined at database creation to enable query-time filtering. You can add fields later with [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema). \ No newline at end of file diff --git a/api-reference/v2/endpoint/test-webhook.mdx b/api-reference/v2/endpoint/test-webhook.mdx index da47f8f3..28c59008 100644 --- a/api-reference/v2/endpoint/test-webhook.mdx +++ b/api-reference/v2/endpoint/test-webhook.mdx @@ -8,8 +8,26 @@ Sends a synthetic `indexing.status_changed` payload to your registered endpoint The test payload carries `"test": true`, and is signed exactly like a real delivery when a signing secret is configured. Use it to check your verifier before relying on it in production. The response reports whether your endpoint accepted it (`delivered`) and the HTTP `status_code` it returned. If no webhook is registered, the call returns `404`. - - Test deliveries do not appear in the delivery history returned by [List Deliveries](/api-reference/v2/endpoint/list-webhook-deliveries). - +Test deliveries do not appear in the delivery history returned by [List Deliveries](/api-reference/v2/endpoint/list-webhook-deliveries). See [Webhooks](/essentials/v2/webhooks) for the full payload shape. + + + +```python Python SDK +result = client.webhooks.test() +print(result.data.delivered, result.data.status_code) +``` + +```typescript TypeScript SDK +const result = await client.webhooks.test(); +console.log(result.data?.delivered, result.data?.statusCode); +``` + +```bash cURL +curl -X POST 'https://api.hydradb.com/webhooks/indexing/test' \ + -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ + -H "API-Version: 2" +``` + + diff --git a/api-reference/v2/endpoint/update-connector-resource.mdx b/api-reference/v2/endpoint/update-connector-resource.mdx index bb3c34af..bd3ecca0 100644 --- a/api-reference/v2/endpoint/update-connector-resource.mdx +++ b/api-reference/v2/endpoint/update-connector-resource.mdx @@ -13,6 +13,26 @@ Updates one configured resource in place. Send `custom_instructions`, `acl`, or +```python Python SDK +client.connectors.update_resource_acl( + "{connector_id}", + "{resource_id}", + request={ + "custom_instructions": "These threads are incident retros. Extract root cause, impact and owner." + }, +) +``` + +```typescript TypeScript SDK +await client.connectors.updateResourceAcl({ + id: "{connector_id}", + resourceId: "{resource_id}", + body: { + custom_instructions: "These threads are incident retros. Extract root cause, impact and owner.", + }, +}); +``` + ```bash cURL curl -X PATCH 'https://api.hydradb.com/connectors/{id}/resources/{resource_id}' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/update-connector.mdx b/api-reference/v2/endpoint/update-connector.mdx index 42cf77e6..0c3097ae 100644 --- a/api-reference/v2/endpoint/update-connector.mdx +++ b/api-reference/v2/endpoint/update-connector.mdx @@ -14,6 +14,27 @@ Only these three fields are read. A connector's `name`, `database`, `collection` +```python Python SDK +client.connectors.update( + "{connector_id}", + custom_instructions=( + "Aurora was renamed to Nimbus in January 2025; they are the same product. " + "Index all Aurora content under Nimbus." + ), + sync_interval_seconds=3600, +) +``` + +```typescript TypeScript SDK +await client.connectors.update({ + id: "{connector_id}", + customInstructions: + "Aurora was renamed to Nimbus in January 2025; they are the same product. " + + "Index all Aurora content under Nimbus.", + syncIntervalSeconds: 3600, +}); +``` + ```bash cURL curl -X PATCH 'https://api.hydradb.com/connectors/{id}' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ diff --git a/api-reference/v2/endpoint/update-metadata-schema.mdx b/api-reference/v2/endpoint/update-metadata-schema.mdx index 273775ec..36e8cbc9 100644 --- a/api-reference/v2/endpoint/update-metadata-schema.mdx +++ b/api-reference/v2/endpoint/update-metadata-schema.mdx @@ -8,12 +8,30 @@ import { Field } from "/snippets/field.jsx"; Use this endpoint to add new fields to a database's `database_metadata_schema`. SDK methods: `client.databases.update_metadata_schema()` (Python) and `client.databases.updateMetadataSchema()` (TypeScript). `GET /databases/{database}/metadata-schema` returns the current schema in the same field shape. - - This endpoint is additive only. It cannot delete fields, rename fields, change the type/flags of existing fields, or add dense/sparse search lanes. - +This endpoint is additive only. It cannot delete fields, rename fields, change the type/flags of existing fields, or add dense/sparse search lanes. +```python Python SDK +response = client.databases.update_metadata_schema( + "acme_corp", + add_fields=[ + {"name": "region", "data_type": "VARCHAR"}, + {"name": "priority", "data_type": "INT64"}, + ], +) +``` + +```typescript TypeScript SDK +const response = await client.databases.updateMetadataSchema({ + database: "acme_corp", + addFields: [ + { name: "region", dataType: "VARCHAR" }, + { name: "priority", dataType: "INT64" }, + ], +}); +``` + ```bash cURL curl -X PATCH 'https://api.hydradb.com/databases/acme_corp/metadata-schema' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ @@ -33,42 +51,6 @@ curl -X PATCH 'https://api.hydradb.com/databases/acme_corp/metadata-schema' \ }' ``` -```python Python -import requests - -response = requests.patch( - "https://api.hydradb.com/databases/acme_corp/metadata-schema", - headers={ - "Authorization": f"Bearer {HYDRA_DB_API_KEY}", - "API-Version": "2", - "Content-Type": "application/json", - }, - json={ - "add_fields": [ - {"name": "region", "data_type": "VARCHAR"}, - {"name": "priority", "data_type": "INT64"}, - ] - }, -) -``` - -```typescript TypeScript -const response = await fetch("https://api.hydradb.com/databases/acme_corp/metadata-schema", { - method: "PATCH", - headers: { - Authorization: `Bearer ${process.env.HYDRA_DB_API_KEY}`, - "API-Version": "2", - "Content-Type": "application/json", - }, - body: JSON.stringify({ - add_fields: [ - { name: "region", data_type: "VARCHAR" }, - { name: "priority", data_type: "INT64" }, - ], - }), -}); -``` - ## Request @@ -142,9 +124,7 @@ A success returns `database`, the deprecated `tenant_id`, and `added_fields` dir | `409` | A field conflicts with an existing field or with another entry in the same request, or the database changed during the request (retry this last case). | | `500` | Saving the schema failed. Retry. | - - Apart from a malformed body and a missing database, this endpoint currently returns `error.code: "INTERNAL_ERROR"` for `400`, `409`, and `500` alike. Branch on the HTTP status and read `error.message`, which names the field and the rule it broke. - +Apart from a malformed body and a missing database, this endpoint currently returns `error.code: "INTERNAL_ERROR"` for `400`, `409`, and `500` alike. Branch on the HTTP status and read `error.message`, which names the field and the rule it broke. ## Related diff --git a/api-reference/v2/endpoint/update-source-metadata.mdx b/api-reference/v2/endpoint/update-source-metadata.mdx index e63310bb..34f6a81d 100644 --- a/api-reference/v2/endpoint/update-source-metadata.mdx +++ b/api-reference/v2/endpoint/update-source-metadata.mdx @@ -12,12 +12,30 @@ Use this endpoint when you know a source ID and need to update its metadata in p PATCH /context/{id}/metadata ``` - - The legacy route `PATCH /context/sources/{source_id}/metadata` still works but is deprecated. Migrate to the route above. Both dispatch to the same handler; `source_id` and `id` name the same value. - +The legacy route `PATCH /context/sources/{source_id}/metadata` still works but is deprecated. Migrate to the route above. Both dispatch to the same handler; `source_id` and `id` name the same value. +```python Python SDK +response = client.context.update_source_metadata( + "policy_main", + database="acme_corp", + collection="team_docs", + database_metadata={"department": "legal", "priority": 7}, + additional_metadata={"author": "Legal Team", "doc_version": 3}, +) +``` + +```typescript TypeScript SDK +const response = await client.context.updateSourceMetadata({ + id: "policy_main", + database: "acme_corp", + collection: "team_docs", + databaseMetadata: { department: "legal", priority: 7 }, + additionalMetadata: { author: "Legal Team", doc_version: 3 }, +}); +``` + ```bash cURL curl -X PATCH 'https://api.hydradb.com/context/policy_main/metadata' \ -H "Authorization: Bearer $HYDRA_DB_API_KEY" \ @@ -37,56 +55,6 @@ curl -X PATCH 'https://api.hydradb.com/context/policy_main/metadata' \ }' ``` -```python Python -import os - -import requests - -response = requests.patch( - "https://api.hydradb.com/context/policy_main/metadata", - headers={ - "Authorization": f"Bearer {os.environ['HYDRA_DB_API_KEY']}", - "API-Version": "2", - "Content-Type": "application/json", - }, - json={ - "database": "acme_corp", - "collection": "team_docs", - "database_metadata": { - "department": "legal", - "priority": 7, - }, - "additional_metadata": { - "author": "Legal Team", - "doc_version": 3, - }, - }, -) -``` - -```typescript TypeScript -const response = await fetch("https://api.hydradb.com/context/policy_main/metadata", { - method: "PATCH", - headers: { - Authorization: `Bearer ${process.env.HYDRA_DB_API_KEY}`, - "API-Version": "2", - "Content-Type": "application/json", - }, - body: JSON.stringify({ - database: "acme_corp", - collection: "team_docs", - database_metadata: { - department: "legal", - priority: 7, - }, - additional_metadata: { - author: "Legal Team", - doc_version: 3, - }, - }), -}); -``` - ## Request @@ -109,9 +77,7 @@ const response = await fetch("https://api.hydradb.com/context/policy_main/metada At least one of `database_metadata`, `additional_metadata`, or `acl` is required. - - This edit endpoint uses `database_metadata` for schema-backed source metadata (deprecated alias: `tenant_metadata`, still accepted, but the canonical field wins if both are sent). The shorter `metadata` field used by ingestion/list examples is not accepted in this PATCH body. `document_metadata` is also not accepted; use `additional_metadata`. - +This edit endpoint uses `database_metadata` for schema-backed source metadata (deprecated alias: `tenant_metadata`, still accepted, but the canonical field wins if both are sent). The shorter `metadata` field used by ingestion/list examples is not accepted in this PATCH body. `document_metadata` is also not accepted; use `additional_metadata`. ## Behavior diff --git a/api-reference/v2/error-responses.mdx b/api-reference/v2/error-responses.mdx index beaa7b66..87384520 100644 --- a/api-reference/v2/error-responses.mdx +++ b/api-reference/v2/error-responses.mdx @@ -33,9 +33,7 @@ HydraDB core endpoints (`/databases`, `/context/*`, `/query`, `/feedback`, and ` | `meta.latency_ms` | Server-side processing time in milliseconds. | | `detail` | Deprecated copy of the error (`detail.error_code`, `detail.message`) kept for older clients. Read `error` instead. | - Use `error.code` for branching and log `meta.request_id` for every failed request. The HTTP status tells you the class of failure; the error code tells you what to do. - ## HTTP status codes @@ -74,17 +72,11 @@ Use `error.code` for branching and log `meta.request_id` for every failed reques | `INTERNAL_ERROR` | `500` | HydraDB hit an unexpected server-side error. A rare auth-layer failure reports `INTERNAL_SERVER_ERROR` instead. | | `SERVICE_UNAVAILABLE` | `503` | A dependency is temporarily unavailable or the service is under load. | - Endpoint pages list the most common codes for that operation. New codes may be added over time, so clients should handle unknown `error.code` values gracefully. - - `PATCH /databases/{database}/metadata-schema` currently returns `INTERNAL_ERROR` for its `400` and `409` errors too, so branch on the HTTP status there. - - -**Deprecated `/tenants` routes keep their pre-rename error codes.** For backward compatibility, a request for a database that does not exist returns `NOT_FOUND` on the deprecated `/tenants` routes (not `DATABASE_NOT_FOUND`), and a duplicate on `POST /tenants` returns `INVALID_INPUT` (not `DATABASE_ALREADY_EXISTS`), whereas the canonical `/databases` routes return `DATABASE_NOT_FOUND` and `DATABASE_ALREADY_EXISTS` as shown above. The HTTP status is identical on both. The route decides the code, so the old `tenant_id` field sent to a `/databases` route still gets the new codes. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). - +**Deprecated `/tenants` routes keep their pre-rename error codes:** for backward compatibility, a request for a database that does not exist returns `NOT_FOUND` on the deprecated `/tenants` routes (not `DATABASE_NOT_FOUND`), and a duplicate on `POST /tenants` returns `INVALID_INPUT` (not `DATABASE_ALREADY_EXISTS`), whereas the canonical `/databases` routes return `DATABASE_NOT_FOUND` and `DATABASE_ALREADY_EXISTS` as shown above. The HTTP status is identical on both. The route decides the code, so the old `tenant_id` field sent to a `/databases` route still gets the new codes. See [Migrating from `tenant_id` and `sub_tenant_id`](/essentials/v2/multi-tenant#7-migrating-from-the-legacy-tenant-and-sub-tenant-fields). ## Ingestion error codes @@ -97,14 +89,14 @@ Many storage- and capacity-related ingestion errors are **transient**: the pipel | `E1002` | The file format is not supported. Returned in the [`POST /context/ingest`](/api-reference/v2/endpoint/ingest-context) response itself, per file, before anything is queued. User message: *"This file format isn't supported. Please upload a PDF, Office document (Word, Excel, PowerPoint), image, CSV or text file."* | **Terminal** (not retried) | | `E6001` | Vector-store storage/indexing error while persisting processed data. The pipeline retries automatically and it usually clears within minutes. User message: *"Failed to store the processed data. Please try again. If the issue persists, contact support@hydradb.com."* | **Transient** (retryable) | - `E1002` is the one ingestion code that does **not** follow the polling advice above. The file is rejected at upload, so it never enters the pipeline and never gets a status record. Polling [`/context/status`](/api-reference/v2/endpoint/source-status) for it returns `FILE_NOT_FOUND`, not `E1002`. Read `error_code` on each item in the upload response instead. See [Supported file formats](/api-reference/v2/endpoint/ingest-context#supported-file-formats) for what is accepted, and note that one rejected file does not affect the other files in the same request. - ## Retry pattern Retry only transient failures: `429`, `500`, and `503`. Use exponential backoff with jitter and keep retries bounded. Two other codes clear on their own, so wait instead of backing off blindly: `409 SOURCE_PROCESSING` (wait for `Retry-After`) and `422 TENANT_INFRA_NOT_READY` (wait until the database is ready). +**The SDKs already retry:** both retry `408`, `429`, and `5xx` responses twice with exponential backoff before raising, and the Python SDK also retries `409`. Raise the count with `request_options={"max_retries": 5}` on a Python call, or `maxRetries` on the TypeScript client or call. The examples below add a wrapper for longer outages or for raw HTTP. + ```typescript TypeScript SDK import { HydraDBError } from "@hydradb/sdk"; diff --git a/api-reference/v2/index.mdx b/api-reference/v2/index.mdx index 8ed50985..ec5e7ae7 100644 --- a/api-reference/v2/index.mdx +++ b/api-reference/v2/index.mdx @@ -3,7 +3,7 @@ title: "API Reference" description: "The four calls most integrations are built on, then every v2 endpoint." --- -Use this reference to look up request fields and response formats. Every request carries `Authorization: Bearer ` and `API-Version: 2` against `https://api.hydradb.com`; the [SDKs](/api-reference/v2/sdks) set both for you. If you are making your first call, follow the [Quickstart](/get-started/v2/quickstart). AI agents can start from the [Agent Integration Guide](/AGENTS) and the [v2 OpenAPI spec](/api-reference/v2/openapi.json). **Context** means the documents, passages, and memories your model uses to answer a question; HydraDB stores and retrieves them, while your model writes the answer. +Use this reference to look up request fields and response formats. New to HydraDB? The [Quickstart](/get-started/v2/quickstart) gets you to a first query in five minutes, and [Core Concepts](/get-started/v2/core-concepts) explains databases, context, and queries. AI agents can start from the [Agent Integration Guide](/AGENTS) and the [v2 OpenAPI spec](/api-reference/v2/openapi.json). ## The four calls @@ -36,62 +36,13 @@ flowchart LR style G fill:#0f172a,stroke:#334155,stroke-width:2px,color:#f8fafc ``` -## Endpoint groups +## Every endpoint -| Group | Purpose | When to reach for it | -|---|---|---| -| [Databases](/api-reference/v2/endpoint/tenants-overview) | Create, monitor, and manage isolated workspaces | First step in any integration, and any time you need usage stats, provisioning status, or to tear down a workspace | -| [Context](/api-reference/v2/endpoint/sources-overview) | Ingest, list, fetch, delete, and inspect knowledge or memories | Every time data flows into HydraDB: document uploads, app sources, user memories, and lifecycle ops | -| [Query](/api-reference/v2/endpoint/query-overview) | Retrieve context with hybrid or text query, and send feedback on the results | Find relevant passages to include in your model prompt | -| [Connectors](/api-reference/v2/endpoint/connectors-overview) | Connect, configure, sync, and manage app connectors such as Slack, GitHub, Google Drive, and Supabase | When you want app data synced into a database without writing ingest code | -| [Webhooks](/api-reference/v2/endpoint/webhooks-overview) | Register a URL that receives a `POST` when indexing finishes, and inspect or retry deliveries | When you would rather be notified than poll `/context/status` | +The Python SDK method is listed for each endpoint; the TypeScript SDK uses the same names in camelCase, such as `deleteCollection`. Every endpoint page shows Python, TypeScript, and cURL. -## Core concepts +### Databases -| Concept | What it means | When you use it | -|---|---|---| -| `database` | Your isolated workspace for data, metadata schema, and query. | Send it on database-scoped calls such as ingestion and query. Formerly `tenant_id`; the `tenant_id` alias is still accepted (deprecated). | -| `collection` | Optional partition inside a database, often a user, team, account, or customer. | Send the same collection on writes and reads. Omitting it uses the default collection, not all collections. Formerly `sub_tenant_id`; the `sub_tenant_id` alias is still accepted (deprecated). Read more about our [multi-tenant architecture](/essentials/v2/multi-tenant) | -| [Knowledge](/essentials/v2/knowledge) | Source material such as PDFs, docs, app pages, tickets, Slack threads, or webpages. | Use `type=knowledge` to search documents and app content. | -| [Memory](/essentials/v2/memories) | Preferences, conversation history, notes, and saved facts. | Use `type=memory` when the content should personalize answers for a specific user or collection. | -| `database_metadata_schema` | Database-level fields you define up front so metadata can be filtered or queried consistently. | Use it for stable fields like department, customer, region, plan, category, or compliance label. | -| `document_metadata` | JSON-stringified per-document metadata array sent during file ingestion. | Use it to attach a source `id`, schema-backed `metadata`, free-form `additional_metadata`, and relations to other sources to each uploaded file, in the same order as `documents`. | -| `ids` | IDs returned by ingestion or visible from `/context/list`. | Use them when polling processing status, inspecting content, listing a specific subset, deleting sources, or inspecting relations. | - -## SDKs - -HydraDB publishes official SDKs for Python and TypeScript/Node. They wrap every endpoint in this reference with typed methods and IDE autocomplete. - -| Language | Package | Install | -|---|---|---| -| **Python** | [`hydradb-sdk` on PyPI](https://pypi.org/project/hydradb-sdk/) | `pip install hydradb-sdk` | -| **TypeScript / Node** | [`@hydradb/sdk` on npm](https://www.npmjs.com/package/@hydradb/sdk) | `npm install @hydradb/sdk` | - -**Quick init:** - - - -```python Python -import os -from hydra_db import HydraDB, AsyncHydraDB - -client = HydraDB(token=os.environ["HYDRA_DB_API_KEY"]) -async_client = AsyncHydraDB(token=os.environ["HYDRA_DB_API_KEY"]) -``` - -```typescript TypeScript -import { HydraDBClient } from "@hydradb/sdk"; - -const client = new HydraDBClient({ - token: process.env.HYDRA_DB_API_KEY, -}); -``` - - - -SDK methods mirror the API: `client..()` maps to the corresponding endpoint. The SDKs are generated from the OpenAPI contract and set `API-Version: 2` automatically. The SDK method names below are the Python names; TypeScript camelCases multi-word names (for example, `delete_collection` is `deleteCollection`). - -## Full endpoint inventory +Create and manage the isolated workspaces your data lives in. [Overview](/api-reference/v2/endpoint/tenants-overview). | Endpoint | Method | SDK method | Purpose | Use when | |---|---|---|---|---| @@ -103,6 +54,13 @@ SDK methods mirror the API: `client..()` maps to the correspondin | [`/databases/collections`](/api-reference/v2/endpoint/list-sub-tenants) | `GET` | `databases.collections` | List active collections | You partition data by user, team, customer, or account and need to inspect those partitions. | | [`/databases/collections`](/api-reference/v2/endpoint/delete-collection) | `DELETE` | `databases.delete_collection` | Delete a collection | You need to permanently remove one collection and its data without deleting the database. | | [`/databases/stats`](/api-reference/v2/endpoint/tenant-stats) | `GET` | `databases.stats` | Get usage statistics | You want to monitor object counts for a database. | + +### Context + +Ingest, check, list, inspect, edit, and delete what you store. [Overview](/api-reference/v2/endpoint/sources-overview). + +| Endpoint | Method | SDK method | Purpose | Use when | +|---|---|---|---|---| | [`/context/ingest`](/api-reference/v2/endpoint/ingest-context) | `POST` | `context.ingest` | Ingest knowledge or memories | You are uploading documents, app sources, or user memories. | | [`/context/status`](/api-reference/v2/endpoint/source-status) | `GET` | `context.status` | Check processing status | You have IDs from ingestion and need to know when they are queryable. | | [`/context/inspect`](/api-reference/v2/endpoint/fetch-content) | `GET` | `context.inspect` | Inspect original source content or presigned URL | You need to display or inspect the original ingested content. | @@ -111,8 +69,22 @@ SDK methods mirror the API: `client..()` maps to the correspondin | [`/context`](/api-reference/v2/endpoint/delete-source) | `DELETE` | `context.delete` | Delete sources or memories | You need to remove one or more knowledge sources or memories by ID. | | [`/context/relations`](/api-reference/v2/endpoint/source-relations) | `GET` | `context.relations` | Inspect entity relationships | You need graph relations for a source or collection. | | [`/context/{id}/subgraph`](/api-reference/v2/endpoint/subgraph) | `GET` | `context.subgraph` | Get the connected subgraph | You need the connected subgraph around one source. | + +### Query + +Retrieve context, and tell HydraDB how a query performed. [Overview](/api-reference/v2/endpoint/query-overview). + +| Endpoint | Method | SDK method | Purpose | Use when | +|---|---|---|---|---| | [`/query`](/api-reference/v2/endpoint/query) | `POST` | `query` | Unified query over knowledge, memories, or both | You need retrieval with `hybrid` or `text` query across `type: "knowledge"`, `type: "memory"`, or `type: "all"`. | | [`/feedback`](/api-reference/v2/endpoint/submit-feedback) | `POST` | `feedback.submit` | Submit query feedback | You want to tell HydraDB whether a query returned what you needed. | + +### Connectors + +Sync data from 70 apps without writing ingestion code. [Overview](/api-reference/v2/endpoint/connectors-overview). + +| Endpoint | Method | SDK method | Purpose | Use when | +|---|---|---|---|---| | [`/connectors/providers`](/api-reference/v2/endpoint/list-connector-providers) | `GET` | `list_providers` | List connector providers | You need the providers you can connect, or the credentials one provider needs. | | [`/connectors`](/api-reference/v2/endpoint/create-connector) | `POST` | `connectors.create` | Create a connector | You are storing credentials for one provider account. | | [`/connectors/{id}/discover`](/api-reference/v2/endpoint/discover-connector-resources) | `GET` | `connectors.discover` | Discover resources | You need the resources the connector's credentials can reach. | @@ -129,6 +101,13 @@ SDK methods mirror the API: `client..()` maps to the correspondin | [`/connectors/{id}/resources`](/api-reference/v2/endpoint/add-connector-resource) | `POST` | `connectors.create_resource` | Add a connector resource | You want to add one new resource to an existing connector. | | [`/connectors/{id}/resources/{resource_id}`](/api-reference/v2/endpoint/delete-connector-resource) | `DELETE` | `connectors.delete_resource` | Delete a connector resource | You want to stop syncing one resource. | | [`/connectors/{id}`](/api-reference/v2/endpoint/delete-connector) | `DELETE` | `connectors.delete` | Delete a connector | You need to remove a connector. | + +### Webhooks + +Get notified when indexing finishes instead of polling. [Overview](/api-reference/v2/endpoint/webhooks-overview). + +| Endpoint | Method | SDK method | Purpose | Use when | +|---|---|---|---|---| | [`/webhooks/indexing`](/api-reference/v2/endpoint/register-webhook) | `POST` | `webhooks.register` | Register a webhook | You want a `POST` to your URL when indexing finishes, instead of polling. | | [`/webhooks/indexing`](/api-reference/v2/endpoint/get-webhook) | `GET` | `webhooks.get` | Get the webhook | You need the current webhook configuration. | | [`/webhooks/indexing`](/api-reference/v2/endpoint/delete-webhook) | `DELETE` | `webhooks.delete` | Delete the webhook | You want to stop receiving indexing notifications. | @@ -139,7 +118,7 @@ SDK methods mirror the API: `client..()` maps to the correspondin ## Conventions -**Authentication:** Every endpoint requires `Authorization: Bearer ` in the request header. Get your key at [app.hydradb.com](https://app.hydradb.com). +**Authentication:** Every endpoint requires `Authorization: Bearer `. Get your key at [app.hydradb.com](https://app.hydradb.com). The [SDKs](/api-reference/v2/sdks) set it, and the version header, for you. **Versioning:** Send `API-Version: 2` with every request. @@ -182,28 +161,8 @@ Errors use the same envelope on every endpoint, including the exceptions above, - **Parameter casing:** The REST API uses snake_case (`database`, `max_results`). The TypeScript SDK uses camelCase keys and method names (`maxResults`, `deleteCollection`) and ignores request keys it does not recognize, including snake_case ones. The Python SDK uses snake_case throughout. See [SDKs](/api-reference/v2/sdks#naming-conventions). -- **Query modes:** `POST /query` supports `query_by: "hybrid"` or `"text"` and `type: "knowledge"`, `"memory"`, or `"all"`. The same `type` enum is used across ingestion, listing, deletion, and query; query additionally accepts `"all"`. - -**Status codes:** Successful responses return `200` (or `202` for async accepts). Errors follow standard HTTP semantics: - -| Code | Meaning | -|---|---| -| `200` | Success | -| `202` | Accepted (async operation queued) | -| `400` | Invalid parameters | -| `401` | Authentication required | -| `402` | Your plan does not allow the request | -| `403` | Forbidden | -| `404` | Resource not found | -| `409` | Conflict (e.g., database already exists) | -| `413` | Request body too large | -| `415` | Unsupported content type | -| `422` | Validation error, or a query sent before the database is ready | -| `429` | Rate limit exceeded | -| `500` | Internal server error | -| `503` | Service unavailable | - -See [Error Responses](/api-reference/v2/error-responses) for response shapes, error codes, and retry patterns. + +**Status codes:** `200` for success and `202` when an async operation is queued. See [Error Responses](/api-reference/v2/error-responses) for response shapes, error codes, and retry patterns. ## Rate limits @@ -214,6 +173,6 @@ Rate limits apply per API key. For production deployments, build retry logic wit Existing v1 endpoints remain available under the v1 API Reference. - **Build something:** [Quickstart](/get-started/v2/quickstart) walks through your first integration in five minutes -- **Understand the model:** [Core Concepts](/get-started/v2/core-concepts) explains databases, collections, content categories, and query results +- **Understand the model:** [Core Concepts](/get-started/v2/core-concepts) explains databases, context, and queries - **Go deeper:** [Usage](/essentials/v2/query) covers each primitive in depth - **Install an SDK:** [Python](https://pypi.org/project/hydradb-sdk/) · [TypeScript](https://www.npmjs.com/package/@hydradb/sdk) diff --git a/api-reference/v2/sdks.mdx b/api-reference/v2/sdks.mdx index 95b83e68..3477949e 100644 --- a/api-reference/v2/sdks.mdx +++ b/api-reference/v2/sdks.mdx @@ -42,9 +42,7 @@ const client = new HydraDBClient({ ``` - **Python:** Both synchronous (`HydraDB`) and asynchronous (`AsyncHydraDB`) clients are available. They share an identical surface; choose based on your application's concurrency model. - ## Naming conventions @@ -56,13 +54,11 @@ The REST API uses **snake_case** for all request and response fields, and the Py | **Python SDK** | snake_case | snake_case | method: `query()`, params: `database`, `query_by` | | **TypeScript SDK** | camelCase for multi-word methods | camelCase | method: `query()`, params: `database`, `queryBy` | - **Use camelCase request fields in TypeScript** (e.g., `client.context.list({ database: "...", pageSize: 50 })`). The SDK converts them to snake_case on the wire and silently drops keys it does not recognize, such as `page_size: 50`. JSON-string values (`documentMetadata`, `memories`, `appKnowledge`) and free-form maps (`metadataFilters`) are sent as written, so keep snake_case inside them. The Python SDK also uses snake_case throughout (e.g., `client.context.list(database="...", page_size=50)`). `database` and `collection` are the current field names (formerly `tenant_id` and `sub_tenant_id`). The old names remain accepted as deprecated aliases for full backward compatibility. - ## SDK method structure @@ -339,9 +335,7 @@ const result = await client.context.ingest({ ``` - **`infer` defaults to `false`.** Set `infer: true` for conversational content where you want HydraDB to extract implicit preferences. Use `infer: false` for content that should be stored verbatim. - ### Verify processing @@ -570,11 +564,9 @@ try { ``` - The Python SDK raises a typed exception per status from `hydra_db.errors`. Each subclasses `ApiError` and carries `status_code`, `headers`, and the parsed response `body`, from which you can read `body["error"]["code"]`. Not every status has its own class: `401` and `429` do not. Catch the base `ApiError` and branch on `status_code`, as above, whenever you need to handle those. - For the full list of error codes and retry patterns, see [Error Responses](/api-reference/v2/error-responses).