diff --git a/aigw/api-reference/inference-api/authentication.mdx b/aigw/api-reference/inference-api/authentication.mdx index 608a26c6..645a2c01 100644 --- a/aigw/api-reference/inference-api/authentication.mdx +++ b/aigw/api-reference/inference-api/authentication.mdx @@ -102,7 +102,7 @@ Read more [here](/aigw/integrations/llms/openai). The AI Gateway supports JWT-based authentication as a secure alternative to API Key authentication. With JWT authentication, clients can authenticate API requests using a JWT token that is validated against a configured JWKS (JSON Web Key Set). -This enterprise-grade authentication method is available as an add-on to any AI Gateway plan. JWT authentication provides enhanced security through: +JWT authentication provides enhanced security through: - Temporary, expiring tokens - Fine-grained permission scopes @@ -112,9 +112,3 @@ This enterprise-grade authentication method is available as an add-on to any AI Learn how to implement JWT-based authentication with the AI Gateway - - - Interested in adding JWT authentication to your AI Gateway plan? - - [Contact our sales team](https://portkey.sh/jwt) to discuss pricing and implementation details. - diff --git a/aigw/api-reference/inference-api/supported-providers.mdx b/aigw/api-reference/inference-api/supported-providers.mdx index 7b766058..188c60b7 100644 --- a/aigw/api-reference/inference-api/supported-providers.mdx +++ b/aigw/api-reference/inference-api/supported-providers.mdx @@ -1,6 +1,5 @@ --- title: "Supported Providers" -mode: "wide" --- | Provider | Chat | Vision | Tools | Prisma AIRS AI Gateway Prompts | Embeddings | Images | Audio | Finetuning | Batch | Files | Moderations | Assistants | Completions | Messages | OCR | diff --git a/aigw/changelog/data-service.mdx b/aigw/changelog/data-service.mdx deleted file mode 100644 index c92c2de7..00000000 --- a/aigw/changelog/data-service.mdx +++ /dev/null @@ -1,536 +0,0 @@ ---- -title: "Data Service" -sidebarTitle: "Data Service [1.9.0]" -rss: true ---- - - -## v1.9.0 ---- - -### Redis - -- Redis improvements for `Azure Redis` with support for multiple auth modes. -- Redis connection support through Discovery URL `REDIS_CLUSTER_DISCOVERY_URL`. - -[Self-Hosting Architecture](/aigw/self-hosting/hybrid-deployments/architecture) - -### Fixes and Improvements -- **Batches**: Validate the endpoint before submitting batch jobs. [Batches Documentation](/aigw/product/ai-gateway/batches) -- **Security**: Added SSRF checks for external calls. - [Custom Hosts Documentation](/aigw/product/ai-gateway/custom-hosts) -- **Prometheus**: Skip Prometheus client initialization when env values are not set. [Prometheus Metrics](/aigw/self-hosting/prometheus-metrics) -- Updated dependencies to patch security vulnerabilities. - - - -## v1.8.1 ---- - -### Fixes and Improvements -- Fixed TLS CA certificate handling. - - - -## v1.8.0 ---- - -### Features -- Added support for [Guardrails on provider batches](/aigw/product/guardrails/guardrails-for-batches). Pass a config containing `input_guardrails` and/or `output_guardrails` via `portkey_options` when creating a 24h provider batch — Prisma AIRS AI Gateway applies input guardrails to redact or filter rows before forwarding to the provider, and runs output guardrails on the upstream output before exposing the final file. - - -**Bull Board is now opt-in and disabled by default.** Set `BULL_BOARD_ENABLED=true` to enable the queue dashboard. - - -### Fixes and Improvements -- Fixed provider batch logs incorrectly using the batch ID as the model name. The actual model is now extracted from the parsed provider output before falling back to the batch ID. - - - -## v1.7.1 ---- - -### Features -- Added support for **Azure Blob Storage** as a log-export storage backend, alongside the existing S3 options. -- Added a **FIPS-compliant Dockerfile** variant. The image runs the data-service with FIPS mode enabled and uses a FIPS-compliant SHA-256 implementation (`@smithy/hash-node`) for AWS request signing. See [FIPS-Compliant Images](/self-hosting/fips-compliant-images) for deployment details. - - - - -## v1.7.0 ---- - - -**Enterprise Gateway 2.8.0+ required.** This data-service version is only compatible with Enterprise Gateway **v2.8.0** or newer. - - -### Improvements -- Remove completed jobs from Redis to reduce memory usage. -- Enhanced authorization checks for Gateway communication. -- Updated dependencies to patch security vulnerabilities. - - - -## v1.6.3 ---- - -### Improvements -- Added Cosign installation and signing step for Docker images. - - - -## v1.6.2 ---- - -### Fixes and Improvements -- Fixed etag handling for small chunk uploads. -- Fixed Loki transport logs. -- Updated dependencies to patch security vulnerabilities. - - - -## v1.6.1 ---- - -### Improvements -- Support multiple auth types for AWS, Azure, and GCS. -- Support pod identity management. -- Updated dependencies to patch security vulnerabilities. - - - -## v1.6.0 ---- - -### Fixes and Improvements -- Updated dependencies to patch security vulnerabilities. -- Fixed `Log Exports` memory leak. - - - -## v1.5.2 ---- - -### Fixes and Improvements -- Relaxed validation for `completion_window` in provider batches -- Updated dependencies to patch security vulnerabilities. - - - -## v1.5.1 ---- - -### Fixes and Improvements -- Updated dependencies to patch security vulnerabilities. - - - -## v1.5.0 ---- - -### Pricing Improvements -- Added support for dedicated batch pricing configurations. The exact pricing will be calculated from Gateway. -This is a breaking change. Requires `Enterprise Gateway` version v2.1.0 or higher. - -### Improvements -- Security updates to dependencies. - - - - -## v1.4.4 ---- - -### Improvements -- Security updates to dependencies. - - - - -## v1.4.3 ---- - -### Batches Improvements -- Added azure responses batches pricing support. - - - - -## v1.4.2 ---- - -### Fixes and Improvements -- Added new batch statuses for batch jobs. - * `file_upload_failed` - the batch job failed to upload the file to the provider after validation - * `batch_start_failed` - the file upload succeeded but the batch job submission to the provider failed - - - - -## v1.4.1 ---- - -### Fixes and Improvements -- Enhanced log exports functionality to read logs based on the path format identifier released in Gateway v1.17.0 - - - - -## v1.4.0 ---- - -**Requires a Helm repo update (>app-1.4.0)** - - -**For air-gapped deployments, `Backend` version v1.5.0 is required as it adds new columns in the analytics store** - - -### Security Patch -- Removed root user from container image (BREAKING CHANGE). The container image used by this chart no longer runs as root. The image now runs processes with a non-root UID and enforces a non-root container securityContext. Requires Helm repo upgrade (>app-1.4.0) to deploy the new image and chart settings. - -### Fixes and Improvements -- Added cost calculation logic for fine-tuning operations across multiple AI providers (OpenAI, Azure OpenAI, and Vertex AI) -- Improved error handling to preserve provider-level file upload failures during batch processing - - - - -## v1.3.0 ---- - -### Fixes and Improvements -- Support KMS key support for custom batch file output. -- Support batch operations for tracking provider batch requests from gateway. - Works with Gateway version 1.16.0 or higher. - - - -## v1.2.8 ---- - -### Fixes and Improvements -- Fixed redundant failed batch status update for batch jobs. -- Fixed - Retry failed job status check and update status accordingly. - - - -## v1.2.7 ---- - -### Fixes and Improvements -- Fixed encoding issues for file uploads to S3. - - - -## v1.2.6 ---- - -### Features -- Azure Redis support for cache with auth modes including Entra and Managed Identity. -- HTTPS Proxy support for all the external calls & HTTPs communication between PODs. -- Added support for virtual key inclusion for custom log if passed in headers. -- Entra and Managed Identity support for Azure Log Store. - -### Fixes and Improvements -- Support Batch output from gateway directly. - - - -## v1.2.5 ---- - -### Fixes and Improvements -- Update dependencies for better performance for file uploads/downloads. - - - -## v1.2.4 ---- - -### Fixes and Improvements -- Fixed issue with custom batches missing cost calculation for some provider models - - - -## v1.2.3 ---- - -### Fine-tuning and Batch Processing -- Added support for configurable `FINETUNE_STATUS_CHECK_INTERVAL` for provider fine-tuning status check operations. -- Added support for configurable `BATCH_STATUS_CHECK_INTERVAL` for provider batch processing status check operations. -- Both values should be in milliseconds. Minimum value is 10000 milliseconds. -- If not provided, will default to 10 seconds. - - - - -## v1.2.2 ---- - -### Observability -- Added support for below Prometheus Counters - - `batch_count` - - `batch_cost` - - `batch_input_tokens` - - `batch_total_tokens` - - `batch_process_time` - - `batch_success_row_count` - - `batch_failure_row_count` - - `batch_row_count` -- With the below labels - - `provider` - - `type` (provider/custom) - -### Fixes and Improvements -- Fixed issue with attributing incorrect created at time stamp for batch processing -- Including error source as `control plane` for management plane failures - - - -## v1.2.1 ---- - -### Data exports -- Added support for Data exports for hybrid deployments. - -### Fixes and Improvements -- Fixed issue with custom batches for small batch files - - - -## v1.2.0 ---- - -### Custom S3 Support -- Added support for `s3_custom` log store option for batches and fine-tunes. - -### Fixes and Improvements -- Fixed issue with STS token generation for AWS. - - - -## v1.1.12 ---- - -### Fixes and Improvements -- Fixed issue with cost calculation for custom batches. - - - -## v1.1.11 ---- - -### Fixes and Improvements -- Fixed issue where queue remains stuck in a queued state during file validation. - - - -## v1.1.10 ---- - -### S3 Upload Improvements -- Added support for passing encryption headers while uploading stream data to S3. -- Added support for both file path and direct value from environment variables for secrets like redis connection. - -### Stream Handling -- Improved stream cleanup for validation processes. - - - -## v1.1.9 ---- - -### File Handling Fixes -- Fixed issue with extra bytes being added to files during processing. - - - -## v1.1.8 ---- - -### File Upload Improvements -- Updated socket timeout for long requests during file uploads to prevent timeouts. - - - -## v1.1.7 ---- - -### Fireworks Fine-tuning Support -- Added support for `Fireworks` fine-tuning operations using Version2. - -### Batch Processing Improvements -- Included response tokens calculation in provider batch output. -- Fixed file loading in memory issues for better performance. - - - -## v1.1.6 ---- - -### S3 SDK Updates -- Upgraded S3 SDK to latest version for fixing issue with S3 streaming. - - - -## v1.1.5 ---- - -### Batch Processing Enhancements -- Improved provider batch output handling. -- Added support for custom batch output paths. -- Increased maximum lines for custom batches to 500k and chunk size to 5MB for better performance. - - - -## v1.1.4 ---- - -### Vertex Embeddings Batches Support -- Added support for `Vertex` batch embeddings. - -### Batch Processing Updates -- Included model information in log objects. -- Implemented custom batch processing output generations. - -### Internal POD to POD HTTPS Support -- Added support for internal POD to POD HTTPS communication. -- This can be enabled by mounting a volume with certificate and key. -- `TLS_KEY_PATH` and `TLS_CERT_PATH` environment variables will be used to fetch the certificate and key from the volume. - - - - -## v1.1.3 ---- - -### Infrastructure Updates -- Streamlined uploaded file location for `Bedrock` operations. - - - -## v1.1.2 ---- - -### Vertex Integration -- Added support for `Vertex` provider options for batches. - -### Infrastructure Updates -- Implemented cluster mode Redis for queues. -- Updated fine-tune status handling. - - - -## v1.1.1 ---- - -### S3 Enhancements -- Made S3 bucket optional for Bedrock batches. -- Added S3 encryption header support for finetunes and batches. -- Implemented SSE file upload support. - -### Logging Improvements -- Added filtering for log exports. -- Implemented end limit for log export records. - -### Performance Optimizations -- Implemented internal memory cache for better performance. - - - -## v1.1.0 ---- - -### Bull Board Integration -- Added Bull Board for visualizing job queues and their status. - -### Batch Job Retry Support -- Implemented retry functionality for batch jobs to handle failures gracefully. - -### Prometheus Metrics Enhancements -- Added Prometheus metrics for batch jobs and fine-tuning operations. - - - - -## v1.0.8 ---- - -### Azure Fine-tuning Support -- Added support for `Azure OpeAI` fine-tuning operations. - - - - -## v1.0.7 ---- - -### Fine-tune v2 -- Implemented version 2 of the [fine-tuning](/aigw/product/ai-gateway/fine-tuning) functionality. - - - -## v1.0.6 ---- - -### Prompt Slug Filter -- Added support for data exports filtering by PromptSlug. - - - - -## v1.0.5 ---- - -### Batch Processing -- Added provider and custom [Batch] (/product/ai-gateway/batches) processing functionality. - - - - -## v1.0.4 ---- - -### Code Quality Improvements -- Fixed dynamic port retrieval from environment variables. - - - -## v1.0.3 ---- - -### Vision Fine-tuning Support -- Added support for vision fine-tuning validation for OpenAI. -- Implemented S3 bucket support for fine-tunes. - -### AWS Integration Improvements -- Fixed assumed role handling for Bedrock fine-tuning dataset URLs. -- Improved S3 bucket path handling for Bedrock fine-tune operations. -- Achieved parity with Enterprise Gateway for data sources. - - - -## v1.0.2 ---- - -### Fine-tuning Enhancements -- Added support for OpenAI job start and Fireworks upload. -- Improved handling of chunk type failures with JSON. - - - -## v1.0.1 ---- - -### Fireworks Fine-tuning Support -- Added support for Fireworks fine-tuning operations. - - - - -## v1.0.0 ---- - -### Initial Release -- Base version of the Data Service with core functionality. - diff --git a/aigw/changelog/enterprise.mdx b/aigw/changelog/enterprise.mdx deleted file mode 100644 index 705462c6..00000000 --- a/aigw/changelog/enterprise.mdx +++ /dev/null @@ -1,3470 +0,0 @@ ---- -title: "Enterprise Gateway" -sidebarTitle: "Enterprise Gateway [2.22.0]" -rss: true ---- - - -Discuss how Prisma AIRS AI Gateway can enhance your organisation's AI infrastructure - - - - -## v2.22.0 - ---- - -### Claude Code Model Discovery via `/v1/models` - -Claude Code and other Anthropic-native clients can now discover models through `/v1/models`. The gateway detects Anthropic-protocol callers via the `anthropic-version` header and returns Anthropic's native model shape (`id`, `type`, `display_name`, `created_at`) instead of the OpenAI-normalized shape, for both single-provider requests and workspace-wide model aggregation. - -[Claude Code Documentation](/aigw/integrations/libraries/claude-code) - -### Context and Service Tier Pricing for Azure - -Azure OpenAI and Azure AI Foundry requests are now costed at the rate matching the request's input token count (short vs. long context) and the service tier it was served on, rather than a single flat rate. Azure OpenAI recognizes `standard`, `batch`, and `priority`; Azure AI Foundry also recognizes `flex`. - -Azure AI Foundry additionally accepts `service_tier` as a request parameter, which was previously dropped before reaching the provider. - -[Azure OpenAI Documentation](/aigw/integrations/llms/azure-openai/azure-openai) · [Azure AI Foundry Documentation](/aigw/integrations/llms/azure-foundry) - -### Provider Updates - -- **Anthropic, Bedrock Mantle & Claude Platform on AWS**: `anthropic-beta` and `anthropic-version` are now picked up from the incoming request headers when they aren't set on the provider or in the config, so client-supplied beta features reach the provider. Anthropic requests also no longer carry an automatic `anthropic-beta: messages-2023-12-15` header. [Anthropic Integration](/aigw/integrations/llms/anthropic) -- **Claude Platform on AWS**: Fixed a streaming regression where SSE events were concatenated into a single unterminated blob, causing strict SSE clients like Claude Code to fall back to non-streaming. [Claude Platform on AWS Integration](/aigw/integrations/llms/claude-platform-aws) -- **Workers AI**: The `thinking` parameter now applies to GLM, QwQ, and Kimi K2.5 models, which expect `enable_thinking` rather than `thinking`. Previously it was silently ignored on these models. [Workers AI Integration](/aigw/integrations/llms/workers-ai) -- **Together AI**: The `thinking` parameter is now supported and translated to Together AI's native `reasoning` field, so you can turn reasoning on or off on models like DeepSeek-V4-Flash using the same syntax as every other provider. [Together AI Integration](/aigw/integrations/llms/together-ai) -- **Azure PII Guardrail**: A new **PII Categories** setting lets you limit detection to specific entity types, so values like organisation names or dates can be left untouched. Leaving it empty continues to detect all default categories. [Azure PII Detection Settings](/aigw/integrations/guardrails/azure-guardrails) -- **Alibaba Content Safety Guardrail**: Added the mainland China service types `query_security_check` and `response_security_check`, alongside the existing international ones. [List of Guardrail Checks](/aigw/product/guardrails/list-of-guardrail-checks) - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Pricing**: Pricing adjustments can now be set per model on an integration, not just for the integration as a whole. A model-level adjustment fully replaces the integration-level one rather than merging with it. [Pricing Adjustments](/aigw/product/model-catalog/pricing-adjustments) -- **Pricing**: Added bundled pricing data for ElevenLabs, so text-to-speech and speech-to-text costs are tracked without a pricing sync. [Model Catalog](/aigw/product/model-catalog) -- **Responses API**: Streaming responses now report reasoning and cached token counts in `usage`, which previously always returned `0`. [Responses API Reference](/aigw/api-reference/responses/create-response) -- **Guardrails**: Checks that are turned off inside a redaction or transformation guardrail — such as PII redaction — are now correctly skipped. Previously they still ran. Guardrail configs referencing an unknown check ID are also now ignored with a warning instead of being passed through. [Guardrails](/aigw/product/guardrails) -- **MCP Gateway**: Fixed a 500 error when loading the OAuth consent screen at `GET /oauth/authorize`. Applies to Prisma AIRS AI Gateway deployments, where the consent screen is served as HTML. [MCP OAuth](/aigw/product/mcp-gateway/authentication/oauth) -- **MCP Gateway**: A repeated OAuth callback no longer shows an `invalid_state` error when the connection has already succeeded. [MCP External OAuth](/aigw/product/mcp-gateway/authentication/external-oauth) -- **MCP Gateway**: OAuth token refresh failures now return the correct `invalid_grant` error code and 400 status per RFC 6749, instead of a generic `unauthorized` 401. [MCP OAuth](/aigw/product/mcp-gateway/authentication/oauth) -- **Security**: Updated dependencies to patch security vulnerabilities - - - - - -## v2.21.0 - ---- - -### ElevenLabs Provider - -ElevenLabs is now available as a provider, so text-to-speech and speech-to-text requests can route through the gateway with full observability and reliability features. Authentication uses ElevenLabs' `xi-api-key` header. Text-to-speech is billed on character count and speech-to-text on transcribed audio duration. - -[LLM Integrations](/aigw/integrations/llms) - -### Gateway-Local JWT Auth: Multiple JWKS URLs & User Attribution - -Gateway-local JWT authentication now accepts a comma-separated list of JWKS URLs, fetching and merging keys from all of them — useful when validating tokens issued by more than one identity provider. - -When a token's email resolves to an existing AI Gateway user in the target workspace, the request is now attributed to that user and inherits their workspace role, instead of being treated as a generic workspace service key. - -[JWT Authentication Documentation](/aigw/product/enterprise-offering/org-management/jwt) - -### Provider Updates - -- **SCX.ai**: Added as a new OpenAI-compatible provider for chat completions -- **Claude Platform on AWS**: Fixed the raw `anthropic-workspace-id` header (and `anthropic-beta`/`anthropic-version`) being dropped when requests use the structured `x-portkey-config` JSON shape instead of individual headers -- **Zhipu**: Usage is now passed through on streaming chat completions -- **Gemini**: Tool result messages are now mapped to the `user` role instead of `function`, matching Gemini's expected conversation format - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **JWT Auth**: Gateway-local JWT authentication now returns a clear 403 when a deployment restricts workspaces but none are configured, or when a resolved workspace can't be found, instead of silently falling back -- **Prisma AIRS Guardrails**: Scan requests now include `source_entity` and `source_entity_type` metadata for improved reporting, and fall back to the default timeout when an invalid (zero, negative, or non-numeric) custom timeout is configured -- **Security**: Updated dependencies to patch security vulnerabilities - - - - - -## v2.20.0 - ---- - -### MCP Gateway Guardrails - -MCP tool calls can now be checked against guardrails before and after execution, reusing the same guardrail checks available for LLM requests. Denied calls return a structured error with the guardrail and check ID for easy debugging. - -UI support is coming soon. - -[MCP Gateway Guardrails Documentation](/aigw/product/mcp-gateway/guardrails) - -### Workload Identity Federation for OpenAI & Anthropic - -OpenAI and Anthropic requests can now authenticate upstream using Workload Identity Federation, exchanging a workload identity token for a short-lived provider credential via OAuth client credentials instead of storing a static API key. - -[LLM Integrations](/aigw/integrations/llms) - -### xAI Image Generation - -The xAI provider now supports image generation, with OpenAI-compatible parameters plus xAI-specific aspect ratio and resolution controls. - -[xAI Documentation](/aigw/integrations/llms/x-ai) - -### Bedrock India Inference Profile & Context-Based Pricing - -Amazon Bedrock requests that use the India cross-region inference profile (model IDs prefixed with `in.`) now resolve to the correct underlying model. - -Bedrock models that charge different rates for short vs. long context are now billed at the tier matching the request's input token count, so long-context requests are no longer priced at the short-context rate. - -[AWS Bedrock Documentation](/aigw/integrations/llms/bedrock/aws-bedrock) - -### Unified Gateway + MCP Server Mode - -Set `SERVER_MODE=unified` to run the AI Gateway and MCP Gateway on a single port. In this mode, MCP endpoints are served under the `/m` path prefix (for example, `https://your-gateway.com/m/...`), while LLM traffic continues to use `/v1/*`. This removes the need to run two services or configure host-based routing to split AI Gateway and MCP Gateway traffic. - -[Self-Hosting Documentation](/aigw/self-hosting/hybrid-deployments/aws/eks) - -### Webhook Guardrails for Proxy Requests - -The webhook guardrail can now run on proxy (`/v1/*`) requests via an opt-in `executeOnProxy` parameter, letting you block passthrough traffic based on a custom webhook check. - -[Guardrails Documentation](/aigw/product/guardrails/list-of-guardrail-checks) - -### Provider Updates - -- **Azure AI Foundry**: Fixed cached input tokens being priced as regular input tokens for models called through `/v1/responses` -- **Azure OpenAI**: Codex models can now be called through the Anthropic-compatible `/v1/messages` endpoint -- **Azure OpenAI**: `entraFederated` and `workload` auth modes no longer send an empty `Authorization` header before token exchange completes -- **Amazon Bedrock**: Fixed a `temperature` validation error for Claude 4+ models called through the Messages API -- **Anthropic**: Long-running streams are no longer dropped during long model prefills — Anthropic's keepalive `ping` events are now passed through when streaming via the OpenAI-compatible `/v1/chat/completions` and `/v1/complete` endpoints -- **Anthropic**: Fixed prompt caching being silently skipped for messages with a single content block when calling Anthropic models through `/v1/responses` -- **fal.ai**: Fixed image-generation pricing falling back to the wrong model name -- **Together AI**: Added `stream_options` support -- **OpenAI**: Fixed a tool-call ID error when converting Anthropic tool calls to the Responses API - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Security**: The JWT guardrail now fails closed by default on invalid signatures, `alg: none`, and expired tokens -- **Reliability**: Fixed `custom_host` leaking from one target to other targets in fallback/load-balance configs -- **JWT Validation**: Changes to your JWT auth configuration now take effect right away instead of waiting for previously cached validation results to expire -- **JWT Validation**: When a token introspection endpoint is configured with a custom `introspectContentType` (such as `application/json`), the request body is now encoded to match that content type instead of always being sent as form-urlencoded. Applies to both the JWT guardrail and MCP Gateway JWT authentication -- **MCP**: Removed a redundant per-user access check for MCP requests authenticated with local JWT auth. Local JWT mode does not carry user attribution, so there was no user to check access for -- **Usage Limits**: Fixed request-count-based usage limits and usage limits on JWT-authenticated keys not being applied — requests continued to be served after the configured limit was reached -- **Prisma AIRS Guardrails**: The scan endpoint is now configurable via the `AIRS_URL` environment variable, and scans now report the model, user, and provider from the live request instead of static config values -- **F5 Guardrails**: Added an opt-in `forwardMetadata` parameter to forward request metadata to the scan request - - - - - -## v2.19.0 - ---- - -### Deepgram Provider - -New Deepgram provider integration for speech-to-text and text-to-speech, with full observability and reliability features. - -[Deepgram Documentation](/aigw/integrations/llms/deepgram) - -### OAuth Client Credentials for Upstream Auth - -OpenAI-compatible integrations can now authenticate to the upstream provider using the OAuth Client Credentials grant (RFC 6749). The gateway acquires and caches the bearer token from your configured token endpoint and forwards it upstream, so no static API key is stored. - -[OpenAI Integration Documentation](/aigw/integrations/llms/openai) - -### Vertex AI OCR - -The `/v1/ocr` endpoint now supports Mistral OCR models served on Google Vertex AI, with pricing tracked per page processed. - -[OCR API Reference](/aigw/api-reference/ocr/create-ocr) · [Vertex AI Documentation](/aigw/integrations/llms/vertex-ai) - -### Qwen: Responses & Messages APIs - -The Qwen provider now supports the Responses and Messages endpoints, along with additional non-OpenAI parameters. - -[LLM Integrations](/aigw/integrations/llms) - -### Headroom Plugin - -New Headroom plugin compresses request context before it reaches the model, reducing input token costs. Configured through the standard guardrails workflow. - -[Headroom Plugin Documentation](/aigw/integrations/guardrails/headroom) - -### Provider-Reported Cost Billing - -Billing now uses the provider-reported cost when available, starting with OpenRouter, for more accurate cost attribution. - -[OpenRouter Documentation](/aigw/integrations/llms/openrouter) - -### MCP Registry Proxy - -The MCP Gateway now proxies MCP Registry requests, exposing registry endpoints (e.g. server listings) directly through the gateway. - -[MCP Registry Documentation](/aigw/product/mcp-gateway/mcp-registry) - -### Meshy & Tripo3D Cost Tracking - -Credit-based cost tracking is now available for the Meshy and Tripo3D 3D-generation providers, with a 24h duplicate-billing protection window for repeated task polls. - -[LLM Integrations](/aigw/integrations/llms) - -### AWS Bedrock STS Session Tags (Env-Gated) - -Per-request AWS session tags can now be passed through to Bedrock via STS, letting you attribute cost and audit trails per application while sharing a single assumed role. This feature is gated behind the `AWS_BEDROCK_STS_SESSION_TAGS_ENABLED` environment variable — set it to `true` on the gateway container to enable session tag passthrough. - -[Session Tags Documentation](/aigw/product/model-catalog/connect-bedrock-with-amazon-assumed-role#attribute-cost-and-usage-with-session-tags) - -### Provider Updates - -- **Anthropic**: `output_config` is now forwarded for Messages requests -- **Anthropic**: `output_tokens_details` is preserved in streaming logs -- **Amazon Bedrock**: Inference-profile prefix is preserved for accurate pricing lookups -- **Amazon Bedrock**: Removed the deprecated `temperature` handling for Claude 4 models -- **Amazon Bedrock (Mantle)**: Fixed Anthropic SSE streaming boundaries -- **Azure AI Foundry**: Additional OpenCode compatibility fixes -- **Vertex AI**: Corrected metadata handling in the Cloudflare environment -- **Dashscope**: `chat_complete_kwargs` are now forwarded to the provider -- **Messages Adapter**: Fixed non-unique tool call IDs and Responses instructions mapping -- **Proxy Requests**: Improved thinking-token attribution - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **MCP**: Issuer URL now follows RFC 8414, extra auth params are supported, and `GET`/`DELETE` requests return `405` since the gateway is stateless -- **Pricing**: Integration `pricing_adjustments` are now applied on the config path -- **Security**: Hardened credential masking in logs and guardrail check handling, and updated dependencies to patch vulnerabilities. - -[MCP Gateway Documentation](/aigw/product/mcp-gateway/mcp-registry) - - - - - -## v2.18.0 - ---- - -### Native OCR Endpoint - -New `POST /v1/ocr` endpoint brings OCR requests under the gateway's retries, fallbacks, load balancing, caching, logging, and pricing. Currently supports Mistral AI and Azure AI Foundry. - -[OCR API Reference](/aigw/api-reference/ocr/create-ocr) · [Mistral AI Documentation](/aigw/integrations/llms/mistral-ai#ocr-document-processing) · [Azure AI Foundry Documentation](/aigw/integrations/llms/azure-foundry#ocr-document-processing) - -### Singulr Guardrail - -New Singulr partner guardrail plugin for request and response scanning, configurable through the standard guardrails workflow. - -[Guardrails Documentation](/aigw/product/guardrails) - -### MCP Tool Call Rate Limits - -Rate limits can now be applied to MCP tool calls, throttling tool invocation volume per the configured limits. - -[MCP Rate Limits Documentation](/aigw/product/mcp-gateway/rate-limits) - -### Guardrail Header Forwarding - -Guardrail checks can now be configured to forward specific client request headers (e.g. `x-session-id`, `traceparent`) to their upstream provider. Credential headers are never forwarded. - -[Guardrails Documentation](/aigw/product/guardrails) - -### Provider Updates - -- **OpenAI & Azure OpenAI**: `prompt_cache_options` parameter can now be passed through to control prompt caching behaviour -- **Azure AI Foundry**: Input items sent via the Responses API now get an explicit `type: message` field, fixing rejected requests from clients (e.g. OpenCode) that omit it -- **Bedrock (Mantle)**: OpenAI models now support the native Messages API, gated behind the `use-responses-api-2026-07-30` beta header -- **Bedrock (Mantle)**: `propertyNames` is stripped from tool schemas for Gemma models to prevent validation errors - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- Updated dependencies to patch security vulnerabilities. - - - - - -## v2.17.0 - ---- - -### Databricks Responses API Support - -Databricks provider now supports the Responses API, enabling OpenAI-compatible Responses format for Databricks-served models. - -[Databricks Documentation](/aigw/integrations/llms/databricks) - -### Anthropic Extended Beta Parameters - -New Anthropic beta parameters can now be passed through to the Anthropic provider: `inference_geo`, `diagnostics`, `fallbacks`, and `context_management`. These parameters enable geo-routing, diagnostic telemetry, fallback behaviour, and managed-context workflows for Claude models. - -[Anthropic Documentation](/aigw/integrations/llms/anthropic) - -### Beta Flag: Use Responses API 2026-07-30 - -New beta flag `use-responses-api-2026-07-30` enables the latest Responses API schema version. Pass this via the `x-portkey-beta` header to opt into the July 2026 Responses API contract. - -[Beta Features Documentation](/aigw/product/ai-gateway/beta-features) - -### Provider Updates - -- **Anthropic**: `thinking_tokens` are now correctly mapped to OpenAI's `reasoning_tokens` in chat completions usage for unified token tracking across providers -- **Segmind**: Security improvements for the endpoint -- **Vertex AI**: Fixed `anyOf` sibling key rejection when using JSON schemas with alternatives -- **Azure**: Improved proxy path validation for `openai.azure.com` URLs -- **Tavily**: Updated client name configuration - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Improvements to rate limit calculation**: reducing unnecessary rate limit consumption for metadata operations -- **Model Config Mapping**: Provider configurations are now correctly mapped to `model_config` for consistent config resolution - - - - - -## v2.16.0 ---- - -### Messages API to Responses API Routing - -Anthropic Messages API requests can now be routed to models that use the OpenAI Responses API — such as OpenAI's codex models — with request and response transformed automatically in both directions. - -[Responses API Documentation](/aigw/api-reference/responses/create-response) - -### Lightning AI Provider - -New Lightning AI provider integration for chat completions. - -[LLM Integrations](/aigw/integrations/llms) - -### Akto Guardrails - -The Akto guardrail is now available for request and response scanning. - -[Akto Documentation](/aigw/integrations/guardrails/akto) - -### MCP Gateway: User Attribution in Logs - -MCP logs now attribute requests to a user, derived from OAuth claims or the API key name, for clearer per-user visibility. - -[MCP Observability Documentation](/aigw/product/mcp-gateway/observability) - -### Provider Updates - -- **Vertex AI**: Added Qwen and xAI Model Garden routing -- **Workers AI**: Added reasoning support for K2.7 and GLM-5.2 -- **Anthropic**: OpenAI `service_tier` now maps to Anthropic's native service tier, avoiding validation errors on unsupported models -- **Modal**: Added support for `chat_template_kwargs` -- **Amazon Bedrock**: Fixed Mantle model routing and Messages API handling, and resolved request validation errors -- **AWS**: Authorization header is now preserved for AWS virtual-key requests -- **Vertex AI**: Cached GCP credentials now refresh before expiry to prevent auth failures - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Pricing**: Accurate cost attribution for OpenAI Responses format, cached tokens, grounded Google Search (per query), and Veo image resolutions -- **MCP Gateway**: Improved handling of empty upstream responses during connection setup - - - - - -## v2.15.0 ---- - -### Fixes and Improvements - -- **Rate Limits**: Post-request rate-limit bucketing now uses the request model instead of the model derived from the response, keeping pre- and post-request buckets consistent for providers that return versioned model names (e.g., Azure) -- **Files**: GCS file endpoints now validate bucket names before use - -[Rate Limits Documentation](/aigw/product/policies/rate-limits) - - - - - -## v2.14.1 ---- - -### Prisma AIRS: Agentic Client Support - -New `scan_scope` and `strip_scaffolding` parameters for the Palo Alto Networks Prisma AIRS guardrail. `scan_scope` controls which messages are scanned (`last_message`, `last_user_message`, `user_messages`, or `all_messages`), while `strip_scaffolding` removes known agent-harness wrappers (e.g., `` blocks, MCP tool instructions) before scanning — reducing false positives for agentic clients like Claude Code, Cursor, and Cline. Tool result content is preserved for injection detection. Fully backwards-compatible with existing configurations. - -[Prisma AIRS Documentation](/aigw/integrations/guardrails/palo-alto-panw-prisma) - - - - - -## v2.14.0 ---- - -### Unified Models Endpoint - -`GET /v1/models` now routes through the gateway's provider-backed model listing handler. The endpoint calls the upstream provider, applies response transforms and returns the normalized model list — replacing the previous pass-through delegation. - -[Models API Documentation](/aigw/api-reference/models/list-models) - -### Guardrails: Soft Deny (HTTP 200) - -New `soft_deny_200` flag on guardrail checks. When enabled, a guardrail denial returns HTTP 200 with the error message formatted as a chat completion response instead of the standard HTTP 446. This keeps sessions alive for clients that treat 4xx status codes as fatal errors (e.g., Claude Code). - -[Guardrails Documentation](/aigw/product/guardrails) - -### Provider Updates - -- **Anthropic**: Multidimensional pricing — cost attribution now differentiates batch (50% discount), fast (speed-priority on Opus models), and standard pricing tiers -- **xAI**: Server-side tool invocation costs (file search, web search) are now tracked for cost attribution -- **Hugging Face**: Fixed model name sanitation - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Guardrails**: Async output guardrails now correctly execute when they are the only output hook configured -- **Proxy**: Provider auth credentials are no longer overwritten by client-forwarded headers on proxy routes -- **MCP Gateway**: Fixed config cache TTL unit mismatch and stale cache invalidation on server recreation -- **Configs**: Config target fan-out is now capped to prevent egress amplification and resource exhaustion from deeply nested configs. Default limits: **100** root-level targets (1,000 for conditional strategy), **50** nested targets per root target - - -Configs exceeding the new target limits will be rejected at request time. If your deployment uses configs with more targets, override the defaults with the `MAX_ROOT_CONFIG_TARGETS`, `MAX_NESTED_CONFIG_TARGETS`, or `MAX_CONDITIONAL_ROOT_CONFIG_TARGETS` environment variables before upgrading. - -- **Caching**: Expired keys are cached slightly longer to reduce management plane network calls during transient expiry -- **Memory**: Improved memory usage and reduced OOM risk under high load -- Updated dependencies to patch security vulnerabilities - - - - - -## v2.13.0 ---- - -### Server-Side MCP Execution (Beta) - -gateway-registered MCP server tools can now be included in the `tools` array of Responses API and Messages API requests using the `@portkey-mcp` prefix. The gateway fetches tool definitions from the MCP server, injects them into the LLM request, and executes tool calls server-side — so providers that don't natively support remote MCP (e.g., AWS Bedrock, Vertex AI) can still use MCP tools. Gated behind the `x-portkey-beta: server-side-mcp-2026-06-01` header. - -[Beta Features Documentation](/aigw/product/ai-gateway/beta-features) - -### Alibaba Cloud AI Guardrails - -New `alibaba-cloud-ai` guardrail plugin for content safety checks, configurable through the standard guardrails workflow. - -[Guardrails Documentation](/aigw/product/guardrails) - -### MCP Gateway: OTel Logs Export - -MCP gateway request logs can now be exported to an OpenTelemetry collector alongside LLM gateway logs. - -[MCP Observability Documentation](/aigw/product/mcp-gateway/observability) - -### OpenTelemetry: Tenant Attributes on Spans - -Gen AI and guardrail OTel spans now carry AI Gateway tenant attributes — org ID, workspace, and environment — so spans can be filtered, grouped, and alerted on per-tenant in your collector without post-processing. - -[OpenTelemetry Documentation](/aigw/product/observability/opentelemetry) - -### Provider Updates - -- **Anthropic**: `thinking.display` parameter is now passed through to the upstream API, enabling clients to control whether extended thinking content is returned in responses -- **Vertex AI**: Added unsupported `thinking token count` beta header to the blocklist, so the header is stripped when forwarding requests to Vertex AI -- **Perplexity AI**: Fixed streaming for Sonar models -- **AI21**: Fixed model name sanitation -- **Workers AI**: Fixed account ID validation -- **Databricks**: Improved serving endpoint name resolution and model routing for accurate pricing and analytics - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Inline Config**: User-passable provider params are now ignored when inline config enforcement is enabled -- **MCP Health**: Health endpoint now returns the gateway version -- **Embeddings**: Response metadata (hooks, usage) is preserved when redacting embedding vectors from logs -- **Fireworks AI Pricing**: Cache token usage is now tracked for cost attribution -- Updated dependencies to patch security vulnerabilities - - - - - -## v2.12.0 ---- - -### OpenTelemetry: Guardrail Execution Spans - -Guardrail checks now emit dedicated OTel spans under the `portkey.guardrail.*` namespace, linked to the parent generation trace, with attributes for phase, verdict, action, deny, async, and per-check events. Gated behind the `EXPERIMENTAL_GEN_AI_OTEL_TRACES_ENABLED` env flag. - -[OpenTelemetry Documentation](/aigw/product/observability/opentelemetry) - -### MCP Gateway: Token Metering - -Token usage from MCP requests (prompts, completions, and tools) is now tracked and reported alongside LLM usage in analytics, logs. - -[MCP Observability Documentation](/aigw/product/mcp-gateway/observability) - -### Agents: Token Metering - -Token usage from agents endpoints (messages and tasks) is now tracked and reported alongside LLM usage in analytics, logs. - -[Observability Documentation](/aigw/product/observability) - -### Azure AI Inference: MAI Image Generation - -Image generation and editing are now supported for MAI models (e.g., MAI-Image-2.5) on Azure AI Inference. - -[Azure AI Foundry Documentation](/aigw/integrations/llms/azure-foundry) - -### Vertex AI: MoonshotAI Models - -Vertex AI integration now supports the `moonshotai` model provider. - -[Vertex AI Documentation](/aigw/integrations/llms/vertex-ai) - -### Provider Updates - -- **GLM**: Chat completions now route system messages to the upstream `system` block instead of the `messages` array, matching the schema GLM models expect -- **Fireworks AI**: Added support for `stream_options`, forwarded `thinking` parameters to the upstream API, and extended multi-dimensional pricing to Fireworks-hosted models -- **Cloudflare Workers AI**: Fixed token tracking for messages-format requests - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Anthropic**: Batch results download URLs validated before fetching -- **AWS Batches**: Bucket names validated before use on batch routes -- **Metrics endpoint**: The `/dataservice/metrics` endpoint is now gated behind the `ENABLE_PROMETHEUS` env flag, matching the `/metrics` endpoint behaviour -- **Inline-block**: `x-portkey-virtual-key` header is allowed when inline configs are blocked -- **Security**: Strengthened request validation across inline-config blocking, batch routes, and provider auth -- Updated dependencies to patch security vulnerabilities - - - - - -## v2.11.2 ---- - -### AI Gateway Models Endpoint Override - -A new `x-portkey-fetch-integrated-models` header and `fetch_integrated_models` config field force `GET /v1/models` to return gateway-configured models even when provider, virtual-key, or config routing signals are present. Without the flag, behaviour is unchanged — provider signals still proxy `/v1/models` to the upstream provider. - -[Models API Documentation](/aigw/api-reference/models/list-models) - -### MCP Gateway: Metrics-Only Logging - -The MCP gateway now honors the metrics-only logging mode (matching the LLM gateway). Request and response bodies and headers are dropped from persisted logs while metrics, routing metadata, and span info are preserved. Response bodies are retained on failures so debugging stays possible. - -[MCP Observability Documentation](/aigw/product/mcp-gateway/observability) - -### Log Store Circuit Breaker [BETA] - -A circuit breaker now protects log writes to S3-compatible object storage. After a configurable number of consecutive failures, the circuit opens and fast-fails log writes for a backoff period instead of retrying on every request — preventing a log store outage from degrading gateway throughput. The circuit probes the endpoint after the backoff elapses and closes automatically on success. - -Configurable via three environment variables on the gateway container: `LOG_STORE_CONNECTION_FAILURE_THRESHOLD` (default: 5 failures), `LOG_STORE_CONNECTION_BACKOFF_MS` (default: 60 s), and `LOG_STORE_CONNECTION_TIMEOUT_MS` (default: 15 s per call). - -### Provider Updates - -- **Amazon Bedrock**: ARN suffixes are ignored during model details lookup, so requests targeting Bedrock model ARNs resolve pricing and validation against the underlying model -- **Vertex AI**: Upstream Vertex AI error details are now surfaced on failures instead of `undefined` - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Stored logs**: Improvements to logging to reduce size of final log -- **Security**: Hardened validations on the Feedback APIs -- Updated dependencies to patch security vulnerabilities - - - - - -## v2.11.1 ---- - -### Amazon Bedrock: Remove `context_management` Pass-Through - -The `context_management` pass-through added in v2.11.0 has been removed for Bedrock Anthropic models. - -[Amazon Bedrock Documentation](/aigw/integrations/llms/bedrock/aws-bedrock) - -### Amazon Bedrock Guardrails: Action Handling - -Bedrock Guardrails now respect the `action` field returned by the guardrail evaluation. Previously, non-blocking actions could still trigger request blocking; the gateway now only blocks when the action is explicitly `blocked`. - -[Guardrails Documentation](/aigw/product/guardrails) - -### Provider Updates - -- **Amazon Bedrock**: Rerank endpoint pricing is now tracked, so rerank requests through Bedrock report accurate cost attribution -- **Amazon Bedrock**: Fixed response transforms for the Messages API on default Converse providers - -[Amazon Bedrock Documentation](/aigw/integrations/llms/bedrock/aws-bedrock) - -### Fixes and Improvements - -- **SSRF**: Improved SSRF validation for air-gapped deployments -- **Security**: Strengthened header validation for first-party requests - -[Custom Hosts Documentation](/aigw/product/ai-gateway/custom-hosts) - - - - - -## v2.11.0 ---- - -### MiniMax Provider - -New `minimax` provider for chat completions with MiniMax's hosted models. - -[Providers Documentation](/aigw/integrations/llms) - -### Modal: Embeddings - -The Modal integration now supports `/v1/embeddings` alongside chat completions, so embedding workloads on Modal-hosted models route through the unified gateway. - -[Modal Documentation](/aigw/integrations/llms/modal) - -### Anthropic: Context Management - -Anthropic chat completions now pass `context_management` through to the upstream API, enabling memory tools and managed-context workflows on Claude models. - -[Anthropic Documentation](/aigw/integrations/llms/anthropic) - -### Anthropic Skills: File Download Fix - -`GET /v1/files/{file_id}` and `GET /v1/files/{file_id}/content` now route correctly for Anthropic, and binary responses (e.g., PPTX generated via Skills + Code Execution) pass through untouched. Previously these requests fell back to list-route behaviour and corrupted binary downloads. - -[Anthropic Files Documentation](/aigw/integrations/llms/anthropic/files) - -### Claude Code OAuth - -When a request from Claude Code/CLI uses an OAuth bearer token, the gateway now forwards `Authorization` to Anthropic instead of injecting `x-api-key`. This unblocks Claude Code's in the newer versions of Claude Code. - -[Claude Code Documentation](/aigw/integrations/libraries/claude-code-anthropic) - -### Responses API: Anthropic Thinking Blocks in Streaming - -Streaming Responses API responses for Anthropic models now include reasoning summary blocks. The streaming transform was dropping `thinking` content blocks, so reasoning never reached clients on streamed requests. - -[Responses API Documentation](/aigw/product/ai-gateway/responses-api) - -### Enforce Inline Config - -A new flag rejects requests that pass an inline `config` JSON object instead of a saved config slug. Use it to require all routing decisions to flow through governed, named configs. - -[Enforce Default Config Documentation](/aigw/product/administration/enforce-default-config) - -### SSRF Hardening (continued) - -Builds on the v2.10.0 SSRF work with broader coverage of edge cases. - -- `TRUSTED_CUSTOM_HOSTS` now matches subdomains of allowlisted hosts -- Provider URL header validation (custom host headers, forward headers) and per-request URL validation extended to Fireworks and Cohere -- Batch output URLs validated before download; concurrency on outbound URL validation capped at 5 -- improvements to`customHost` header lookup -- SSRF hardening improvements - -[Custom Hosts Documentation](/aigw/product/ai-gateway/custom-hosts) - -### JWT Authentication - -- **JWT auth**: Management plane requests now forward the actual JWT instead of the gateway's effective auth token, fixing MCP gateway failures behind JWT auth -- **Local JWT scopes**: Locally injected JWT scopes are honored for upstream MCP authorization decisions - -[MCP JWT Authentication Documentation](/aigw/product/mcp-gateway/authentication/jwt) -[JWT Authentication Documentation](/aigw/product/enterprise-offering/org-management/jwt) - -### Provider Updates - -- **Amazon Bedrock**: Image URL handling refactored to support S3 URIs for image, video, and document inputs on Amazon Nova models -- **Amazon Bedrock**: Fixes to assumed-role authentication for Bedrock models. -- **Azure OpenAI**: Fixes to cost attribution for Azure video create requests. - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **`notNull` guardrail**: Skip the null-content check when the response carries tool calls, so tool-only assistant turns no longer trip the guardrail -- Updated dependencies to patch security vulnerabilities - - - - - -## v2.10.0 ---- - -### SSRF Hardening - - -- In production deployments, arbitrary custom hosts are no longer trusted by default -- URL and hostname length limits (2048 characters) apply, and localhost is blocked unless explicitly allowed -- If you use private or internal `custom_host` targets today, set `TRUSTED_CUSTOM_HOSTS` in your gateway environment (comma-separated hosts) before upgrading to avoid disruption. - -[Custom Hosts Documentation](/aigw/product/ai-gateway/custom-hosts#trusted-custom-hosts-allowlist) - - -- Outbound provider requests are now validated at DNS resolution time: resolved addresses are checked before a connection is opened, and traffic to private networks, reserved ranges, and cloud metadata endpoints is blocked. This closes SSRF and DNS rebinding gaps for upstream LLM and integration calls. -- Custom host and provider URL checks are hardened per RFC 1918, with stricter rules for link-local addresses, cloud metadata hosts (including subdomains), common internal service ports, and tunneled or ambiguous IPv6 forms. IPv6 handling is more reliable for mixed and compressed address formats. - -### Claude Platform on AWS - -New `claude-platform-aws` provider for Anthropic's full-platform capabilities through AWS, with `accessKey`, `assumedRole`, and `serviceRole` auth modes against an Anthropic AWS workspace. - -[Claude Platform (AWS) Documentation](/aigw/integrations/llms/claude-platform-aws) - -### Cato Networks Guardrail - -New Cato Networks guardrail plugin for AI security inspection of LLM prompts and responses via Cato's analyze API, with monitor, anonymize, or block actions. - -[Cato Networks Documentation](/aigw/integrations/guardrails/cato) - -### Multi-Dimensional Pricing - -Pricing engine now resolves cost across multiple dimensions (region, tier, and other model attributes) for OpenAI, Amazon Bedrock, and Vertex AI, so usage is attributed to the right rate when the same model has different prices per region or service tier. - -[Pricing Adjustments Documentation](/aigw/product/model-catalog/pricing-adjustments) - -### Prometheus: Config Name and API Key Name Labels - -Two new env flags add the request's config slug and API key name as labels on every custom metric: `PROMETHEUS_INCLUDE_CONFIG_NAME_LABEL` and `PROMETHEUS_INCLUDE_API_KEY_NAME_LABEL`. Both default to `false`. `config_name` falls back to `N/A` when an inline config object is sent rather than a slug. - -[Prometheus Metrics Documentation](/aigw/self-hosting/prometheus-metrics) - -### Proxy Path: `allowedRequestTypes` Guardrail - -The `default.allowedRequestTypes` guardrail now runs as a `beforeRequestHook` on proxy requests, so allowed/blocked endpoint policies apply to the proxy path. `otel` is added to the plugin manifest allow list so OpenTelemetry export can be permitted explicitly via guardrail config. - -[Guardrails Documentation](/aigw/product/guardrails) - -### CrowdStrike AIDR: Request Metadata Forwarding - -`guard_chat_completions` requests to CrowdStrike AIDR now include `user_id`, `user_name`, `llm_provider`, and `model` extracted from request context metadata for richer policy decisions and audit. - -[CrowdStrike AIDR Documentation](/aigw/integrations/guardrails/crowdstrike-aidr) - -### Vertex AI: zai-org Models - -Vertex AI integration adds support for the `zai-org` model provider on chat completions, including pricing resolution and provider validation. - -[Vertex AI Documentation](/aigw/integrations/llms/vertex-ai) - -### Vertex AI: `service_tier` Header Forwarding - -OpenAI-style `service_tier` is now forwarded as the upstream Vertex header for tiered inference and usage attribution. - -[Vertex AI Documentation](/aigw/integrations/llms/vertex-ai) - -### Provider Updates - -- **Cloudflare Workers AI**: Handle the OpenAI-compatible response format used by newer models (Kimi K2.5, K2.6) — chat and text completions, streaming chunks, tool calls, usage, and `finish_reason` now parse from both legacy (`result.response`) and new (`result.choices[].message.content`) shapes -- **Together AI**: Capture usage from streaming chunks so cost is reported correctly on streamed responses -- **Anthropic Messages**: Don't close the `tool_use` block on `finish_reason` — tool-call argument deltas arriving after the finish chunk are no longer dropped - -[Providers Documentation](/aigw/integrations/llms) - -### Fixes and Improvements - -- **Rate limit errors**: Workspace, API key, virtual key, and integration rate-limit errors now surface human-readable slugs instead of masked IDs (and fix `exceded` → `exceeded`) -- Updated dependencies to patch security vulnerabilities - - - - - -## v2.9.0 ---- - -### `default_params` and `drop_params` for Config Targets - -Config targets accept two new fields alongside `override_params`: `default_params` (inject only when the client hasn't set the field) and `drop_params` (remove fields by dot-notation path, including nested keys and array indices). Execution order: `default_params` → `override_params` → `drop_params`. All three inherit through nested targets. - -[Configs Documentation](/aigw/product/ai-gateway/configs#default-and-drop-params) - -### New Providers - -- **Fal AI** (`fal-ai`): image, video, and audio generation models -- **BytePlus** (`byteplus`): chat completions -- **Scenario AI** (`scenario-ai`): proxy passthrough with Basic auth (`apiKey` + optional `scenarioApiSecret`) - -[Providers Documentation](/aigw/integrations/llms) - -### Lakera Guard - -New guardrails partner via Lakera's `/v2/guard` API. Supports `beforeRequestHook` and `afterRequestHook`, redacts PII when only `pii/*` detectors fire, blocks on any other policy hit. Optional `project_id` selects a Lakera policy; optional `apiBase` targets regional endpoints. - -[Guardrails Documentation](/aigw/product/guardrails) - -### Vertex AI: Multi-Region Endpoints - -Vertex AI requests can now target per-region endpoints for batches, embeddings (Gemini Embedding 2), Anthropic Claude on Vertex, and Gemini models. - -[Vertex AI Documentation](/aigw/integrations/llms/vertex-ai) - -### Vertex AI: Anthropic Batches - -The unified `/v1/batches` API supports Anthropic Claude models on Vertex AI. - -[Vertex AI Batches Documentation](/aigw/integrations/llms/vertex-ai/batches) - -### Azure AI Foundry: Native Responses API (Agents) - -Azure AI Foundry can use the native `/v1/responses` endpoint behind a flag, including agent references. Chat completions, Responses API transformation, agent-reference Responses, and Anthropic chat completions all flow through the native path. `/v1/messages` is not supported for Anthropic on Azure. - -[Azure AI Foundry Documentation](/aigw/integrations/llms/azure-foundry) - -### Amazon Bedrock Provider: `service_tier` - -Bedrock provider requests can pass OpenAI-style `service_tier` where supported, with correct mapping and usage attribution for tiered Bedrock inference. - -### Realtime: Custom Host Routing for WebSockets - -Realtime WebSocket connections honor the configured custom host slug, matching HTTP routing. Previously routed to the provider default host regardless of the custom host setting. - -[Realtime API Documentation](/aigw/product/ai-gateway/realtime-api) - -### MCP Gateway - -- **Capability enforcement on `tools/call`**: `tools/call`, `prompts/get`, and `resources/read` repopulate the disabled-capabilities set from the management plane on cold cache instead of silently allowing requests through -- **External fetch path**: MCP upstream calls use the external-service fetch path for consistent SSRF and TLS controls -- **Scope handling**: Fixed scope handling for upstream MCP requests - -[MCP Gateway Documentation](/aigw/product/mcp-gateway) - -### Pricing - -Synced model pricing across providers. - -### Fixes and Improvements - -- **Perplexity streaming**: Skip SSE keepalive comment lines (e.g., `: ping - `) before JSON parsing — long Sonar streams no longer end without a `data: [DONE]` event (SEV1) -- **JWT/JWKS**: External JWKS fetches go through the external-service agent so `HTTPS_PROXY` / `TLS_CA` are honored -- **TLS**: `TLS_CA` applied for outbound calls to external services -- **Stream parsing**: SSE comment chunks skipped to avoid spurious parse errors -- Updated dependencies to patch security vulnerabilities. - - - - - -## v2.8.0 ---- - - -**Deploy data-service v1.7.0+ before** upgrading this gateway. - - -### Tavily Online Guardrail Plugin - -New Tavily-backed online guardrail plugin for web-grounded checks in your guardrail pipelines. - -### Vertex AI: Gemini TTS (`createSpeech` / `generateSpeech`) - -Vertex AI / Gemini text-to-speech is available through the unified speech APIs — `createSpeech` and `generateSpeech` — so you can drive TTS through the same gateway routes as other multimodal workloads. - -### Amazon Bedrock: `service_tier` - -Bedrock requests can specify OpenAI-style `service_tier` where supported, with correct mapping and billing alignment for tiered Bedrock inference. - -### Amazon Bedrock Mantle: 1-Hour Cache Cost - -Prompt cache pricing for Bedrock Mantle now includes the additional cost path for **1-hour** cached context, improving cost attribution for longer-lived cache windows. - -### Azure: Workload Identity (OpenAI / Foundry and Redis) - -- **LLM routes**: Workload identity coverage is extended for Azure OpenAI and Azure AI Foundry, including gateway-specific configuration paths for federated, keyless access. -- **Azure Cache for Redis**: The gateway adds **workload identity** as an authentication option for Azure Redis, alongside existing token and connection-string flows. - -### Usage Limit Policies: New Policy Type - -Administrators can attach a **new usage limit policy type** for finer-grained control over consumption limits (aligned with management-plane policy models). - -### MCP Gateway: CORS Controls - -New environment flags let you tune **CORS response headers** for MCP endpoints in self-hosted deployments. The default MCP CORS configuration **no longer sets `Access-Control-Allow-Credentials: true`**, reducing accidental credentialed cross-origin exposure when you rely on simple bearer flows. - -### Prompt Security - -Prompt security checks are **extended** with additional coverage and configuration hooks (provider- and deployment-specific hardening). - -### Provider Beta Headers - -Beta header forwarding and filtering refreshed across providers so new upstream beta programs reach the right integration without manual header surgery. - -### Circuit Breaker Target Resolution - -Circuit breaker extraction now supports **provider slug references** when resolving targets, so breaker rules stay stable across renames and shared provider catalogs. - -### Pricing and Cost Fixes - -- **Anthropic**: Corrected **prompt cache** pricing in cost calculations. -- **Catalog**: Pricing tables synced with the latest provider rates. - -### Compliance: FIPS (DHI) - -Deployment hardening for **FIPS** environments (DHI-related paths) improves compatibility with strict crypto policies in regulated installs. - -### Fixes and Improvements - -- **Messages API streaming**: Fixed **content block** tracking in the stream adapter so multi-block assistant output is logged and forwarded consistently. -- **Security**: Enhanced authorization checks for internal service communication. -- **Data plane / custom host**: Enforces a configurable **maximum response size** from custom-host upstreams to protect the gateway from oversized payloads. -- **Redis**: **Auth refresh** fixes for long-lived connections; token rotation and reconnect behaviour are more reliable. - - - - - -## v2.7.1 ---- - -### Anthropic: Catch `overloaded_error` on Streaming Responses - -Anthropic can return `overloaded_error` as the first chunk of a streaming (`200 OK`) response, which previously slipped past retry/fallback logic. When enabled on an Anthropic integration, the gateway now inspects the first streaming chunk and — if it's an `overloaded_error` — converts the response to HTTP **529**, which is a [retriable status code](/aigw/product/ai-gateway/automatic-retries), so your configured retries and fallbacks trigger as expected. - -Opt-in via the integration flag `anthropicInspectStreamForOverloadedError: true` (applies to Anthropic integrations only). - - - - - -## v2.7.0 ---- - -### Vertex AI Rerank - -Vertex AI is now supported on the AI Gateway's unified `/v1/rerank` endpoint, backed by Google's Discovery Engine ranking API (`semantic-ranker-default`, `semantic-ranker-fast`, `semantic-ranker-512`). - -[Vertex AI Rerank Documentation](/aigw/integrations/llms/vertex-ai/rerank) - -### Amazon Bedrock Mantle (OpenAI-Compatible Bedrock Endpoint) - -New `bedrock-mantle` provider for AWS's OpenAI-compatible Bedrock inference engine (`bedrock-mantle.{region}.api.aws`). Supports `/v1/chat/completions`, `/v1/responses`, and Anthropic-native `/v1/messages` via the AI Gateway's unified API, with `apiKey`, `assumedRole`, and `serviceRole` auth modes. - -[Amazon Bedrock Mantle Documentation](/aigw/integrations/llms/bedrock-mantle) - -### Anthropic `speed` Parameter on `/messages` - -`/v1/messages` requests now pass through Anthropic's top-level `speed` parameter (`fast` / `standard`). OpenAI's `service_tier` → Anthropic `speed` mapping continues to work for `/chat/completions`. - -[Anthropic `service_tier` → `speed` mapping](/aigw/integrations/llms/anthropic#service-tier) - -### Centralized `anthropic-beta` Header Filtering - -`anthropic-beta` headers are now filtered and remapped per provider (Azure AI, AWS Bedrock, Bedrock Converse, Vertex AI, Databricks). Unsupported beta flags are stripped before being forwarded upstream; provider-renamed betas (e.g., Bedrock's `tool-search-tool-2025-10-19`) are rewritten automatically. Unknown beta values are passed through for forward compatibility. - -### Required Metadata Key-Value Pairs: `matchType` - -The [Required Metadata Key-Value Pairs](/aigw/product/guardrails/list-of-guardrail-checks#access-control-management) guardrail now accepts a `matchType` parameter to control how values are compared: - -| `matchType` | Behaviour | -|---|---| -| `exact` (default) | Metadata value must equal the expected value | -| `contains` | Metadata value must contain at least one expected value | -| `containsAll` | Metadata value must contain all expected values | -| `regex` | Expected values are treated as regex patterns | - -### Output Guardrails with `strict_openai_compliance` - -Output guardrails (`after_request_hooks`) now execute even when `x-portkey-strict-open-ai-compliance: true`. Hook results are still only appended to the stream in non-strict mode, but the guardrail checks themselves run for every request. - -### Google / Vertex Embeddings for Semantic Cache - -`SEMANTIC_CACHE_EMBEDDING_PROVIDER` now accepts `google` and `vertex-ai` in addition to `openai`. Embedding requests are routed to the Gemini / Vertex embeddings API to populate the semantic cache. - -[Semantic Cache Documentation](/aigw/product/ai-gateway/cache-simple-and-semantic#set-up-semantic-caching-self-hosted) - -### MCP Integrations: Secret Manager Support - -MCP server integrations can now reference centrally-managed secrets via `secret_mappings` in the MCP integration config. The gateway resolves the referenced secret values at connection time before forwarding to upstream MCP servers. - -### Secret Mappings: JSON-typed Values - -Secret mappings now accept a `value_format` field (`string` | `json`). `json` parses stringified JSON (or passes objects through) into the target field — useful for injecting structured credentials like service-account JSON into provider configs. - -### Pricing Adjustments (Integration-level Multipliers) - -Integrations can now carry a `pricing_adjustments.multiplier` config. Per-token multipliers (request / response / cached / audio / reasoning, plus `default`) are applied on top of the base pricing during cost calculation — enabling markup / discount pricing per integration without editing the shared pricing catalog. - -### Signed Container Images (Cosign / Sigstore) - -Published Docker images (release and prerelease) are now signed with [Sigstore Cosign](https://docs.sigstore.dev/cosign/overview) and can be verified with the AI Gateway's public key before deployment. - -### Fixes and Improvements - -- **Vertex AI (Anthropic on Vertex)**: Fixed `cache_control` sanitization — unsupported block-level cache markers are stripped while preserving valid ones on `/messages` and `/chat/completions` -- **Vertex AI (Anthropic on Vertex)**: `output_config` is dropped from `/messages` requests (Vertex rejects it); newly-denylisted `anthropic-beta` values are filtered before being forwarded -- **Anthropic on Azure / Copilot**: `max_tokens: null` now falls back to the provider default instead of being forwarded as `null`; `top_p` is automatically dropped when both `top_p` and `temperature` are provided -- **DeepSeek**: Fixed `reasoning_effort` transform — now correctly sent as `thinking: { type: "enabled" | "disabled" }`, and the original `reasoning_effort` is also forwarded -- **Gemini / Vertex**: `finish_reason` is now reported as `tool_calls` (instead of `stop`) when the response contains tool calls -- **Gemini**: `thought` flag is preserved on inline-image parts when strict OpenAI compliance is disabled -- **`/messages` adapter**: Thinking blocks without a `signature` are stripped when adapting cross-provider responses (prevents Claude rejecting conversation history from Gemini / OpenAI models) - - - - - -## v2.6.2 ---- - -### MCP Gateway: Header Renaming in `forward_headers` - -`forward_headers` entries now accept `{ from, to }` objects to rename a client header before it reaches the upstream. Protected headers remain blocked on both sides. - -[MCP Forwarding Headers Documentation](/aigw/product/mcp-gateway/authentication/forwarding-headers) - -### Azure Entra ID: Federated Authentication - -New `entraFederated` auth mode for Azure OpenAI and Azure AI Foundry exchanges an AWS web identity token for an Entra ID access token — keyless Azure access from AWS workloads (EKS/IRSA). - -[Azure Authentication Modes Documentation](/aigw/integrations/llms/azure-openai/authentication) - -### Gemini Embedding 2 Preview Support - -`gemini-embedding-2-preview` is now available on the unified `/v1/embeddings` route for Vertex AI with text, image, video, and audio input. - -[Vertex AI Embeddings Documentation](/aigw/integrations/llms/vertex-ai/embeddings) - -### Fixes and Improvements - -- **Gemini / Vertex thinking**: Restored `thinking_budget` forwarding on Vertex; `thinking.type = "enabled"` with a zero `budget_tokens` now correctly disables thinking -- **Pinecone vector store**: Fixed client initialization to use the configured `vector_store_api_key` -- **Passthrough configs**: Clearer errors when a [passthrough target](/aigw/product/ai-gateway/configs#passthrough-targets) can't resolve a provider from the incoming request -- **Revert**: Removed the multi-layer Vertex AI pricing that caused zero cost attribution in v2.6.1 - - - - - -## v2.6.1 ---- - -This release contains a bug which led to zero cost attribution for most requests. - -This has been fixed in v2.6.2 - - -### Anthropic Batch API via `/v1/batches` - -Run asynchronous Claude jobs through the AI Gateway's unified `/v1/batches` endpoints. - -[Anthropic Batches Documentation](/aigw/integrations/llms/anthropic/batches) - -### Mask Sensitive Headers in Logs - -New `x-portkey-sensitive-headers` header (SDK aliases: `sensitive_headers` / `sensitiveHeaders`) masks matching header values in request, response, and OTel trace logs. - -[Sensitive Headers Documentation](/aigw/api-reference/inference-api/headers#mask-sensitive-headers-in-logs) - -### OpenAI-compatible `reasoning_effort` for Bedrock - -`reasoning_effort` (`minimal` / `low` / `medium` / `high` / `none`) now works for Bedrock Claude reasoning models — mapped to adaptive thinking on Claude 4.6, and to `thinking.budget_tokens` on earlier reasoning models. - -[Bedrock `reasoning_effort` Documentation](/aigw/integrations/llms/bedrock/aws-bedrock#using-reasoning-effort-parameter) - -### Agent Gateway (Preview) - -New experimental proxy route `/v1/agent/:agentServerId/*` forwards requests to A2A-registry agent servers with AI Gateway auth and observability. Reach out to preview. - -### Fixes and Improvements - -- **Vertex AI metadata labels**: Hardening for the v2.6.0 metadata-to-labels feature — empty values are skipped, keys/values are lowercased and trimmed to Google Cloud's 64-char limits, and keys are sanitized -- **Structured outputs adapter**: `/messages` requests with `output_config.format.type = "json_schema"` now correctly map to `response_format.json_schema` when adapted to `/chat/completions` -- **Streaming logs**: AWS event-stream chunk parse failures are now logged with the truncated payload and error details instead of being silently swallowed -- **Hugging Face**: Default base URL updated from `api-inference.huggingface.co` to `router.huggingface.co` -- **Error messages**: Invalid virtual key errors now surface the specific key that failed validation - - - - - -## v2.6.0 ---- - -### Vertex AI: Metadata to Labels Mapping - -The AI Gateway [metadata](/aigw/product/observability/metadata) is now automatically mapped to Vertex AI resource labels across `chat/completions`, `embeddings`, `batches`, and `fine-tuning` — unlocking Google Cloud–native cost attribution. Keys are sanitized and prefixed (default `pk_gateway_`, configurable via `METADATA_MAP_KEY_PREFIX`). - -[Vertex AI Custom Metadata Labels](/aigw/integrations/llms/vertex-ai#custom-metadata-labels) - -### Custom Host per Custom Model - -Custom models in the [Model Catalog](/aigw/product/model-catalog/custom-models) can now carry their own `custom_host` in the model config, letting different custom models under the same integration route to different upstreams. A header-level `custom_host` still takes priority. - -### Anthropic Enhancements - -- **Data URL file inputs**: `/chat/completions` accepts base64 data URLs (`data:;base64,`) in `file.file_data` and forwards them as Anthropic document content -- **`cache_control` passthrough**: `/chat/completions` now forwards a top-level `cache_control` parameter to Anthropic, matching `/messages` behaviour. For block-level caching, continue using [prompt caching](/aigw/integrations/llms/anthropic/prompt-caching) -- **Strict structured outputs**: Object-type `response_format` JSON schemas automatically get `additionalProperties: false` - -### MCP Gateway Reliability - -- **Upstream token refresh**: Pooled MCP connections now refresh expired upstream OAuth tokens silently mid-session; upstream-invalidated sessions trigger pool invalidation and retry. New log attributes: `gateway_version`, `mcp.auth.upstream_token_refreshed` -- **Configurable request timeout**: `MCP_REQUEST_TIMEOUT_MS` sets the outbound timeout for every MCP call (default: SDK's 60 s) -- **Connection pool**: `MCP_POOL_*` env vars now load with correct precedence; default `maxLifetimeMs` lowered from 30 min → 10 min; improved stream and session cleanup - -### Fixes and Improvements - -- **Perplexity streaming**: Fixed `sonar-deep-research` — SSE parsing handles `[DONE]` and non-JSON chunks, structured deltas are emitted correctly, and parse errors propagate instead of being silently swallowed - - - - - -## v2.5.1 ---- - -### Output Guardrails for Streaming Responses - -Output guardrails (`output_guardrails` / `after_request_hooks`) now work with streaming responses. The AI Gateway accumulates all stream chunks, parses the complete response after the stream ends, and runs the configured output guardrails. Results are delivered as an additional SSE chunk at the end of the stream. - -- Works with all streaming endpoints: `/chat/completions`, `/completions`, and `/messages` -- Requires `x-portkey-strict-open-ai-compliance: false` header -- For `/chat/completions`: results sent as `data: {"hook_results":{"after_request_hooks":[...]}}` -- For `/messages` (Anthropic): results sent as `event: hook_results` SSE event - -[Guardrails Documentation](/aigw/product/guardrails#streaming-responses) | [Guardrails Capabilities](/aigw/product/guardrails/capabilities#streaming) - - - - - -## v2.5.0 ---- - -### Gateway-Local JWT Authentication - - -Requires Backend v1.13.0 or higher for Air Gapped deployments - - -The gateway can now validate JWT tokens locally without calling the management plane, reducing authentication latency for self-hosted deployments. - -SET `JWT_ENABLED=ON` to enable gateway-local JWT authentication. - -- Supports all existing JWT claims: `portkey_oid`, `portkey_workspace`, `scope`, `defaults`, `usage_limits`, `rate_limits` -- Organisation ID can be resolved from the token (`portkey_oid` / `organisation_id`) or from `ORGANISATIONS_TO_SYNC` when a single org is configured -- JWT-authenticated requests work with rate limit and usage limit policies - -[JWT Authentication Documentation](/aigw/product/enterprise-offering/org-management/jwt) - -### Headers-to-Metadata Injection - -New `HEADERS_TO_METADATA` environment variable allows automatically injecting request header values into metadata. This enables enriching observability data with upstream context (e.g., caller identity, trace IDs, or environment tags) without requiring clients to set `x-portkey-metadata`. - -Configure with a comma-separated list of header names: -``` -HEADERS_TO_METADATA=x-request-id,x-caller-service,x-environment -``` - -Header values are matched case-insensitively and injected into the request metadata (lower case keys) alongside any existing metadata from `x-portkey-metadata`. - -[Metadata Documentation](/aigw/product/observability/metadata) - -### Rate Limit Policies: Weekly Window and Endpoint Type Conditions - -- **Requests per week (`rpw`)**: Rate limit policies now support a weekly window in addition to per-minute, per-hour, and per-day -- **`endpoint_type` condition**: Rate limit policies can now target specific endpoint types (e.g., `chatComplete`, `embed`, `complete`) using the `endpoint_type` condition key - -[Usage & Rate Limit Policies Documentation](/aigw/product/enterprise-offering/budget-policies) - -### New Guardrail: Required Metadata Key-Value Pairs - -Added a new native guardrail that validates specific metadata key-value pairs before processing a request. Supports `all`, `any`, and `none` operators for flexible matching. - -| Check Name | Parameters | Supported Hooks | -|---|---|---| -| Required Metadata Key-Value Pairs | `metadataPairs` (object), `operator` (all/any/none) | `beforeRequestHook` | - -[Guardrail Checks Documentation](/aigw/product/guardrails/list-of-guardrail-checks) - -### Lasso Security Guardrail: v3 API Upgrade - -Upgraded the Lasso Security integration from v2 to v3 Classify API with the following improvements: -- **Findings-based detection**: Responses now include structured findings with name, category, action (BLOCK/AUTO_MASKING/WARN), and severity -- **Session and user tracking**: Supports `sessionId` and `userId` for conversation-level analysis -- **Multi-format support**: Handles `chatComplete`, `messages`, `complete`, and `embed` request types -- **Custom API endpoint**: Supports configurable `apiEndpoint` for self-hosted Lasso deployments - -[Lasso Security Documentation](/aigw/integrations/guardrails/lasso) - -### Image Format Support: HEIC and HEIF - -Added `image/heic` and `image/heif` to the list of supported image MIME types for multimodal requests and the Inline Image URLs guardrail. - -### Fixes and Improvements -- **Bedrock**: - - Improved tool calling support — Anthropic-format tool calls in Bedrock now correctly handle tool result content, cache control blocks, and strict role alternation - - Tool descriptions are now optional in Bedrock Converse API requests -- **Vertex AI**: - - Fixed model allowlist checking for proxy context-cache create routes - - Fixed schema conversion for tools -- **Google**: - - Fixed Google provider's embedding endpoint to correctly map the OpenAI-compatible `dimensions` parameter to Google's `output_dimensionality` parameter - - Fixed schema conversion for tools -- **Mistral AI**: - - Added `response_format` parameter support for Mistral AI, enabling structured JSON output and strict JSON schema enforcement. - - [Mistral AI Documentation](/aigw/integrations/llms/mistral-ai) -- Security dependency updates - - - - - -## v2.4.4 ---- - -### OpenTelemetry Semconv 1.40.0 Alignment - -The experimental GenAI OTel export now follows [semconv 1.40.0](https://opentelemetry.io/docs/specs/semconv/) for inference and embeddings spans. Key changes: - -- **Structured messages**: `gen_ai.input.messages`, `gen_ai.output.messages`, and `gen_ai.system_instructions` replace the previous indexed `gen_ai.prompt.{index}.*` format. Messages include multimodal `parts` with normalized types (text, tool_call, tool_call_response, image, audio, file, reasoning) -- **Normalized tool semantics**: `tool_use` (Anthropic), `tool_calls` (OpenAI), and `function_call` are all mapped to the standard `tool_call` type -- **Normalized finish reasons**: Provider-specific finish reasons (e.g., `tool_use`) are mapped to standard values (e.g., `tool_call`) -- **Provider normalization**: Provider names use semconv-compliant identifiers (e.g., `aws.bedrock`, `azure.ai.openai`, `gcp.gemini`). `gen_ai.system` is retained as a backward-compatible alias for `gen_ai.provider.name` -- **Embeddings support**: Spans for embedding requests include `gen_ai.request.encoding_formats` and `gen_ai.embeddings.dimension.count` -- **New attributes**: `gen_ai.operation.name`, `gen_ai.conversation.id`, `gen_ai.output.type`, `gen_ai.request.choice.count`, `gen_ai.request.top_k`, `gen_ai.usage.cache_creation.input_tokens`, `gen_ai.usage.cache_read.input_tokens` -- **Span updates**: Span kind changed from `SERVER` to `CLIENT`; span name now follows `{operation} {model}` format (e.g., `chat gpt-4o`) -- **New environment variable**: `EXPERIMENTAL_GEN_AI_OTEL_RESOURCE_ATTRIBUTES` allows setting custom OTLP resource attributes as comma-separated `key=value` pairs - -[OTel Export Documentation](/aigw/product/enterprise-offering/otel/complete-logs) - -### Anthropic Files API Support - -Added support for Anthropic's [Files API](https://docs.anthropic.com/en/docs/build-with-claude/files) (beta) through the gateway. You can now upload, list, retrieve, delete, and fetch content for files using the unified `/v1/files` endpoint. Uploaded files can be referenced in chat completions using `file_id` instead of re-uploading content with each request. - -[Anthropic Documentation](/aigw/integrations/llms/anthropic#files-api) - -### Prompt Caching TTL Support - -`cache_control` now supports an optional `ttl` field with values `"5m"` (default) or `"1h"` for both Anthropic and Bedrock providers. The 1-hour TTL is useful for agentic workflows and long conversations where follow-up requests may exceed the default 5-minute window. - -- [Anthropic Prompt Caching](/aigw/integrations/llms/anthropic/prompt-caching#cache-ttl-options) -- [Bedrock Prompt Caching](/aigw/integrations/llms/bedrock/prompt-caching#cache-ttl-options) - -### Fixes and Improvements -- **Anthropic**: Reverted `strict` parameter support for tool definitions (Anthropic tool input schema supports only a limited subset of JSON Schema) -- **Anthropic**: Fixed cache replay stream serialization for tool_use content blocks -- **Gemini/Bedrock**: Fixed cost logging when Gemini models respond in Anthropic Messages SSE stream format -- **Bedrock**: Fixed MCP logging hardcoded log path to use configurable paths -- **Azure AI Inference**: Added cache token pricing support -- **Vertex AI**: Added pricing support for embeddings on proxy routes -- **Budget Tracking**: Fixed race conditions in budget tracking pipeline to prevent data loss and double-counting -- **API Key Rotation**: Improved API key cache invalidation to use set-based lookups for more reliable key rotation -- **mTLS**: Fixed mTLS certificate handling with other internal services -- Updated model pricing configurations - - - - - -## v2.4.3 ---- - -### MCP Gateway Observability - -Enhanced MCP Gateway logging to capture all MCP events (not just tool calls). Logs now include caller identity, authorization scope and decision, and gateway `trace_id` for end-to-end traceability across MCP tool calls. - -### MCP Connection Pool Improvements - -Fixed MCP upstream connection pool invalidation when server configurations change. Connections are now properly refreshed based on configuration hashes, and OAuth session management has been hardened with masked token logging and improved user identity resolution. - -### Fixes and Improvements -- **Prometheus**: Added missing environment variable support for `PROMETHEUS_INCLUDE_MODEL_LABEL`, `PROMETHEUS_INCLUDE_METADATA_LABELS`, `PROMETHEUS_HISTOGRAM_BUCKETS`, `PROMETHEUS_HTTP_DURATION_BUCKETS`, and `PROMETHEUS_GC_BUCKETS` (these were documented in v2.2.5 but not wired into the env config) -- Updated model pricing configurations - - - - - -## v2.4.2 ---- - -### Bedrock Structured Outputs (Converse API) - -Bedrock structured outputs now use Bedrock's native `outputConfig` with `json_schema` text format via the Converse API, replacing the previous pass-through approach. This provides more reliable structured output enforcement for Claude models on Bedrock. - -[Bedrock Structured Outputs Documentation](/aigw/integrations/llms/bedrock/structured-outputs) - -### Fixes and Improvements -- **Anthropic**: Preserved `additionalProperties` field in tool `input_schema` conversion -- **Vertex AI**: Fixed metadata labels not being forwarded correctly for some request types -- **Azure OpenAI**: Fixed embedding pricing calculation -- **Azure Redis**: Fixed cluster mode connection handling for Azure Managed Redis -- **Akto Guardrail**: Updated schema configuration -- **Hooks**: Fixed sensitive header masking (Authorization, AI Gateway API key) in guardrail hook log responses -- Security dependency updates (zlib vulnerability patch, audit fixes) -- Updated model pricing configurations - - - - - -## v2.4.1 ---- - -### Fixes and Improvements -- **Zscaler Guardrail**: Fixed block reason handling to correctly parse and return block reasons from Zscaler AI Guard responses - - - - - -## v2.4.0 ---- - -### DeepInfra: Tool Calling & Expanded API Support - -DeepInfra now supports tool calling (function calling), including `tools`, `tool_choice`, and `parallel_tool_calls` parameters. Additionally, the **completions** and **embeddings** endpoints are now supported for DeepInfra. - -[DeepInfra Documentation](/aigw/integrations/llms/deepinfra) - -### Vertex AI: Metadata to Labels - -AI Gateway request metadata is now forwarded as Vertex AI resource labels. This enables governance and compliance for enterprises using GCP. - -### Streaming Usage for DeepSeek - -DeepSeek now respects `stream_options`. - -### Fixes and Improvements -- **Together AI**: Added cost logging support for video generation requests (`/v2/videos`) -- **Anthropic**: Added `strict` parameter support in tool definitions -- **OpenAI & Azure OpenAI**: Fixed `response_format` being sent to non DALL-E image generation models (e.g., `gpt-image-1`) that don't support this parameter -- **AWS SageMaker**: Fixed AssumeRole authentication support -- **AWS Bedrock**: Claude Code support for AWS Gov Models -- Security dependency updates - - - - - -## v2.3.0 ---- - -### New Guardrail: Akto Agentic Security - -Added [Akto](/aigw/integrations/guardrails/akto) as a new guardrails partner. Akto Agentic Security provides advanced threat detection and security scanning for LLM inputs and outputs, protecting against prompt injection, sensitive data leakage, and other security threats. Supports `beforeRequestHook` and `afterRequestHook` hooks with configurable timeout (default: 5000ms). - -### DeepSeek Tool Calling and Reasoning - -Added tool calling and reasoning support for the [DeepSeek](/aigw/integrations/llms/deepseek) provider: -- **Tool Calling**: `tools`, `tool_choice`, and `stream_options` parameters are now supported on `deepseek-chat` -- **Reasoning**: `reasoning_effort` parameter maps to DeepSeek's thinking mode for `deepseek-reasoner`. The `reasoning_content` field is returned in streaming responses - -### Azure AI Foundry Rerank - -Added [rerank](/aigw/integrations/llms/azure-foundry#rerank) support for Azure AI Foundry using Cohere models (e.g., `cohere.Cohere-rerank-v4.0-pro`). The `cohere.` prefix is automatically stripped before forwarding to the provider. - -### Vertex AI Enterprise Web Search - -Added support for [enterprise web search](/aigw/integrations/llms/vertex-ai#grounding-with-enterprise-web-search) grounding on Vertex AI. Pass the `enterpriseWebSearch` or `enterprise_web_search` tool in the `tools` array. Cost attribution distinguishes between standard Google Search and enterprise web search. - -### Vertex AI Cross-Cloud Workload Identity Federation - -Added support for **AWS-to-GCP Workload Identity Federation** for Vertex AI, enabling cross-cloud authentication from AWS-hosted deployments. New environment variables: -- `GCP_WIF_AUDIENCE`: The WIF audience for the GCP project -- `GCP_WIF_SERVICE_ACCOUNT_EMAIL`: The GCP service account email to impersonate - -### Bedrock Batch Embeddings - -Added batch embeddings support for Bedrock. You can now run batch inference jobs for embeddings in addition to chat completions. - -### Bedrock Guardrails Custom Host - -Added `customHost` support for the [Bedrock guardrail](/aigw/integrations/guardrails/bedrock-guardrails) plugin, enabling connections to custom or private Bedrock endpoints. - -### Fixes and Improvements -- **Bedrock**: Fixed region handling for Anthropic models when using the `/messages` route -- **Bedrock**: Added `au` (Australia) prefix support for model config resolution -- **Bedrock/Vertex AI**: Improved filtering logic for unsupported `anthropic-beta` header values -- **Bedrock**: Updated validation checks for batch `completion_window` -- Updated model pricing configurations - - - - - -## v2.2.6 ---- - -### Fixes and Improvements -- **Debug Logging**: Organisation-level `enforce_debug_log_setting` now correctly ignores the `x-portkey-debug` request header when enforcement is enabled. Previously, the header could override the org-level setting even when enforcement was turned on -- **Azure AI Foundry**: Fixed token counting for Anthropic models (e.g., Claude Opus) accessed through Azure AI Foundry via the Messages API. The token counter now correctly handles Anthropic-style usage fields (`input_tokens`/`output_tokens`) in addition to OpenAI-style fields (`prompt_tokens`/`completion_tokens`) - - - - - -## v2.2.5 ---- - -### Secondary Redis for BullMQ Workers - -Added support for a **dedicated Redis instance** for BullMQ worker queues, separate from the primary Redis used for caching and rate limiting. This avoids Redis Cluster slot limitations that can affect BullMQ job processing. - -New environment variables: - -| Variable | Description | -|----------|-------------| -| `REDIS_WORKERS_URL` | Full Redis URL for the workers instance | -| `REDIS_WORKERS_HOST` | Redis host (alternative to URL) | -| `REDIS_WORKERS_PORT` | Redis port (alternative to URL) | -| `REDIS_WORKERS_USERNAME` | Redis username | -| `REDIS_WORKERS_PASSWORD` | Redis password | -| `REDIS_WORKERS_TLS_ENABLED` | Enable TLS for the workers Redis connection | -| `REDIS_WORKERS_AUTH_MODE` | Authentication mode (e.g., `EntraID`, `ManagedIdentity`) | - -When not configured, workers continue to use the primary Redis connection. - -### Configurable Prometheus Metrics - -Prometheus histogram buckets and metric labels are now configurable to control cardinality: - -- **`PROMETHEUS_INCLUDE_MODEL_LABEL`**: Opt-in to include the `model` label in metrics (default: `false`). Enabling this can cause high cardinality -- **`PROMETHEUS_INCLUDE_METADATA_LABELS`**: Opt-in to include custom metadata labels (default: `false`) -- **`PROMETHEUS_LABELS_METADATA_ALLOWED_KEYS`**: Comma-separated list of allowed metadata keys when metadata labels are enabled -- Custom histogram bucket environment variables for LLM latency, HTTP duration, and GC duration metrics -- Default bucket counts reduced for lower cardinality while preserving observability - -### Provider Updates -- **Perplexity**: Added as a native [Responses API](/aigw/api-reference/responses/create-response) provider -- **OpenRouter**: Added embeddings endpoint support -- **ZhipuAI (Z.ai)**: Added cost attribution for chat and image generation models - -### Fixes and Improvements -- **Groq**: Fixed streaming tool calls by adopting a minimal pass-through response transform -- **Together AI**: Fixed streaming delta content returning `undefined` instead of `null` for missing content, which caused issues with tool call responses -- **OpenAI**: Fixed token config for embedding models to correctly handle units -- Added retries for ClickHouse queue inserts to improve analytics data reliability -- Updated model pricing configurations - - - - - -## v2.2.4 ---- - -### Bedrock Anthropic Citations - -Added support for Anthropic's citations feature on Bedrock for chat completions API. - -### Zscaler AI Guard - -Added [Zscaler AI Guard](/aigw/integrations/guardrails/zscaler) as a new guardrails partner. Zscaler AI Guard enforces Detections Policies to perform security checks including Data Loss Prevention (DLP) and prompt injection protection on both inbound prompts and outbound model responses. - -### Speech SSE Streaming - -OpenAI and Azure OpenAI [text-to-speech](/aigw/product/ai-gateway/multimodal-capabilities/text-to-speech#sse-streaming) requests now support Server-Sent Events (SSE) streaming. Set `stream_format: "sse"` to receive audio data as a stream of events with proper usage logging. - -### ZhipuAI Image Generation - -ZhipuAI (Z.ai) now supports [image generation](/aigw/integrations/llms/zhipu#image-generation) through the CogView model family (e.g., `cogview-4-250304`). Supported parameters include `prompt`, `model`, `n`, `size`, `response_format`, and `user`. - -### Fixes and Improvements -- **OpenAI & Azure OpenAI**: Added support for the `prompt_cache_retention` parameter in chat completions and the Responses API -- **Vertex AI**: Added `x-portkey-vertex-auth-type` support in provider options for configuring authentication type (e.g., workload identity) -- **Gemini & Vertex AI**: Fixed tool message content handling when content is an array of text parts (OpenAI format) -- **Logging**: `user_id` is now automatically recorded in analytics from the API key's associated user. -- Fixed stream consumption errors when the downstream client disconnects mid-stream -- Updated model pricing configurations - - - - - -## v2.2.3 ---- - -### New Provider: Databricks - -Added support for Databricks Model Serving as a new provider with support for chat completions, completions, and embeddings. - -[Databricks Documentation](/aigw/integrations/llms/databricks) - -### Together AI Reasoning Support - -Together AI now supports reasoning/thinking models with `reasoning_effort` parameter and `content_blocks` in the response. When `strictOpenAiCompliance` is set to `false`, the response includes structured `content_blocks` with both thinking content and the final text response. Both streaming and non-streaming modes are supported. - -[Together AI Documentation](/aigw/integrations/llms/together-ai#reasoning--thinking-support) - -### Fixes and Improvements -- **Unified Messages & Responses API**: Fixed issues when a request config had a combination of native and non-native providers (e.g., load balancing across Anthropic and OpenAI). The adapter decision is now made per-provider, ensuring correct behaviour for mixed configs. -- **Responses API**: Groq, OpenRouter, and xAI now use the native `/responses` endpoint supported by each provider for Responses API requests -- **Anthropic**: Fixed `max_tokens` handling in chat completions — `max_tokens` now takes precedence over `max_completion_tokens` when both are provided -- Updated model pricing configurations for Bedrock - - - - - -## v2.2.2 ---- - -### Claude 4.6 Support - -Full support for Claude 4.6 features across Anthropic, Bedrock, and Azure AI Foundry: - -- **Adaptive Thinking**: Support for `thinking: { type: "adaptive" }` with `reasoning_effort` parameter -- **`output_config`**: Passthrough for structured outputs (`output_config.format`) and reasoning control (`output_config.effort`) -- **`response_format` → `output_config` Mapping**: `response_format` with `json_schema` is automatically mapped to Anthropic's `output_config.format` -- **`text_editor_20250728`**: Support for the latest text editor tool version -- **New Stop Reasons**: `refusal` and `model_context_window_exceeded` mapped to OpenAI-compatible `finish_reason` values - -### Responses API Improvements - -- **Native Provider Support**: Added x-ai (Grok), Groq, OpenRouter, and Azure AI Foundry as native Responses API providers -- **Schema Compliance**: Response objects now include all required fields per the OpenAI Responses API spec -- **Prompt Caching**: `cache_control` preserved on content items and tools in the Responses API adapter -- **Thinking Passthrough**: `thinking` parameter now works through the Responses API adapter - -### Fixes and Improvements -- **Gemini/Vertex AI**: Fixed structured output handling by switching to `responseJsonSchema` for proper JSON schema support -- **Gemini/Vertex AI**: Improved thought signature handling for tool calling in multi-turn conversations -- **Vertex AI**: Fixed global region handling and improved error responses for batch operations -- **Messages API**: Fixed override params handling when checking native provider support in fallback/retry configurations -- **HTTP**: Fixed timeout configuration to use separate HTTP agents, preventing interference with data service connections -- Updated model pricing configurations -- Updated depencies to patch security vulnerabilities - - - - - -## v2.2.1 ---- - -### Fixes and Improvements -- **Workspace Alerts**: Fixed threshold validation for integration workspace alerts -- **Gemini/Vertex AI**: Fixed JSON schema handling issues for structured outputs -- **OpenAI**: Fixed cost attribution for image editing requests -- Updated model pricing configurations across multiple providers - - - - - -## v2.2.0 ---- - -### Messages API Now Works with All Providers - -The Anthropic Messages API (`/v1/messages`) now works with **all providers** through a universal adapter, not just Anthropic, Bedrock, and Vertex AI. Use the Messages API format with OpenAI, Google, and more providers seamlessly. - -### OpenTelemetry Enhancements -- Metadata arrays are now flattened into individual span attributes for better observability and easier querying in your tracing backend - -### Provider Updates -- **Gemini/Vertex AI**: Added support for the `media_resolution` parameter to control the resolution of media inputs when processing images and videos - -### Fixes and Improvements -- Fixed model detection for batch requests to correctly identify and attribute the model being used -- Updated model pricing configurations across multiple providers - - - - - -## v2.1.0 ---- - -### Responses API Now Works with All Providers - -The OpenAI Responses API (`/v1/responses`) now works with **all 70+ providers**, not just OpenAI and Azure OpenAI. Use the Responses API format with Anthropic, Google, Bedrock, and more. - -**Note:** Responses API-only features like `previous_response_id` state management and built-in tools (`web_search`, `file_search`, `computer_use`) are only supported on OpenAI and Azure OpenAI. - -[Responses API Documentation](/aigw/api-reference/responses/create-response) - -### Provider Updates -- **Vertex AI**: Added option to skip cost attribution for Provisioned Throughput (PTU) deployments. Configure via: - - Model Catalog Integration settings: `vertex_skip_ptu_cost_attribution: true` - - Virtual Key API: `vertexSkipPtuCostAttribution: true` -- **Bedrock**: Fixed handling of `cache_control` blocks for Anthropic models to remove unsupported `scope` parameter -- **Anthropic**: Improved handling of beta headers and version headers across all routes - -### Batch Pricing -- Added support for dedicated batch pricing configurations. When batch-specific pricing is not available, the system defaults to 50% of standard pricing for cost attribution on batch requests. - Works with `Data Service` version v1.5.0 or higher -Older versions of Data Service continue to work for batch pricing with the default 50% cost attribution. - -### Fixes and Improvements -- **MCP Gateway**: Fixed user identity forwarding not working correctly -- **JWT Plugin**: Fixed validation failures when identity providers like Microsoft Entra ID don't include the `alg` field in their JWKS keys. The plugin now correctly falls back to `RS256` when the configured algorithms list is empty. -- **OpenAI**: Fixed cost attribution not being reflected for embeddings requests -- **OpenAI Streaming**: Fixed `service_tier` and `system_fingerprint` handling in stream concatenation and cache streaming -- **Gemini/Vertex AI**: Improved error handling for `MALFORMED_FUNCTION_CALL` responses. Error details are now returned in `finish_message` field when `strictOpenAiCompliance` is disabled. -- Various error logging improvements and internal dependency updates - - - - - -## v2.0.0 ---- - -### MCP Gateway is now Generally Available -The AI Gateway's MCP Gateway is now GA, providing enterprise-grade infrastructure for the Model Context Protocol. Key features include: -- **Centralized MCP Server Management**: Add and manage internal and external MCP servers from a single registry -- **OAuth 2.1 & API Key Authentication**: Secure access with OAuth for interactive use or API keys for programmatic access -- **Workspace-Level Provisioning**: Control which teams and workspaces can access specific MCP servers and tools -- **Full Observability**: Monitor and debug all MCP tool calls with complete context, logs, and traces - -[MCP Gateway Documentation](/aigw/product/mcp-gateway/quickstart) - -### Unified Rerank API -Introduced a unified `/rerank` endpoint that provides a consistent interface for reranking across multiple providers. This allows you to switch between reranking providers (Cohere, Jina, Voyage AI) without changing your application code. - -### Usage and Rate Limit Policy Enhancements -Added new condition keys for fine-grained budget and rate limit policies: -- `virtual_key`: Match by virtual key slug -- `provider`: Match by provider (e.g., `openai`, `anthropic`) -- `config`: Match by gateway config slug -- `prompt`: Match by prompt template slug -- `model`: Match by model with wildcard support (e.g., `@openai/gpt-4o`, `@anthropic/*`) - -[Budget Policies Documentation](/aigw/product/enterprise-offering/budget-policies) - -### New Plugins -- **Inline Image URLs Plugin**: Convert external image URLs to inline base64 data for VPC-SC environments where external URLs are not accessible -- Updated support for multiple partner guardrails - -### Dynamic Model Pricing for Air Gapped Deployments -- Dynamically fetch model pricing without requiring image updates -- Set `MODEL_CONFIGS_PROXY_FETCH_ENABLED=ON` in **Gateway** to fetch pricing from the Backend service -- The **Backend** service resolves pricing via configurable sources (proxy service, log store, or local files) — see [Backend Changelog (v1.7.0)](/changelog/backend#private-deployment-pricing) for Backend-side environment variables -- See the [Air-Gapped Model Pricing guide](/self-hosting/airgapped/model-pricing) for full setup instructions - -### Infrastructure Updates -- **GCP Workload Identity**: Added support for Workload Identity IAM-based access for GCS and VertexAI, enabling keyless authentication in GKE environments - -### Provider Updates -- **Gemini/Vertex AI**: Added explicit caching support for Gemini models -- **Gemini/Vertex AI**: Improved thought signature handling for seamless multi-turn conversations with Gemini models -- **Gemini/Vertex AI**: Added support for minimal thinking mode, mapping `reasoning_effort: minimal` to `budget_tokens: 1024` for Gemini 2.5 models -- **Gemini/Vertex AI**: Fixed response schema mapping for structured outputs -- **Bedrock**: Fixed handling when documents are present by adding required text block -- **Bedrock**: Fixed handling of empty tool arguments in multi-turn conversations - -### Performance Improvements -- **Budget Tracking**: Implemented in-memory cache key tracking with periodic sync to Redis. Budget increments are now accumulated in memory and synced every 10 seconds, significantly reducing Redis calls on the hot path - -### Fixes and Improvements -- **Models Endpoint**: The `/v1/models` endpoint now accepts API keys with `completions:write` permission (for upstream provider routing) in addition to `virtual_keys:list` (for the AI Gateway models endpoint) -- **Logging**: Response body logging is now skipped for all embeddings requests (previously only skipped for specific models), reducing log storage -- **OpenAI**: Fixed JSON body parsing for `/v1/vector_stores/{id}/files` endpoints which was incorrectly returning "Missing required parameter" errors -- Updated token counting logic for chat completions to align with OpenAI specification -- Fixed cost attribution: costs are no longer attributed for Anthropic's `count_tokens` endpoint -- Fixed cost attribution: costs are no longer attributed when usage object is not present in Responses API -- Fixed handling of self-referencing JSON schemas in structured outputs -- Fixed `prompt_tokens` returning `0` in streaming responses for Vertex AI Anthropic models -- Internal dependency updates - - - - - -## v1.17.15 ---- - -### New Guardrails -- **CrowdStrike AIDR**: Added partner plugin integration with CrowdStrike AI Detection and Response for scanning LLM inputs and outputs. Supports blocking or redacting content based on configured rules. - -### Fixes and Improvements -- Fixed sequential guardrail checks execution - - - - - -## v1.17.14 ---- - -### Fixes and Improvements -- Improved management plane sync handling for multi-organisation deployments -- Internal dependency updates - - - - - -## v1.17.13 ---- - -### Provider Updates -- **Gemini/Vertex AI**: Fixed `reasoning_effort` parameter mapping for Gemini 2.5 models. Now correctly maps to `thinking_budget` (token-based) instead of `thinkingLevel`: - - `low`: 1,024 tokens - - `medium`: 8,192 tokens - - `high`: 24,576 tokens - - Gemini 3.0+ models continue to use `thinkingLevel` mapping - -### Fixes and Improvements -- **Bedrock**: Fixed `anthropic_beta` parameter handling to properly parse comma-separated string values into arrays (e.g., `"beta1, beta2"` now correctly converts to `["beta1", "beta2"]`) -- **Bedrock**: Updated `anthropic_version` parameter handling for Anthropic models - - - - - -## v1.17.12 ---- - -### Responses API Hooks Support -- Hooks and guardrails now fully support the `/v1/responses` endpoint -- Apply input/output guardrails, custom webhooks, and other hooks to Responses API requests - -### New Azure Guardrails -- **Shield Prompt**: Detects jailbreak and prompt injection attacks using Azure AI Content Safety Prompt Shields API - - Analyzes system prompts and user messages for potential attacks - - Supports both API key and Entra ID authentication - - [Documentation](/aigw/integrations/guardrails/azure-guardrails#azure-shield-prompt) -- **Protected Material**: Detects copyrighted or protected text content in LLM outputs using Azure AI Content Safety API - - Identifies known protected/copyrighted material in model responses - - Helps ensure compliance with intellectual property requirements - - [Documentation](/aigw/integrations/guardrails/azure-guardrails#azure-protected-material) - -### Sequential Guardrails Execution -- Added `sequential` flag for guardrails to execute checks in order rather than in parallel -- Useful when guardrail results depend on previous checks or when order matters for compliance - -### Provider Updates -- **xAI**: Added support for **Realtime Voice Agent API** (Grok Voice API), enabling real-time voice interactions through WebSocket connections. The API is OpenAI Realtime API compatible, making it easy to integrate with existing voice agent workflows. Includes pricing configuration for the `grok-2-voice` realtime model with full cost attribution. -- **Gemini/Vertex AI**: Added support for **Google Maps grounding** with Gemini and Vertex AI models -- **OpenAI & Azure OpenAI**: Added support for new `gpt-image-1` parameters: `moderation`, `output_format`, `output_compression`, `background`, `partial_images`, `stream` -- **Azure OpenAI**: Added `x-portkey-azure-entra-scope` header to specify custom authentication scopes for Entra ID and Managed Identity auth (e.g., `https://ai.azure.com/.default` for Azure AI Foundry) -- **OpenAI & Azure OpenAI**: Added `output_expires_after` parameter support for batch creation. Azure OpenAI additionally supports blob-based batch inputs via `input_blob` and `output_folder` parameters. -- **Anthropic**: Added full support for Anthropic's citations feature in responses -- **Anthropic on Bedrock**: Anthropic models on Bedrock now use native Messages API format when calling `/v1/messages` route, enabling features like citations that aren't available through the Converse API - -### Pricing Updates -- **Gemini Thinking Tokens**: Updated pricing to reflect that Gemini thinking tokens are no longer charged separately -- **Vertex AI Embeddings**: Added cost attribution for embedding models on proxy routes - -### Fixes and Improvements -- Fixed Oracle provider configuration mapping -- Improved Anthropic beta header handling across Vertex AI and Bedrock -- Fixed minor issues with `Regex Replace` guardrail - - - - - -## v1.17.11 ---- - -### Prometheus Metrics Toggle -- Added ability to **disable Prometheus metrics** via environment variable -- Set `ENABLE_PROMETHEUS=false` to disable the `/metrics` endpoint and metrics collection middleware -- Useful for deployments where Prometheus metrics are not needed or when using alternative monitoring solutions -- Enabled by default for backward compatibility - - - - - -## v1.17.10 ---- - -### OpenAI-Compatible Response IDs -- **Google and Vertex AI**: Response IDs are now generated in OpenAI-compatible format for improved compatibility with OpenAI SDKs and tooling - - Chat completion IDs: `chatcmpl-{random}` (previously `portkey-{uuid}`) - - Tool call IDs: `call_{random}` (previously `portkey-{uuid}`) -- This change improves interoperability when using OpenAI-compatible clients with Google/Vertex AI models - -### Pricing Updates -- Added pricing configurations for new models across multiple providers - - - - - -## v1.17.9 ---- - -### Provider Updates -- **Azure OpenAI**: Added support for the `/v1/images/edits` endpoint for image editing -- **Gemini/Vertex AI**: Added support for the `image_config` parameter to control image generation settings. Users can now specify `aspect_ratio` and `image_size` for Gemini's image generation capabilities. -- **Gemini/Vertex AI**: Fixed web search cost attribution for grounding calls to correctly detect grounding chunks in responses when strict openai compliance flag is not set or set to true - - - - - -## v1.17.8 ---- - -### Pricing Updates -- **OpenAI Web Search**: Updated web search cost calculation to align with OpenAI's latest pricing model, consolidating context-based pricing (`web_search_low_context`, `web_search_medium_context`, `web_search_high_context`) into a single `web_search` metric - -### Fixes and Improvements -- **F5 Guardrails**: Fixed API endpoint and improved blocking logic. Requests are now blocked when the scan outcome is `blocked` or `flagged`, and redaction is only applied when explicitly enabled. -- **Streaming**: Fixed stream response log handling to correctly include annotations in responses - - - - - -## v1.17.7 ---- - -### New Guardrail: Blocked Tools -- Added a new guardrail plugin to control which AI tools can be used in requests. Block specific tool types or function names using blocklists or allowlists. -- Supports blocking tool types: `function`, `web_search_preview`, `web_search`, `file_search`, `code_interpreter`, `computer_use`, `mcp` -- Supports blocking specific function names by name -- Can use either blocklist (block specific tools) or allowlist (only allow specific tools) approach - -### Redis Cluster Discovery -- Added support for dynamic Redis cluster endpoint discovery via an HTTPS URL -- Added support for static Redis cluster endpoints configuration -- New environment variables: - - `REDIS_CLUSTER_ENDPOINTS`: Comma-separated list of static Redis cluster endpoints (e.g., `10.0.1.1:6379,10.0.1.2:6379`) - - `REDIS_CLUSTER_DISCOVERY_URL`: HTTPS URL that returns comma-separated Redis cluster endpoints - - `REDIS_CLUSTER_DISCOVERY_AUTH`: Optional authorization header for the discovery endpoint - - `REDIS_CLUSTER_DISCOVERY_REFRESH_INTERVAL`: Cache refresh interval in milliseconds (default: 5 minutes) -- Supports IPv4, IPv6, and hostname formats for endpoints - -### Semantic Cache Improvements -- Added configurable embedding dimensions for semantic cache -- New environment variable: `SEMANTIC_CACHE_EMBEDDING_DIMENSIONS` to set custom embedding dimensions (default: 1536) - -### Provider Updates -- **Anthropic**: Fixed `anthropic-beta` header to be allowed on all routes for the Anthropic provider -- **Gemini/Vertex AI**: Fixed response handling to correctly concatenate multiple text parts in responses - -### Fixes and Improvements -- Excluded batch and files GET requests from budget exhausted checks -- Updated F5 guardrails to use `prompts` endpoint instead of `scan` -- Moved the rate limiter implementation to fixed window rate limiter to avoid continuous token refills. - - - - - -## v1.17.6 ---- - -### New Provider -- **Oracle Cloud Infrastructure (OCI)**: Added support for Oracle OCI Generative AI service with request signing authentication. Supports Cohere and Meta Llama models via Oracle's inference API. - - [Documentation](/aigw/integrations/llms/oracle) - -### New Features -- **Sticky Load Balancing**: Added sticky session support for load balancing configurations. Ensures consistent routing based on configurable hash fields (e.g., user ID, session ID) with configurable TTL. - - [Documentation](/aigw/product/ai-gateway/load-balancing#sticky-load-balancing) - -### Provider Updates -- **Gemini/Vertex AI**: Added `reasoning_effort` parameter support for controlling thinking behaviour. Maps OpenAI's `reasoning_effort` (`minimal`/`low`/`medium`/`high`) to Gemini's `thinkingLevel` (`low`/`high`). - - [Documentation](/aigw/integrations/llms/gemini#using-reasoning_effort-parameter) -- **Azure OpenAI**: Added support for v1 preview API version for Azure OpenAI endpoints -- **Azure OpenAI**: Added pricing support for batch `/responses` endpoint with deployment - -### Fixes and Improvements -- Fixed Vertex AI model preference handling for batch pricing requests -- Fixed Realtime API to update request body with model from query parameters -- Fixed OpenTelemetry status code conversion to HTTP status codes -- Fixed embedding providers to accept array format for input -- Removed unnecessary text check condition for afterRequestHook in HooksManager -- Updated Dockerfile npm version to 11.6.4 to fix glob vulnerability - - - - - -## v1.17.5 ---- - -### Security -- **SSRF Protection**: Added request validation to prevent Server-Side Request Forgery (SSRF) attacks via configuration URLs - -### OpenTelemetry Improvements -- **W3C Traceparent Fix**: Fixed `traceparent` header parsing to correctly set `parent_span_id` from the incoming span and generate a new `span_id` for the gateway's span. This improves trace linking in distributed tracing tools. -- **Span Kind**: Added `span.kind = SERVER` attribute to exported spans -- **Span Name**: Automatically sets `span_name` from HTTP method and path (e.g., `POST /v1/chat/completions`) when `traceparent` is provided -- [Documentation](/aigw/product/observability/traces#w3c-trace-context-support) - -### Semantic Cache Improvements -- Added configurable similarity metric support for Milvus and Pinecone vector stores (COSINE, L2, IP) -- Fixed semantic cache logic for Milvus to use correct similarity threshold comparisons - -### Realtime API Improvements -- Added support for provider authentication via query parameters (e.g., `model=@provider/model`) -- Fixed model parameter sanitization in Realtime API handler - -### Fixes and Improvements -- Fixed OpenAI responses stream parsing logic for `createModelResponse` function -- Updated Jest version to address security vulnerabilities - - - - - -## v1.17.4 ---- - -### Provider Updates -- **Azure AI Foundry**: Added support for Anthropic Claude models via Azure AI Foundry, including: - - Native `/messages` endpoint support for Anthropic-native features (extended thinking, prompt caching, native streaming) - - Support for Claude 4.5 models: `claude-opus-4-5-20251101`, `claude-sonnet-4-5-20250929`, `claude-haiku-4-5-20251001` - - [Documentation](/aigw/integrations/llms/azure-foundry#using-anthropic-models-on-azure-ai-foundry) - -### Fixes and Improvements -- Fixed cross-region support for Bedrock and Sagemaker - user-specified region now takes priority over credentials region -- Fixed Azure OpenAI to use deployment name for responses API -- Fixed TLS configuration in agent store to use 'connect' instead of 'tls' -- Improved Redis URL handling with fallback to `redis:6379` for empty URLs - - - - - -## v1.17.3 ---- - -### Fixes and Improvements -- Fixed guardrails execution for unified `/messages` endpoint - - - - - -## v1.17.2 ---- - -### Provider Updates -- **VertexAI and Google**: - - Added cost calculation support for `gemini-2.5-flash-image` and `gemini-3-pro-image-preview` models - - Added support for the `thought_signature` parameter. [Documentation](/aigw/integrations/llms/vertex-ai#thought-signatures-tool-calling-verification) -- **Bedrock**: Added support for the `x-portkey-aws-region` header to set the region for Bedrock and SageMaker integration with `serviceRole` auth type - -### Fixes and Improvements -- Added the traces URL from the incoming request to the analytics objects created from OpenTelemetry spans - - - - - -## v1.17.1 ---- -### Hook Results in Streaming Responses -- Added support for returning hook results in streaming responses for `/chat/completions`, `/completions`, `/embeddings`, and `/messages` endpoints. This enables you to stop streaming or display redacted content when a request violates guardrail policies, allowing clients to enforce content policy guidelines in real-time. -- [Documentation](/aigw/product/guardrails#streaming-responses) - -### New Plugins -- **F5 Guardrails**: Added partner plugin integration with content moderation and PII detection/redaction capabilities. - -### Provider Updates -- **Azure OpenAI**: Added support for `model-router` -- **TogetherAI**: Added support for image generation cost calculation -- **VertexAI**: - - Added support for video generation (Veo models) cost calculation - - Fixed MIME type mapping for opus format: Changed from `audio/ogg` to `audio/opus` for better accuracy - - Added MIME type mapping for aliases: 'ogg', 'pcm', 'aac', and 'm4a' alongside their existing 'x-' prefixed variants -- **OpenAI and Azure OpenAI**: Added support for video generation (Sora models) cost calculation -- **Bedrock**: Fixed `count_tokens` endpoint errors for Anthropic models -- **Anthropic**: Added proper header forwarding for `anthropic-beta` and `anthropic-version` headers from the original request - -### Fixes and Improvements -- Added and updated model pricing configurations across multiple providers -- Fixed issues in `/log/exports` endpoints -- Fixed an issue where the provider identifier was not being logged in the analytics store for some cases -- Updated internal dependencies to patch security vulnerabilities -- Reduced model pricing configuration cache TTL for faster pricing updates - - - - - -## v1.17.0 ---- - -**Requires a Helm repo update (>app-1.4.0)** - - -**For air-gapped deployments, `Backend` version v1.5.0 is required as it adds new columns in the analytics store** - - -### Security Patch -- Removed root user from container image (BREAKING CHANGE). The container image used by this chart no longer runs as root. The image now runs processes with a non-root UID and enforces a non-root container securityContext. Requires Helm repo upgrade (>app-1.4.0) to deploy the new image and chart settings. - -### Usage and Rate Limit Policy -- Introduced usage limits and rate limit policies, which allow organisations to apply flexible budget and rate limit controls based on dynamic conditions (API keys, metadata, workspace, etc.). -- More details: [Documentation](/aigw/product/enterprise-offering/budget-policies) and [API Reference](/aigw/api-reference/usage-limit-policies/list-usage-limits-policies) - -### Logging Enhancements -- Added support for OpenTelemetry W3C trace context headers (`traceparent` and `baggage`) to enable integration with distributed tracing systems. -- When these standard headers are present, they are automatically parsed and mapped to the AI Gateway's internal tracing headers. - -### Finetuning Cost Tracking -- Added cost calculation logic for fine-tuning operations across multiple AI providers (OpenAI, Azure OpenAI, and Vertex AI). - -### Hourly-Based Log Object Prefix -- Introduced support for hourly-based S3 file path prefixes. -- Set `LOG_STORE_FILE_PATH_FORMAT` environment variable to toggle between `v1` (flat structure - existing format) and `v2` (time-hierarchical) path formats. By default, it will use the existing format (v1). -- `v1` (default) - The object path will follow this structure: `30//.json` -- `v2` - The object path will follow this structure: `30///////.json` -Changing the `LOG_STORE_FILE_PATH_FORMAT` environment variable only affects newly written logs. Previously written logs will retain their original path format and are not migrated to the new structure. -The new v2 prefix structure is currently not supported for air-gapped deployments where `LOG_STORE` is set to `control_plane`. Support will be added in a future release. - -### Provider Updates -- **OpenAI and Azure OpenAI**: Fixed cached tokens cost calculation for `responses` endpoint -- **Bedrock**: Added handling for empty `tools` array in `messages` endpoint -- **Google and VertexAI**: Allow `0` and empty string values for parameters - -### Fixes and Improvements -- Added performance improvements for file uploads -- Added logging for gateway-managed POST `/batches` and `/fine_tuning/jobs` requests -- Updated `gcs` and `s3_custom` log store implementation to use unsigned-payload for PUT operations - - - - -## v1.16.7 ---- - -### Provider Updates -- **Google VertexAI**: Added mime type mapping for multiple audio formats - `opus`, `flac`, `pcm16`, `x-aac`, `x-m4a`, `mpeg`, `mpga`, `mp4`, `webm` - -### Fixes and Improvements -- Improved cache handling for Azure Managed Identity and Entra tokens - - - - -## v1.16.6 ---- - -### Token Sum Prometheus Metric -- Added a new Prometheus metric `llm_token_sum` to track the total LLM tokens consumed - -### New Providers -- **Modal**: Added support for Modal Labs as a new LLM provider - -### JWT Plugin Enhancements -- Added token introspection endpoint validation as an alternative to JWKS -- Implemented flexible claim validation with multiple match types (exact, contains, containsAll, regex) -- Added claim extraction and injection into request context as headers - -### Provider Updates -- **Cohere**: Updated the Cohere integration to use the v2 `chat` and `embed` endpoints - -### Fixes and Improvements -- Added an environment variable (`SKIP_DATAPLANE_CONFIG_CHECK = 'true'`) to allow configs with multiple targets for unified batches file upload endpoints -- Implemented cost and token calculation for trace spans following the GenAI OTel semantic conventions -- Added and updated model pricing configs for multiple providers - - - - -## v1.16.5 ---- - -### Provider Updates -- **VertexAI**: Added support for Computer Use. See [Documentation](/aigw/integrations/llms/vertex-ai#computer-use-browser-automation-preview). -- **VertexAI**: Added support for `anthropic-beta`. Please pass `anthropic-beta` or `x-portkey-anthropic-beta` header to enable this feature. -- **OpenAI**: Added support for `conversation` and `modalities` parameters. -- **Anthropic**: Add role to `assistant` in chat completion stream response. - -### HTTP Proxy Support for Websockets -- **HTTPS_PROXY**: Added support for HTTP Proxy support for websockets. - -### Experimental GenAI Otel Support -- The semantic convention is defined here: https://github.com/open-telemetry/semantic-conventions/blob/main/model/gen-ai/spans.yaml -- [Documentation](/aigw/product/observability/opentelemetry#experimental-features) - -### Robust Cors Support -- Added support for robust CORS configuration for the Gateway. -- The following environment variables will be used to configure the CORS configuration: - - `CORS_ALLOWED_ORIGINS`: The allowed origins for CORS requests - - `CORS_ALLOWED_METHODS`: The allowed methods for CORS requests - - `CORS_ALLOWED_HEADERS`: The allowed headers for CORS requests - - `CORS_ALLOWED_EXPOSE_HEADERS`: The exposed headers for CORS requests - - - - -## v1.16.4 ---- - -### Provider Updates -- **AWS Bedrock**: Added support for `bearer tokens` for authentication. - -## Infrastructure Updates -- **EKS Pod Identity**: Added support for EKS Pod Identity in conjuction with IRSA for EKS clusters. - - - - - -## v1.16.3 ---- - -### Provider Updates -- **OpenAI and Azure OpenAI**: Add support for streaming audio transcription requests. -- **VertexAI**: Enable custom model support for batch inference. - - - - - -## v1.16.2 ---- - -### Provider Updates -- **Google and VertexAI**: Fixed cost calculation for grounding requests to check `groundingChunks` array in the response. Cost attribution now occurs when at least one grounding chunk is present - - - - - -## v1.16.1 ---- - -### New Guardrails -- **Add Prefix**: Add a configurable prefix to the user's input before sending to the model -- **Allowed Request Types**: Control which request types (endpoints) can be processed. Use either an allowlist or blocklist approach - -### Fixes and Improvements -- Handle 0 value for the Bedrock `temperature` parameter - - - - - -## v1.16.0 ---- - -### Enhancements For Unified Batches -- Track provider batches automatically through data-service for cost tracking - Requires `Data Service` version 1.3.0 to be deployed. - -### AWS Service Role Auth -- Added support for AWS Service Role authentication for AWS Bedrock models. -- In this mode, the Gateway will use its own service role to invoke Bedrock models. - -### Configurable Outbound Request Timeout -- Introduced `REQUEST_TIMEOUT` environment variable to configure fetch timeouts for outbound LLM requests -- The value should be in milliseconds. Default is 300000 (5 minutes) -- NOTE: This timeout is only applicable for individual LLM requests that are made by the Gateway - -### NO_PROXY support -- Added support for `NO_PROXY` environment variable to bypass outbound requests to the specified hosts. -- This is useful when you want to bypass outbound requests to certain hosts in conjuction with `HTTPS_PROXY`. - -### TLS support for Management Plane Calls for Air-gapped Deployments -- Added support for TLS configuration for management plane calls for air-gapped deployments. -- Use `TLS_KEY`, `TLS_CERT` , `TLS_CA` environment variables to configure the TLS certificate, key and CA certificate respectively. - -### Provider Updates -- **AWS Bedrock**: - - Added `global` profile support for Bedrock models - - Added `name` field mapping for `/messages` endpoint document object - - Streamlined token counting endpoint to use invoke instead of converse mode for better compatibility -- **Azure**: - - Improved Azure entra caching to improve performance - -### Fixes and Improvements -- Enhancements for token rate limiter to handle small limit values -- Fixed status code and timestamp handling for OTel `http/protobuf` transport protocol -- Fixed an issue with `multipart/form-data` requests when the data contained a `model` field -- Added support for multiple checks of same type under a single guardrail - - - - - -## v1.15.8 ---- - -### Provider Updates -- **Anthropic**: - - **Fixed** `count_tokens` endpoint mapping -- **AWS Bedrock**: - - Added empty `usage` object in `message_start` event for unified messages endpoint stream response to ensure compliance with Anthropic and avoid Claude Code errors - -### Fixes and Improvements -- Added handling for empty span ID in OTel log transformation -- Added support for fetching base model from inference profile in unified batches output handler - - - - - -## v1.15.7 ---- - -### Provider Updates -- **AWS Bedrock**: - - **Fixed** `cache_control` parameter mapping for tools in `/messages` endpoint. - - **Fixed** streaming response handling for `tool_use`, `thinking`, and `redacted_thinking` content_block_start events in `/messages` endpoint. - - - - - -## v1.15.6 ---- - -### Features -- Support pricing for image editing models, currently supported for providers: `openai`, `azure-openai`, and `azure-foundry` -- Support `input_audio` parameter for vertex provider. -- Allow pass through parameters for bedrock batch create endpoint. - -### Fixes -- Handle batch fetch errors gracefully for batch output endpoint. - - - - - -## v1.15.5 ---- - -### OTel Exporter Enhancements -- Added support for `http/protobuf` transport protocol. -- Set `OTEL_EXPORTER_OTLP_PROTOCOL` environment variable to `http/protobuf` to use this. - -### Fixes -- **Prisma AIRS Guardrail**: Fixed an issue with guardrail registration during initialization. - - - - - -## v1.15.4 ---- - -### Provider Updates -- **OpenAI and Azure-OpenAI**: - - Fixed `parallel_tool_calls` parameter mapping for `/responses` API. - - Added support for additional parameters for `/responses` API:: `max_tool_calls`, `safety_identifier` and `top_logprobs` - - - - - -## v1.15.3 ---- - -### Provider Updates -- **Vertex AI**: Handle empty responses returned by the provider -- **AWS Bedrock**: - - Added support for `APAC` cross region inference profiles - - Added support for `performance_config` parameter which will be passed as-is to the provider as `performanceConfig` parameter -- **Azure Foundry and Github**: Updated the parameter mapping to support all the latest OpenAI compatible chat completions parameters -- **OpenAI and Azure-OpenAI**: Updated the tokenizer to support streaming request token calculation for latest gpt-5 models - - - - - -## v1.15.2 ---- - -### New Features -- KMS Support for file uploads to bedrock/AWS. -- Support custom scope for entra auth to use with deprecated azure serverless models. -- Custom Header support for OTel Export of analytics data - -### Improvements and Fixes -- Support Inference Profiles when uploading files to Bedrock for batches & finetuning. -- Vertex `global` region support. -- Cache cleanup for azure entra and managed identity authentication modes. -- Wait for upstream websocket to be connected for Realtime APIs. - - - - - -## v1.15.1 ---- - -### Improvements and Fixes -- Fixed incorrect the AI Gateway 429 error for token-based rate limiting when used with passthrough requests - - - - - -## v1.15.0 ---- - -### Conditional Router Enhancements -- Conditional router config strategy now supports conditions on request path -- [Documentation Link](/aigw/product/ai-gateway/conditional-routing#more-examples-using-conditional-routing) - -### Unified finish_reason -- Unified `finish_reason` across all providers. By default, the value is mapped to an OpenAI-compatible value. If `x-portkey-strict-openai-compliance` is set to false, the original provider-returned value is retained - -### Gemini 2.5 Flash Image Model -- Gemini 2.5 Flash Image model is now supported -- [Documentation Link](/aigw/integrations/llms/vertex-ai#multiple-modalities-on-chat-completions-endpoint) - -### Unified Count Tokens Endpoint -- Introduced unified endpoint for counting tokens across AWS Bedrock, Vertex AI, and Anthropic - -### Metadata-Based Model Access Guardrail -- Introduced a new guardrail to restrict model access based on metadata key-value pairs - -### New Base Providers -- **Meshy** -- **Tripo3D** - -### Provider Updates -- **Vertex AI**: - - Added support for Mistral models - - Added support for `task_type` and `dimensions` parameters in Vertex AI batch embeddings -- **AWS Bedrock**: Added `video` support in chat completions - - - - - -## v1.14.4 ---- - -### Improvements and Fixes -- Resolved Authorization header conflict for passthrough requests. This issue occurred when the AI Gateway API key was sent in the Authorization header instead of the `x-portkey-api-key` header -- Minor bug fixes for the Azure Foundry provider (azure-ai) in Entra and managed auth modes. Note that this does not affect the Azure OpenAI provider (azure-openai) - - - - - -## v1.14.3 ---- - -### Improvements and Fixes -- Fixed backward compatibility issue for the models endpoint when used with default configs. The new models endpoint will only be used when the incoming request does not have a provider, virtual_key, or a config (default or explicitly sent) with any of these fields. Otherwise, the request will be proxied to the upstream provider endpoint (same as the old flow) - - - - - -## v1.14.2 ---- - -### Regex Replace Guardrail -- Added a new guardrail that can replace regex patterns with a specified string - - - - - -## v1.14.1 ---- - -### Provider Updates -- **Anthropic**: Handle tool index when multiple tools are returned in streaming response -- **OpenAI and Azure OpenAI**: Updated stream handling to log newly introduced fields of the `usage` object -- **AWS Bedrock**: Fixed token calculation error for Bedrock messages response when cache tokens were returned by the provider - -### Improvements and Fixes -- Updated pricing configurations for multiple providers and models -- Fixed an issue where metadata labels for Prometheus metrics were getting dropped - - - - - -## v1.14.0 ---- - -### Performance Improvements -- Multiple performance improvements including: - - Removed redundant JSON operations - - Upgraded Hono framework for better performance - - Removed redundant LLM cache key creation - -### Provider Updates -- **DashScope**: Updated the supported parameters -- **Vertex AI**: Added `timeRangeFilter` support for Google Search tool -- **Fireworks**: - - Handle non-ASCII characters in file upload - - Removed unnecessary response transforms to reduce processing time -- **OpenAI and Azure OpenAI**: Added new parameters for GPT-5 compatibility -- **OpenRouter**: Return reasoning messages, if returned by the model - -### Improvements and Fixes -- Return the correct `timeout` value in webhook guardrail response - - - - - -## v1.13.3 ---- - -### Improvements and Fixes -- Allow AI Gateway API key in `Authorization` header for the unified models endpoint - - - - - -## v1.13.2 ---- - -### Unified Models Endpoint -- Released unified models API which follows OpenAI API specification to list all available models that can be used through the AI Gateway -- [Documentation Link](/aigw/api-reference/models/list-models) - - - - - -## v1.13.1 ---- - -### Analytics Enhancements -- Added new analytics data point to capture granular processing time for requests - -### OpenAI Chat Completions Improvements -- Improvements to eliminate extra processing time for OpenAI chat completions responses - - - - - -## v1.13.0 ---- - -### Unified Messages API -- Released unified messages API which follows Anthropic's messages API specification -- Available for AWS Bedrock, Anthropic, and Vertex AI models -- [Documentation Link](/aigw/product/ai-gateway/universal-api#using-the-anthropic%E2%80%99s-%2Fmessages-route) - - - - - -## v1.12.0...v1.12.5 ---- -### NOTE: -- All builds between v1.12.0 and v1.12.5 were part of the Model Catalog early rollout - -### Model Catalog -- Released Model Catalog support along with the latest API specification updates -- [Documentation Link](/aigw/product/model-catalog) - -### Workspace Budget Limits -- You can now enforce budget and rate limits at the workspace level -- [Documentation Link](/aigw/product/administration/enforce-workspace-budget-and-rate-limits) - -### Circuit Breaker -- Introduced circuit breakers which can be added per-strategy in configurations -- [Documentation Link](/aigw/product/ai-gateway/circuit-breaker) - -### Automatic User Attribution -- Introduced automatic `_user` metadata attribution when User API keys are used - -### Provider Updates -- **DeepSeek**: Added support for `response_format` parameter - -### New Base Providers -- **Qdrant** -- **DashScope** - -### Improvements and Fixes -- Removed headers from `webhook` guardrail response to avoid returning sensitive details - - - - -## v1.11.11 ---- - -### Replication for Analytics Data -- Added support for analytics data replication -- This is only applicable for air-gapped deployments - - - - -## v1.11.10 ---- - -### Provider Updates -- **Fireworks**: Added support for `prompt_cache_max_len` parameter - -### Improvements and Fixes -- Enhancements for unified batch output handling - - - - -## v1.11.9 ---- - -### S3 Log Store Enhancements -- Added support for Object Lock and Retention-enabled buckets -- Required environment variable: `LOG_STORE_OBJECT_LOCK_RETENTION_ENABLED="true"` - -### Unified Finish Reason -- Unified `finish_reason` values across Anthropic and Bedrock models -- If `x-portkey-strict-openai-compliance` is set to `false`, the provider-returned value will be retained - -### Provider Updates -- **Vertex AI**: Added support for cost calculation for fine-tuned models -- **AWS Bedrock**: - - Added support for **Computer Use** tool for Anthropic models - - Handle backward compatibility for Titan G1 embeddings model `encoding_format` parameter - - Fixed token calculation for Bedrock cache read and write tokens -- **Cohere**: Handled null values for embeddings `encoding_format` parameter -- **Azure AI**: `max_completion_tokens` parameter will now be forwarded as-is instead of being mapped to `max_tokens` -- **OpenAI and Azure OpenAI**: Added support for `web_search_options` parameter for chat completions endpoint - -### Improvements and Fixes -- Better handling for management plane synchronization when multiple Gateways are deployed - - - - -## v1.11.8 ---- - -### Unified Batches Improvements -- Optimizations to handle large batch output files -- Use `usage` object returned by Anthropic during batch output processing - -### Unified Finish Reason -- Unified `finish_reason` values across Anthropic and Bedrock models -- If `x-portkey-strict-openai-compliance` is set to `false`, the provider-returned value will be retained - -### Provider Updates -- **AWS Bedrock**: Return `finish_reason` in error response - -### Improvements and Fixes -- Updated pricing configurations for multiple providers and models -- Better error logging for Gateway exceptions - - - - -## v1.11.7 ---- - -### Provider Updates -- **Azure OpenAI**: Added Azure Managed Identity support for Azure Containers - -### Improvements and Fixes -- Fixed an issue with Bedrock Signature calculation for passthrough requests -- Better error logging for Azure Managed Identity errors - - - - -## v1.11.6 ---- - -### Provider Updates -- **Groq**: Added support for `service_tier` parameter in the Groq provider configuration -- **Anthropic**: Added support for Anthropic's prompt caching for tool results and tool use -- **Anthropic**: Fixed multi turn tool calling when arguments to the tool call is empty - -### Improvements and Fixes -- Fixed an issue with Auth enabled Aws Redis Cache with Password and cluster mode -- Handled Webhook Guardrail errors and return verdict with the correct status and error - - - - -## v1.11.5 ---- - -### Guardrails -- Added support for metadata keys plugin to enforce metadata keys from the request. - - - - -## v1.11.4 ---- - -### Provider Updates -- **Bedrock**: Added support for `AssumedRole` for bedrock application inference profiles -- **Bedrock Multimodal Embeddings**: Added support for multimodal embeddings for providers `cohere` and `titan`. -- **Azure Foundry**: Added support for `createTranscription`,`createTranslation`, `imageGeneration`, `batch` and `files` endpoints. -- **Anthropic**: Added Support for computer use tool. -- **Anthropic**: Added support for `file_url` and `mime_type` for `file` content parts in Anthropic requests. -- **VertexAI**: Added support for Gemini/Vertex Thinking mode. - -### Cache Improvements -- Added support for Azure Redis with auth modes `EntraID` and `ManagedIdentity` - -### Fixes And Improvements -- Improvements for Redis Cache - - Added support for separate username and password for Redis Cache. Use `REDIS_USERNAME` and `REDIS_PASSWORD` environment variables. - - Added support for Azure Redis Cache. Use `CACHE_STORE` with `azure-redis` as value. - - Added support for Managed Identity for Azure Managed Redis. - - You can pass `AZURE_REDIS_AUTH_MODE` and `AZURE_REDIS_MANAGED_CLIENT_ID` for a different auth setup. - - Defaults to `AZURE_AUTH_MODE` and `AZURE_MANAGED_CLIENT_ID` if not provided - - Added support for Entra ID for Azure Redis Cache. - - You can pass `AZURE_REDIS_AUTH_MODE` and `AZURE_REDIS_ENTRA_CLIENT_ID`, `AZURE_REDIS_ENTRA_CLIENT_SECRET`, `AZURE_REDIS_ENTRA_TENANT_ID` for a different auth setup. - - Defaults to `AZURE_AUTH_MODE` and `AZURE_ENTRA_CLIENT_ID`, `AZURE_ENTRA_CLIENT_SECRET`, `AZURE_ENTRA_TENANT_ID` if not provided -- **HTTPS Proxy** - - Added HTTPS Proxy support for all the external calls. - - Pass `HTTPS_PROXY` environment variable to enable this feature. -- Added support for virtual key inclusion for custom log if passed in headers. -- Fixed issue with proxy calls not working with configs for some providers. - - - - -## v1.11.3 ---- - -### Observability -- Prometheus Metrics are migrated to use endpoints instead of path for all the metrics - -### Fixes And Improvements -- Added a global error handler for all the unhandled exceptions to prevent server crashes. -- Updated JWT Plugin to validate `iat` field - - - - -## v1.11.2 ---- - -### Fixes And Improvements -- Fixed IRSA Web Identity token handling issue for Log Store that was introduced in v1.11.1 - - - - -## v1.11.1 ---- - -### Provider Updates -- **OpenAI**: Added support for `background` and `service_tier` parameters -- **Azure**: Added support for custom hosts and private links for Azure Plugins -- **Bedrock**: Added native support for `inference profiles` [Ref](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-profiles-use.html) - -### OTel Traces Collector endpoints -- Added new endpoint `/v1/otel/v1/traces` to collect any OTel traces as AI Gateway traces - -### Log Exports available on Data Plane -- Log exports are now available on Data Plane -- Export logs without them being sent to Management Plane -- Please note that, `Dataservice` is required for log exports to work via Data Plane. - -### Fixes And Improvements -- Fixed cache bug where Bedrock and Vertex requests were getting responses from the wrong models -- Added support for fetching environment variables from mounted paths - - - - -## v1.11.0 ---- - -### Provider Updates -- **Bedrock**: Fixed cache token calculation for streaming requests - -### Fixes And Improvements -- Added mTLS support for Gateway to internal services in air-gapped deployments - - - - -## v1.10.23 ---- - -### Provider Updates -- **Vertex**: Added support for `dimensions` parameter in multi-modal embeddings - -### Plugins -- **JWT Plugin**: Added JWT authentication for runtime validation - -### Fixes And Improvements -- Added pricing support for `gpt-image-1` model - - - - -## v1.10.22 ---- -### Enhanced Anthropic PDF Support -- Added transformation logic to support OpenAI-spec-compatible `file` content parts in request. -- Introduced two new the AI Gateway parameters for the `file` content parts: `file_url` and `mime_type`. -- Applicable for Anthropic, Bedrock-Anthropic and VertexAI-Anthropic models. -- [Docs](/aigw/integrations/llms/anthropic#processing-pdfs-with-claude) - -### OTel Configurations -- Added two new environment variables: - - `OTEL_SERVICE_NAME`: Sets the `service.name` resource attribute value. - - `OTEL_RESOURCE_ATTRIBUTES`: Comma-separated `key=value` pairs which will be sent as individual resource attributes. - -### New Providers -- **Ncompass**: Supports chat completions endpoint. -- **Lepton**: Supports chat completions, completions and transcriptions endpoints. -- **Snowflake Cortex**: Supports chat completions endpoint. - -### Provider Updates -- **Groq**: Handled an exception that occurred when `stream_options` was included in the request, because the response transformer was not handling `usage` chunk mapping as expected. -- **Workers AI**: Added support for `/images/generations` route. -- **Openrouter**: Added support for `usage` request parameter and response mapping. -- **Deepinfra**: Handled `ping` event returned in stream and mapped `usage` field returned in the response. -- **VertexAI**: Now returns 400 (instead of 500) for empty model validation errors. - -### Fixes And Improvements -- For providers except OpenAI and Azure-OpenAI, updated the value of `object` field in chat completions response from `chat_completion` to `chat.completion` for OpenAI spec compliance. -- **Proxy (Passthrough) Requests**: Fixed endpoint construction logic, which was affecting a few provider-route combinations. -- **Unified Batches**: Fixed issue where embeddings batch output was being returned as empty. -- **Prometheus**: Fixed issue where `model` label was set as N/A for Bedrock requests. - - - -## v1.10.21 ---- - -### AWS Bedrock Prompt Caching -- Added support for AWS Bedrock prompt caching. -- [Docs Link](/aigw/integrations/llms/bedrock/prompt-caching#how-bedrock-prompt-caching-works) - -### VertexAI Gemini 2.5 Thinking Param Support -- Added support for the thinking settings parameters for VertexAI. - -### General Purpose File Upload For VertexAI -- Documentation coming soon. - -### Provider Updates -- **Azure-OpenAI and Azure Foundry**: Added cost calculation support for OpenAI finetuned models. -- **Bedrock**: - - Handled tool role messages with empty content to avoid validation errors. - - Added `response_format` support for Deepseek partner models -- **VertexAI**: - - Handled the unsupported `$schema` property in tools properties JSON Schema. - - Inference support for fine-tuned Gemini models. -- **Groq**: Added support for translations, transcriptions and speech endpoints. - -### Fixes And Improvements -- Fixed `blocklist` handling for Azure Content Safety guardrail. -- Fixed Fireworks dataset upload validation error. - - - -## v1.10.20 ---- - -### OpenAI Embeddings Latency Improvements -- Improved response handling for OpenAI Embeddings resulting in significant reduction in response processing latency. - -### Strict Metadata Enforcement -- Updated the preference for metadata logging. The new order is `Workspace Default Metadata > API Key Default Metadata > Incoming Request Metadata`. -- This provides better control to organisation and workspace admins. Values set by admins cannot be overridden by request level metadata fields. - -### Strict Default Config Enforcement -- Added support to disable default config override for API keys. If config override is not allowed and user tries to send a new config in request as well, Gateway will throw a 400 error. - -### Provider Updates -- **VertexAI**: If a batch record failed on the provider's end, the error will be retained in the final batch output file. -- **AzureOpenAI**: Fixed URL path construction logic for non-completions requests like batches and files where an extra `/v1` was getting added in the final URL, causing request failures. - -### Fixes And Improvements -- Fixed an edge case where Batches, Files and Fine-tune endpoint threw an error when the passed config had `targets` field with a single virtual key in it. - - - -## v1.10.19 ---- - -### OTel Metrics Push -- Added support for pushing AI Gateway Clickhouse analytics (traces and spans) to OTel collector. -- The following environment variables will be used to configure OTel collector: - - `OTEL_PUSH_ENABLED` - - `OTEL_ENDPOINT` - -### Milvus Vector Store for Semantic Caching -- Added support for Milvus vector store for semantic caching. -- The following Vector stores are now supported: - - `pinecone` - - `milvus` -- The following environment variables will be used to configure the Milvus vector store: - - `VECTOR_STORE` - - `VECTOR_STORE_ADDRESS` - - `VECTOR_STORE_API_KEY` - - `VECTOR_STORE_COLLECTION_NAME` - -### Azure Guardrails Support -- Added support for [Azure Content Safety](https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/harm-categories?tabs=warning) API -- Added support for [PII detection](https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview?tabs=text-pii) with Azure Language Service - -### Prompts Render Endpoint -- Prompts Render endpoint is now a part of the Gateway. It is available at `/v1/prompts/:id/render`. - -### Provider Updates -- **Vertex AI**: Added support for `dimensions` for embeddings - -### Minor Enhancements -- **Prometheus Metric**: Added `portkey_processing_time_excluding_last_byte_ms` metric which provides the AI Gateway processing time excluding the LLM last byte diff latency (`llm_last_byte_diff_duration_milliseconds`). - - - -## v1.10.18 (Redacted) ---- -## Redaction notice -This release introduced a critical bug in budget enforcement. -We are redacting this release and will be releasing a patch with out Workspace Budget and related changes. - -### Workspace Level Usage and Rate Limits -- Organisations can now enforce usage limits for each workspace -- Organisations can now enforce rate limits for each workspace - -### OTel Metrics Push -- Added support for pushing AI Gateway Clickhouse analytics (traces and spans) to OTel collector. -- The following environment variables will be used to configure OTel collector: - - `OTEL_PUSH_ENABLED` - - `OTEL_ENDPOINT` - -### Milvus Vector Store for Semantic Caching -- Added support for Milvus vector store for semantic caching. -- The following Vector stores are now supported: - - `pinecone` - - `milvus` -- The following environment variables will be used to configure the Milvus vector store: - - `VECTOR_STORE` - - `VECTOR_STORE_ADDRESS` - - `VECTOR_STORE_API_KEY` - - `VECTOR_STORE_COLLECTION_NAME` - -### Azure Guardrails Support -- Added support for [Azure Content Safety](https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/harm-categories?tabs=warning) API -- Added support for [PII detection](https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview?tabs=text-pii) with Azure Language Service - -### Provider Updates -- **Vertex AI**: Added support for `dimensions` for embeddings - - - - -## v1.10.17 ---- -### OpenAI and Azure OpenAI Response API -- Added end-to-end support for the Response API. -- Implemented caching for stream requests. -- Introduced cost calculation for tools like `web_search`, `file_search` and `code_execution`. - -### Azure AI Foundry Enhancements -- Updated the existing Azure Inference integration to directly accept endpoints from the Azure Foundry dashboard. - -### Retry Enhancements -- Introduced a new retry setting `use_retry_after_header`. When set to `true`, if the provider returns the `x-retry-after` or `x-retry-after-ms` headers, Gateway will use these headers for retry wait times instead of applying the default exponential backoff for 429 responses. - -### Configurable Default Cache TTL -- Default max cache TTL can now be set at the organisation level. - -### Provider Updates -- **Azure OpenAI**: Added support for `logprobs` and `top_logprobs` request parameters. -- **Perplexity**: Added support for `response_format` and `search_recency_filter` request parameters. -- **AWS Bedrock**: Handled empty assistant tools messages containing only newline characters (`\n\n`). - -### Improvements -- Gateway will now populate the `model` field in responses for the `/chat/completions` API if the providers do not natively return this field, ensuring alignment with the OpenAI signature. - - - - -## v1.10.16 ---- - -### Improvements -- **File Upload**: - - Handle file upload failures for Bedrock in some scenarios -- **Unified Batch API**: - - Return `error_file_id` content in the batch output for failed file uploads for OpenAI and Azure OpenAI providers. - - - - -## v1.10.15 ---- - -### Improvements -- **File Upload**: - - Support for uploading large files to Providers and Data Service -- Allow users to pass custom mime-types in the request body. For example: - -```json - { - "model": "gemini-1.5-pro", - "messages": [ - { - "role": "system", - "content": "You are a helpful assistant!" - }, - { - "role": "user", - "content": [ - { - "type": "text", - "text": "What's in this image?" - }, - { - "type": "image_url", - "image_url": { - "url": "", - "mime_type": "image/jpeg" - } - } - ] - } - ] - } -``` - - - - - -## v1.10.14 ---- -### Enforce Organisation And Workspace Guardrails -- It is now possible to enforce guardrails at organisation and workspace levels, which will be applied to all requests. -- Documentation: [Workspace-Level Guardrails](/aigw/product/administration/enforce-workspace-level-guardrails), [Organisation-Level Guardrails](/aigw/product/administration/enforce-organisation-level-guardrails) - -### Unified Finetuning APIs for Fireworks -- Extended the existing unified finetuning APIs to support Fireworks provider. - -### Pricing Updates -- Added support for calculating **Perplexity search** cost and **Gemini grounding** cost. - -### Updated Unified API Signature For Anthropic Extended Thinking -- Updated the unified API signature for Extended thinking which was introduced in v1.10.12 to ensure that OpenAI compliant field of the response remain untouched regardless of strict_open_ai_compliance flag. -- More Details: - - [Anthropic](/aigw/integrations/llms/anthropic#extended-thinking-reasoning-models) - - [AWS Bedrock](/aigw/integrations/llms/bedrock/aws-bedrock#extended-thinking-reasoning-models-beta) - - [VertexAI](/aigw/integrations/llms/vertex-ai#extended-thinking-reasoning-models-beta) - -### Unified Batches API Improvements -- ```custom_id``` will be preserved in the VertexAI batch output. -- Fixed some issues with batches cost calculation. - -### Logging Updates -- Non-OpenAI compliant fields like groundingMetadata (Gemini Grounding), citations (Perplexity Search) and extended thinking response will now be logged for stream responses. Previously, these fields were not logged specifically for streaming response. - -### Provider Updates -- **Fireworks**: Added support for ```logprobs``` and ```top_logprobs``` parameters. - -### Fixes and Improvements -- Added new environment variable (```AWS_ENDPOINT_DOMAIN```) which can be used to override the default value (```amazonaws.com```) -- Fixed an edge case where before_request_hook failures were not getting flagged with 246 response status code for cached and non-cached stream responses. - - - - -## v1.10.13 ---- -### Unified Batches APIs for VertexAI Embedding -- Added support for batch processing of embeddings with Vertex AI. - -### Provider Updates -- **AWS Bedrock** - - Multi-Turn Conversation With Tools: - - Handled assistant messages where content is set as null and tool_calls are passed. -- **OpenAI** - - Fixed an edge case (introduced in the previous version) which was causing issues in cost calculation of fine-tuned models. - -### Fixes and Improvements -- Fixed batch pricing calculation issue for VertexAI and Anthropic Bedrock models. -- Fixed an edge case where the `x-portkey-retry-attempt-count` response header was set to ```-1``` even when no retries were configured. -- Improved handling to skip stream mode detection for irrelevant request types. For example: stream mode detection should not happen for any GET requests as it is not supported. -- Removed redundant AWS credential fetch failures at boot time. - - - - - -## v1.10.12 ---- -### Real-Time Model Pricing Sync -- Model pricing configs are no longer coupled with gateway builds. -- For hybrid deployments, model pricing configs will be fetched from the management plane. - -### Unified API Signature For Anthropic Thinking -- Introduced a unified API signature to support single-turn and multi-turn conversations with Anthropic Extended Reasoning across Anthropic, AWS Bedrock and VertexAI. -- More Details: - - [Anthropic](/aigw/integrations/llms/anthropic#extended-thinking-reasoning-models) - - [AWS Bedrock](/aigw/integrations/llms/bedrock/aws-bedrock#extended-thinking-reasoning-models-beta) - - [VertexAI](/aigw/integrations/llms/vertex-ai#extended-thinking-reasoning-models-beta) - -### Prometheus Metric Updates -- Added a new metric (```llm_last_byte_diff_duration_milliseconds```) to track LLM last byte latency for chunked JSON responses. -- Added a new label (```stream```) for all metrics. Possible values: 0/1 - -### Guardrails Updates -- **AWS Bedrock**: Added handling to flag regex patterns returned by the guardrail. - -### Provider Updates -- **Azure OpenAI**: Mapped the correct model name from multi-deployment virtual keys. - -### Fixes and Improvements -- The AI Gateway 500s are now logged in the console for debugging. - -### Internal POD to POD HTTPS Support -- Added support for internal POD to POD HTTPS communication. -- This can be enabled by mounting a volume with certificate and key. -- `TLS_KEY_PATH` and `TLS_CERT_PATH` environment variables will be used to fetch the certificate and key from the volume. - - - - - -## v1.10.11 ---- - -### Provider Updates -- **AWS Bedrock** - - Added support for encryption key usage when uploading files to S3. -- **VertexAI**: - - Minor updates to streamline the unified spec for batches and fine-tune APIs. - - Updated pricing for gemini-2.0-flash-lite models. - - Added support for ```webm``` mimeType. -- **Openrouter** - - Mapped the usage object for streaming responses. -- **Azure Inference** - - Replaced `extra-parameters: ignore` with `extra-parameters: drop` due to deprecation by Azure. -- **OpenAI and Azure OpenAI** - - Update pricing for GPT 4.5 models - - - - -## v1.10.10 ---- - -### Unified Finetuning APIs for VertexAI -- Extended the existing unified finetuning APIs to support VertexAI. -- The File-upload and transformations will be done according to the provider requirements. - -### Body Params Support in Conditional Router -- Added support for using ```params``` to specify body fields in [conditional router](/aigw/product/ai-gateway/conditional-routing) queries. Previously, only metadata-based routing was supported. - -### Streaming Cache Responses Optimization -- Increased stream chunk content size from 1 token to 125 tokens for cached responses. This reduces the number of chunks significantly (e.g., 2000 tokens now stream in ~16 chunks instead of 2000 chunks). -- Improved last chunk delivery time. -- In addition to latency improvements, this update reduces unnecessary network overhead caused by the large number of chunks. - -### AWS IRSA-based Authentication Updates -- Switched from the default global STS endpoint to regional STS endpoints (for Bedrock and S3 requests) to ensure proper token generation when the global STS is unavailable from the instance. - -### Provider Updates -- **Anthropic**: - - Better error handling for ```error``` type stream chunks returned by the provider. - - Pricing updates for Claude 3.7 models across Anthropic, Bedrock and VertexAI. - - - - -## v1.10.9 ---- - -### Redis Cache Optimization -- Updated cache implementation to avoid redundant Redis calls to improve overall performance. - -### VertexAI Service Account Token Caching -- Implemented caching for Vertex service account token. Previously, tokens were being regenerated on every request despite having 1-hour validity. -- This will reduce VertexAI request latency by 50-100ms per request. - -### Provider Updates -- **Google and VertexAI** - - Handled tool call response parsing when there is one part tool call and one part text. - - Made the default/empty usage object compliant with OpenAI for streaming response. - - - - -## v1.10.8 ---- - -### Mutator Webhooks -- The existing ```webhook``` plugin now has mutation capability. -- This can be used for use-cases like BYO-PII redaction guardrail. - -### Configurable Timeouts for Guardrails -- It is now possible to set timeout values for Guardrail execution. The current default value is 5 seconds. -- ```timeout``` parameter can be used for all the guardrails that make a fetch call internally. -- It is also possible to store this timeout value in management plane while creating/updating a Guardrail on UI. - -### Provider Updates -- **AzureOpenAI**: Added support for ```stream_options``` parameter. - - - - -## v1.10.7 ---- - -### Fixes and Enhancements -- **Fix**: Allow empty body in POST and PUT requests. Gateway was adding empty object as a default body for POST and PUT requests. This caused issues for APIs like POST assistants cancel or POST batches cancel where the upstream provider does not accept body at all. - - - - -## v1.10.6 ---- - -### Unified Batches APIs for AzureOpenAI -- Extended the unified batches APIs to support AzureOpenAI batching. - -### Provider Updates -- **Deepseek Models**: Added support for Deepseek models across multiple inference providers like Fireworks, Groq and Together. - -### Fixes and Enhancements -- **Chore**: Allow budget exhausted user API keys to view logs. Management plane uses user API keys to fetch UI logs from the Gateway. Budget exhaustion of these keys should not have blocked logs view. - - - - -## v1.10.5 ---- - -### JWT Auth -- Added support for JWT based authentication and authorization. -- Customers can configure their JWKS endpoint or the JWKS JSON. - -### Unified Batches APIs for VertexAI -- Extended the unified batches APIs to support VertexAI batching. - -### Provider Updates -- **Google and VertexAI**: Updated the Grounding implementation to support their new API signatures. [Docs Link](/aigw/integrations/llms/vertex-ai#grounding-with-google-search) -- **AWS Bedrock**: Handle edge cases for AWS Bedrock file uploads. - -### Fixes and Enhancements -- **Logging**: Added exception details like ```cause``` and ``name`` in logs for provider level fetch failures. -- **Caching**: Enabled caching even when the ```debug``` flag is set to false. - - - - -## v1.10.4 ---- - -### PII Redaction Guardrails -- Added PII Redaction Guardrails through multiple guardrail providers: - - AI Gateway Managed - - AWS Bedrock - - Pangea - - Patronus - - Promptfoo -- If any entities were redacted from request/response, the guardrail result object in the final response will contain a flag named ```transformed``` set to true. - -### Request Metadata Logging Updates -- Workspace metadata will now logged on individual request level. - -### New Providers -- Replicate: Now supported for proxy (passthrough) requests. - -### Fixes and Enhancements -- **Guardrails**: Added ability to override default guardrail credentials (stored in management plane) with custom credentials at runtime. - - - - -## v1.10.3 ---- - -### AzureOpenAI Unified Finetuning Support -- Extended the unified finetuning APIs to support AzureOpenAI provider. - -### AWS Bedrock Guardrails -- AWS Bedrock Guardrails are now supported for request/response checks. -- [Here](https://docs.google.com/document/d/1sCeuGi5p03wh56WmHpJvMhi7XV9N68vz-wzYq1RH_OQ/edit?usp=sharing) is a short document which can be used to setup this with the AI Gateway. - -### Virtual Keys for Custom Models/Providers -- It is now possible to configure custom host and custom headers directly in the virtual keys. -- If your custom model's API signature matches any of our existing providers, you can create a virtual key with your custom settings. -- While this functionality was already available, it has now been integrated directly into virtual keys for more streamlined configuration. - -### Prometheus Metric Updates -- Updated the units for LLM request duration histogram metrics to milliseconds. The label has been renamed from ```llm_request_duration_seconds``` to ```llm_request_duration_milliseconds``` -- Added a new metric named ```portkey_request_duration_milliseconds``` to track the AI Gateway's processing latency. - -### New Providers -- Milvus DB: Supported as a passthrough provider. - -### Provider Updates -- **VertexAI and Google Improvements** - - Added ```logprobs``` support compatible with OpenAI format via ```logprobs``` and ```top_logprobs``` parameters - - Added support for experimental Gemini Thinking Models. - - Added tool parameters JSON schema handling to ignore/skip fields which are not compatible with these 2 providers. -- **Anthropic**: Added ```total_tokens``` in stream response to make it compliant with OpenAI spec. - - - - -## v1.10.2 ---- - -### Provider Updates -- **VertexAI**: VertexAI requests that sent the virtual key and config headers separately were failing with a provider 401 error. This was happening specifically for VertexAI requests where the virtual key was sent as a separate header along with a config header. - - - - -## v1.10.1 ---- - -## Unified Finetune APIs -- Added unified finetune APIs for OpenAI, AzureOpenAI, Bedrock and Fireworks. - -### Fixes and Enhancements -- **Code Detection Guardrail Updates**: Added checks for verbose identifiers to detect python and js markdown code blocks. Example: check for python and javascript along with py and js identifiers. - - - - -## v1.10.0 ---- - -### Unified Batches and Files API -- Added unified batching APIs for OpenAI, AWS Bedrock and Cohere -- [Docs Link](/aigw/product/ai-gateway/batches#batches) - -### Improved Batch Management for Analytics Data Inserts -- Improved Clickhouse batch management to prevent log drops. -- Notable reduction in memory usage growth and spikes compared to previous builds. -- We also recommend changing ANALYTICS_STORE env to ```control_plane``` (for hybrid deployments) so that batching/retries can managed by the AI Gateway. - -### Gateway Docker Image Size Reduction: -- Made some updates to the image build process, reducing the size (compressed) from ~275MB to ~75MB. - -### VertexAI Self-Deployed Models (a.k.a Endpoints in Vertex): -- You can now use self-deployed models from VertexAI. This update also supports Vertex-Huggingface models. - -### Shorthand Format For Guardrails In Config: -- Added ```input_guardrails``` and ```output_guardrails``` fields in config which accept array of guardrail slugs. - -### Guardrail Output Explanation -- Guardrails responses now include an ```explanation``` property to clarify why checks passed or failed. -- This property is currently only available for default checks. - -### OpenAI ```developer``` Role Support Across All Providers: -- For OpenAI and AzureOpenAI, the role will be mapped as expected. -- For other providers, the developer role is mapped to the system role (or its equivalent). - -### New Partner Guardrails -- Mistral (mistral.moderateContent): Guard against different type of contents like ```hate_and_discrimination```, ```violence_and_threats```, etc. -- Pangea (pange.textGuard): Guard against malicious content and other undesirable data. - -### Provider Updates -- **Cohere**: Removed unsupported ```stream``` parameter from the Bedrock-Cohere integration. - -### Fixes and Enhancements -- **Image Cost Calculation**: Updated the image calculation logic to handle different quality, size, etc. combinations. -- **ValidURL Guardrail**: Updated the URL extraction logic to handle more edge cases. -- **Prompt Render Error Message**: Prompt render API ```(/render)``` is a management plane API. Added detailed message to highlight this in case a user tries to use this API on their deployed Gateway. - - - - -## v1.9.5 ---- - -### Gemini Grounding Mode Support -- Added Gemini grounding mode support in OpenAI compatible tools format. -- [Docs Link](/aigw/integrations/llms/vertex-ai#grounding-with-google-search) - -### Provider Updates -- **Groq**: Fixed ```finish_reason``` mapping for streaming response. -- **AWS Bedrock**: fixed the index mapping for tool call streaming response. -- **VertexAI**: fixed final ```model``` param mapping for VertexAI Meta partner models. - -### Fixes and Enhancements -- **Proxy (Passthrough) Requests**: fixed audio/* content-type passthrough request handling. - - - - -## v1.9.4 ---- - -### Enhanced Request/Response Logging -- Added comprehensive logging for all request/response phases: - - Original request - - Transformed request - - Original response - - Transformed response - -### Prometheus Metrics Standardization -- Standardized all Prometheus metric labels to use a consistent set: - - `method` - - `route` - - `code` - - `custom_labels` - - `provider` - - `model` - - `source` - -### Provider Updates -- **Ollama and Groq** - - Added support for ```tools```. - - - - -## v1.9.3 ---- - -### Allow All S3-compatible Log Stores -- Added a new LOG_STORE type named ```S3_CUSTOM``` which can be used to integrate any S3-compatible storage service for request logging. -- The custom host for the storage provider can be set in ```LOG_STORE_BASEPATH```. - -### New Provider - AWS Sagemaker -- AWS Sagemaker models can now be used through Gateway as passthrough requests. -- Unified API signature is not yet possible because Sagemaker inherits the request body structure from the underlying model. -- [Docs Link](/aigw/integrations/llms/aws-sagemaker) - - - - -## v1.9.2 ---- - -### Proxy (Passthrough) Request Enhancements -- Added streamlined support for virtual keys and configs in proxy (passthrough) requests. - -### Prompt Labels -- Added support for labelled prompt cache invalidation whenever an update happens on management plane side. -- NOTE: Prompt labels is a management plane change and has no major updates in Gateway apart from cache key invalidation for labelled prompt keys. -- Docs Link - -### S3 Integration Enhancements -- Allow sub-paths in bucket name for logs. - -### Provider Updates -- **Perplexity**: Allow ```citations``` in response if strict_open_ai_compliance flag is set to false. -- **AWS Bedrock** - - Stringify the response tool arguments to make it OpenAI compliant. - - Merge successive user messages to avoid Bedrock errors. -- **Openrouter**: Handle cost calculation when input model is ```openrouter/auto```. -- **Google**: Fix the mapping for ```code``` in error response. - - - - -## v1.9.1 ---- - -### Provider Updates -- **OpenAI and AzureOpenAI** - - For Realtime APIs, the socket close event now retains the original close reason returned by the provider. - - Added support for newly released ```prediction```, ```store```, ```metadata```, ```audio``` and ```modalities``` parameters. -- **AWS Bedrock**: Fixed an issue where an extra newline character was being returned in the AWS Bedrock response. - - - - -## v1.9.0 ---- - -### Dynamic Budgets and Auto Expiry for API Keys and Virtual Keys -- Introduced support for setting dynamic budgets and auto-expiry for API keys and virtual keys. - -### Realtime API Integration -- Added Realtime APIs integration for OpenAI and AzureOpenAI. -- [Docs Link](/aigw/product/ai-gateway/realtime-api) - -### Provider Updates -- **VertexAI**: Fixed structured outputs integration for VertexAI when using JS SDK. The SDK was adding extra fields in the JSON schema that were incompatible with Vertex's API requirements. - - - - - -## v1.8.4 ---- - -### Provider Updates -- **Azure OpenAI**: Added ```encoding_format``` and ```dimensions``` as supported params. - -### Fixes & Enhancements -- Updated the default behaviour to use IMDS/Service account role for Bedrock and S3. - - - - -## v1.8.3 ---- - -### Fixes & Enhancements -- Fixed implementation conflicts of existing AWS AssumeRole implementation with the newly released IRSA (IAM Roles for Service Accounts) Assume Role and IMDS (Instance Metadata Service) Assume Role auth approaches. - - - - - -## v1.8.2 ---- - -### Fixes and Enhancements -- Added a new Prometheus metric to track LLM-only latency. Label name: ```llm_request_duration_seconds``` - - - - - -## v1.8.1 ---- - -### Management Plane Log Store -- Added a new log and analytics store named ```control_plane```. -- Setting LOG_STORE and ANALYTICS_STORE environment variables as ```control_plane``` will route all logs and analytics to the management plane and will eliminate the need of having Clickhouse connection on Gateway. - -### - - - - -## v1.8.0 ---- - -### Bedrock Converse API integration -- Bedrock's /chat/completions have been updated to use Bedrock converse API. -- This enables features like tool calls, vision, etc. for many bedrock models. -- This also removes the hassle of maintaining chat templating logic for llama and mistral models. - -### VertexAI Image Generation -- Added support for Vertex Imagen models. - -### Stable Diffusion v2 Models -- StabilityAI introduced v2 models with a new API signature. Gateway now supports both v1 and v2 models, with internal transformations for different API signatures. -- Supported for both stability-ai and bedrock providers. -- New models: Stable Image Ultra, Core, 3.0 and 3.5. - -### Pydantic SDK Integration for Structured Outputs -- Done for GoogleAI and VertexAI (follows OpenAI) -- We previously added support for structured outputs through REST API. However, SDKs using Pydantic were not supported due to extra fields in the JSON schema. -- Added a dereferencing function that converts JSON schemas from the library to Google-compatible schemas. - -### OpenAI and AzureOpenAI Prompt Cache Pricing -- Added support for handling prompt caching pricing for required models. - -### New Providers -- Lambda (`lambda`): Supports chat completions and completions. - -### Provider Updates -- **Perplexity**: Added the missing [DONE] chunk for stream calls to comply with OpenAI's spec. -- **VertexAI**: Fixed provider name extraction logic for meta models, so users can send it like other partner models (e.g., meta.``). -- **Google**: Added structured outputs support (similar to Vertex-ai). - -### Fixes & Enhancements: -- Exclude files, batches, threads, etc. (all passthrough) from ```llm_cost_sum``` prometheus metric to avoid unnecessary labels. - diff --git a/aigw/guides/use-cases/librechat-web-search.mdx b/aigw/guides/use-cases/librechat-web-search.mdx index 24d4313c..fea60dab 100644 --- a/aigw/guides/use-cases/librechat-web-search.mdx +++ b/aigw/guides/use-cases/librechat-web-search.mdx @@ -254,7 +254,3 @@ Use semantic caching for frequently asked current events questions to reduce API - **Monitor costs** using the AI Gateway's analytics dashboard - **Create specialized configs** for your team's specific use cases - **Set up access controls** with user-specific API keys - -## Support -- Support: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) -- Documentation diff --git a/aigw/help-center/mcp-gateway-troubleshooting.mdx b/aigw/help-center/mcp-gateway-troubleshooting.mdx index af311226..9286f636 100644 --- a/aigw/help-center/mcp-gateway-troubleshooting.mdx +++ b/aigw/help-center/mcp-gateway-troubleshooting.mdx @@ -157,9 +157,8 @@ Users have to re-authorize a connected server far more often than expected. **What to know:** access tokens are short-lived (about 1 hour) and refresh tokens are long-lived (about 30 days); the gateway refreshes them automatically, so you shouldn't need to reconnect often. **Fix:** -1. Make sure you're on the latest gateway version (especially for self-hosted deployments). +1. Make sure you're on the latest gateway version (especially for hybrid deployments). 2. If the server is a Google one, see the next section. -3. If reconnects continue, [contact support](#still-stuck) with the server slug and timestamps so the team can review the token refresh logs. ### A Google MCP server asks you to reconnect every hour @@ -212,24 +211,60 @@ If you run the gateway yourself (instead of `aigw.portkey.ai/m`): 3. **Set `MCP_GATEWAY_BASE_URL`** to your public gateway URL (for example `https://`). The gateway uses this to construct callback and discovery URLs—if it's wrong, OAuth and discovery fail. 4. Make sure your ingress forwards the `Host` header so the server advertises its public URL (not an internal address). ---- +With `SERVER_MODE=unified` (gateway 2.20.0 or later), both gateways share one port. Use `https:///m/{slug}/mcp` as the client URL in step 1. `/.well-known/*` and `/oauth/*` stay at the root, not under `/m`. -## Enterprise: `external_auth_config` returns 403 +### Client registration fails with "Invalid API Key. Error Code: 03" -`external_auth_config` (where the gateway brokers your own identity provider directly) is available on **self-hosted/enterprise** deployments only and is blocked on managed SaaS. +An API key works, but connecting without one fails before any login page opens. Claude Code reports it as: -**Fix:** on SaaS, use `oauth_metadata` with a pre-registered OAuth client instead. See [External OAuth](/aigw/product/mcp-gateway/authentication/external-oauth). +``` +SDK auth failed: Dynamic Client Registration rejected (HTTP 401): +{"status":"failure","message":"Portkey Error: Invalid API Key. Error Code: 03", ...} +``` + +**Why:** that response body comes from the AI Gateway, not the MCP Gateway. The MCP Gateway's `/.well-known/*`, `/oauth/register` and CORS preflight (`OPTIONS`) endpoints do not require an API key. When the MCP Gateway itself rejects a request, it answers `{"error":"unauthorized", ...}` with a `WWW-Authenticate` header. So the OAuth requests are reaching the AI Gateway service. API-key access still works because only the `/{slug}/mcp` path is routed correctly. + +**Confirm:** these checks are for `SERVER_MODE=all` or `mcp`, where the MCP Gateway has its own port (`8788`). In unified mode one port serves both paths, so this routing fault does not occur. + +1. Fetch the authorization server metadata and check that every URL in it uses your public MCP Gateway host: + + ```bash + curl https:///.well-known/oauth-authorization-server/{slug}/mcp + ``` + +2. Register a test client at the `registration_endpoint` it returns. Expect `201` with a `client_id`: + + ```bash + curl -i -X POST -H 'Content-Type: application/json' \ + -d '{"client_name":"probe","redirect_uris":["http://localhost:33418/callback"],"token_endpoint_auth_method":"none","grant_types":["authorization_code","refresh_token"],"response_types":["code"]}' + ``` + +3. If that returns the `401` above, repeat it from inside the gateway pod, which bypasses the load balancer and ingress: + + ```bash + kubectl exec -- wget -qO- --header 'Content-Type: application/json' \ + --post-data '{"client_name":"probe","redirect_uris":["http://localhost:33418/callback"],"token_endpoint_auth_method":"none","grant_types":["authorization_code","refresh_token"],"response_types":["code"]}' \ + http://localhost:8788/oauth/register + ``` + + A JSON body with a `client_id` here means the gateway is fine and the Service, load balancer or ingress is sending MCP traffic to the AI Gateway port (`8787`). + +**Fix:** + +- Set `MCP_GATEWAY_BASE_URL` to the public MCP Gateway URL, including the scheme and no path, then restart the pods. It is read at startup. +- With `SERVER_MODE=all`, route **every** path on the MCP Gateway host to the MCP port (`8788`), not only `/{slug}/mcp`. `/.well-known/*` and `/oauth/*` must reach it too. +- Or switch to `SERVER_MODE=unified`, which serves everything on one port and needs no host-based routing. +- In the client, remove the server and add it again with no API key header, so it starts a fresh OAuth flow. --- -## Still stuck? +## Hybrid only: `external_auth_config` returns 403 + +`external_auth_config` (where the gateway brokers your own identity provider directly) is available on **hybrid** deployments only and is blocked on managed SaaS. -Contact [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) with: +**Fix:** on SaaS, use `oauth_metadata` with a pre-registered OAuth client instead. See [External OAuth](/aigw/product/mcp-gateway/authentication/external-oauth). -- Your workspace ID -- The MCP server slug -- The full error message (and `request_id` if shown) -- The MCP client you're using (Claude Desktop, Cursor, etc.) +--- ## Related diff --git a/aigw/help-center/you-do-not-have-enough-permissions.mdx b/aigw/help-center/you-do-not-have-enough-permissions.mdx index 56ffc772..be1e2fa2 100644 --- a/aigw/help-center/you-do-not-have-enough-permissions.mdx +++ b/aigw/help-center/you-do-not-have-enough-permissions.mdx @@ -21,7 +21,7 @@ This means your API key is missing the required **permission scope**. - Test requests in Playground, Prompt Studio, or Model Catalog + Test requests in Playground or Model Catalog Chat completions, embeddings, files, batches, or other inference calls @@ -123,12 +123,12 @@ Enable the specific scopes needed when creating or editing an API key: ## Playground and UI test requests -**Most common cause of AB03.** Start here if Playground, Prompt Studio, or Model Catalog test requests fail. +**Most common cause of AB03.** Start here if Playground or Model Catalog test requests fail. ### Why it happens -Playground, Prompt Studio, and Model Catalog test requests require a **Workspace User API key**. Service keys don't work — the UI needs to authenticate your individual session. +Playground and Model Catalog test requests require a **Workspace User API key**. Service keys don't work — the UI needs to authenticate your individual session. ### Common triggers @@ -187,7 +187,6 @@ The `completions.write` scope gates **all** data-plane endpoints — not just ch | `/v1/fine-tuning/jobs` | Fine-tuning jobs | | `/v1/realtime` | Realtime WebSocket | | `/v1/models` | List models (also works with `virtual_keys.list`) | -| `/v1/prompts/:id/completions` | Prompt completions | | `/v1/gateway/tokenize` | Tokenization | | `/v1/decisions` | Typed judgments | @@ -394,14 +393,6 @@ If none of the above scenarios match, work through these checks: **Org security settings override scopes.** Organisation admins can restrict members from viewing API keys, providers, or configs via **Admin Settings > Security**. When active, affected users get AB03 even with correct scopes and key type. Check with your org admin if everything looks correct but you still see the error. -### 5. Still stuck? - -Contact [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) with: -- Workspace ID -- Full error response (including `request_id`) -- Endpoint called -- API key type (Admin / Workspace User / Workspace Service) - --- ## Complete scope reference diff --git a/aigw/integrations/agents.mdx b/aigw/integrations/agents.mdx index 58348ca3..5dbb15f1 100644 --- a/aigw/integrations/agents.mdx +++ b/aigw/integrations/agents.mdx @@ -126,14 +126,10 @@ Access a dedicated section to view records of action executions, including param azure -### 6\. Prompt Management - -Use the AI Gateway as a centralized hub to store, version, and experiment with your agent's prompts across multiple LLMs. Easily modify your prompts and run A/B tests without worrying about the breaking prod. - -### 7\. [Continuous Improvement](/aigw/product/observability/feedback) +### 6\. [Continuous Improvement](/aigw/product/observability/feedback) Improve your Agent runs by capturing qualitative & quantitative user feedback on your requests, and then using that feedback to make your prompts AND LLMs themselves better. -### 8\. [Security & Compliance](/aigw/product/enterprise-offering/security) +### 7\. [Security & Compliance](/aigw/product/enterprise-offering/security) Set budget limits on provider API keys and implement fine-grained user roles and permissions for both the app and the AI Gateway APIs. diff --git a/aigw/integrations/agents/agentcore.mdx b/aigw/integrations/agents/agentcore.mdx index 94d6f346..29596218 100644 --- a/aigw/integrations/agents/agentcore.mdx +++ b/aigw/integrations/agents/agentcore.mdx @@ -3,7 +3,7 @@ title: "AWS AgentCore" description: "Run gateway-powered agents inside Amazon Bedrock AgentCore" --- -Amazon Bedrock AgentCore is AWS's agentic platform for executing, scaling, and governing AI agents. Because AgentCore can host any OpenAI-compatible framework (Strands, OpenAI Agents, LangGraph, Google ADK, custom code), you can plug Prisma AIRS AI Gateway in as the LLM gateway to unlock multi-provider routing, deep observability, and enterprise guardrails without changing your agent logic. +Amazon Bedrock AgentCore is AWS's agentic platform for executing, scaling, and governing AI agents. Because AgentCore can host any OpenAI-compatible framework (Strands, OpenAI Agents, LangGraph, Google ADK, custom code), you can plug Prisma AIRS AI Gateway in as the LLM gateway to unlock multi-provider routing, deep observability, and guardrails without changing your agent logic. **What you get with this integration** - **Unified gateway** for 3,000+ models while keeping AgentCore's runtime, gateway, and memory services intact diff --git a/aigw/integrations/agents/agno-ai.mdx b/aigw/integrations/agents/agno-ai.mdx index b67f1a27..059114c7 100644 --- a/aigw/integrations/agents/agno-ai.mdx +++ b/aigw/integrations/agents/agno-ai.mdx @@ -5,7 +5,7 @@ description: "Use Prisma AIRS AI Gateway with Agno to build production-ready aut ## Introduction -Agno is a powerful framework for building autonomous AI agents that can reason, use tools, maintain memory, and access knowledge bases. The AI Gateway enhances Agno agents with enterprise-grade capabilities for production deployments. +Agno is a powerful framework for building autonomous AI agents that can reason, use tools, maintain memory, and access knowledge bases. The AI Gateway enhances Agno agents with production-grade capabilities. The AI Gateway transforms your Agno agents into production-ready systems by providing: @@ -14,7 +14,7 @@ The AI Gateway transforms your Agno agents into production-ready systems by prov - **Built-in reliability** with fallbacks, retries, and load balancing - **Cost tracking and optimization** across all agent operations - **Advanced guardrails** for safe and compliant agent behaviour -- **Enterprise governance** with budget controls and access management +- **Governance** with budget controls and access management Learn more about Agno's agent framework and core concepts @@ -248,57 +248,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in Agno Agents - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your Agno agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It's where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - -1. Iteratively develop prompts before using them in your agents -2. Test prompts with different variables and models -3. Compare outputs between different prompt versions -4. Collaborate with team members on prompt development - -This visual environment makes it easier to craft effective prompts for each step in your Agno agent's workflow. - - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your request - - - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are an AI assistant helping with {{task_type}}. - -User question: {{user_input}} - -Please respond in a {{tone}} tone and include {{required_elements}}. -``` - -When rendering, simply pass the variables: - - - - -### 4. Guardrails for Safe Agents +### 3. Guardrails for Safe Agents Guardrails ensure your Agno agents operate safely and respond appropriately in all situations. @@ -352,7 +302,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your Agno agents using the AI Gateway's metadata system. @@ -379,7 +329,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Agents +### 5. Caching for Efficient Agents Implement caching to make your Agno agents more efficient and cost-effective: @@ -445,7 +395,7 @@ Semantic caching considers the contextual similarity between input requests, cac -### 7. Model Interoperability: using different LLMs +### 6. Model Interoperability: using different LLMs One of the AI Gateway's key strengths is providing access to 3,000+ LLMs through a unified interface. Here's how to use different providers with Agno: @@ -585,19 +535,19 @@ if __name__ == "__main__": ) ``` -## Set Up Enterprise Governance for Agno AI agents +## Set Up Governance for Agno AI agents -**Why Enterprise Governance?** +**Why Governance?** If you are using Agno AI agents inside your orgnaization, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. -**Enterprise Implementation Guide** +**Implementation Guide** The AI Gateway allows you to use 3,000+ LLMs with your Agno AI agents setup, with minimal configuration required. Let's set up the core components in the AI Gateway that you'll need for integration. @@ -662,7 +612,7 @@ Here's a basic configuration to load-balance requests to OpenAI and Anthropic: } ``` -Create your config on the [Configs page](https://stratacloudmanager.paloaltonetworks.com/) in Strata Cloud Manager. You'll need the config ID for connecting to Cline's setup. +Create your config on the [Configs page](https://stratacloudmanager.paloaltonetworks.com/) in Strata Cloud Manager. You'll need the config ID for connecting your Agno agents. Configs can be updated anytime to adjust controls without affecting running applications. @@ -690,7 +640,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw ### Step 4: Deploy & Monitor -After distributing API keys to your engineering teams, your enterprise-ready Cline setup is ready to go. Each developer can now use their designated API keys with appropriate access levels and budget controls. +After distributing API keys to your engineering teams, your Agno setup is ready to go. Each developer can now use their designated API keys with appropriate access levels and budget controls. Apply your governance setup using the integration steps from earlier sections Monitor usage in Strata Cloud Manager: - Cost tracking by engineering team @@ -702,7 +652,7 @@ Monitor usage in Strata Cloud Manager: -### Enterprise Features Now Available +### What's Available **Agno AI agents now has:** - Departmental budget controls @@ -717,7 +667,7 @@ Monitor usage in Strata Cloud Manager: -The AI Gateway adds production-grade features to Agno agents including comprehensive observability (traces, logs, analytics), reliability (fallbacks, retries, load balancing), access to 3,000+ LLMs, cost management, and enterprise governance - all without changing your agent logic. +The AI Gateway adds production-grade features to Agno agents including comprehensive observability (traces, logs, analytics), reliability (fallbacks, retries, load balancing), access to 3,000+ LLMs, cost management, and governance - all without changing your agent logic. diff --git a/aigw/integrations/agents/autogen.mdx b/aigw/integrations/agents/autogen.mdx index 1eab3510..201c0bca 100644 --- a/aigw/integrations/agents/autogen.mdx +++ b/aigw/integrations/agents/autogen.mdx @@ -215,57 +215,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in Autogen Agents - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your Autogen agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It's where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - -1. Iteratively develop prompts before using them in your agents -2. Test prompts with different variables and models -3. Compare outputs between different prompt versions -4. Collaborate with team members on prompt development - -This visual environment makes it easier to craft effective prompts for each step in your Autogen agent's workflow. - - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your request - - - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are an AI assistant helping with {{task_type}}. - -User question: {{user_input}} - -Please respond in a {{tone}} tone and include {{required_elements}}. -``` - -When rendering, simply pass the variables: - - - - -### 4. Guardrails for Safe Autogen Agents +### 3. Guardrails for Safe Autogen Agents Guardrails ensure your Autogen agents operate safely and respond appropriately in all situations. @@ -319,7 +269,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your Autogen agents using the AI Gateway's metadata system. @@ -346,7 +296,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Agents +### 5. Caching for Efficient Agents Implement caching to make your Autogen agents more efficient and cost-effective: @@ -379,7 +329,7 @@ Simple caching performs exact matches on input prompts, caching identical reques -### 7. Model Interoperability: using different LLMs +### 6. Model Interoperability: using different LLMs One of the AI Gateway's key strengths is providing access to 3,000+ LLMs through a unified interface. Here's how to use different providers with Autogen: @@ -420,19 +370,19 @@ model_client = OpenAIChatCompletionClient( -## Set Up Enterprise Governance for Autogen agents +## Set Up Governance for Autogen agents -**Why Enterprise Governance?** +**Why Governance?** If you are using Autogen agents inside your organisation, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. -**Enterprise Implementation Guide** +**Implementation Guide** The AI Gateway allows you to use 3,000+ LLMs with your Autogen setup, with minimal configuration required. Let's set up the core components in the AI Gateway that you'll need for integration. @@ -525,7 +475,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw ### Step 4: Deploy & Monitor -After distributing API keys to your engineering teams, your enterprise-ready Autogen setup is ready to go. Each developer can now use their designated API keys with appropriate access levels and budget controls. +After distributing API keys to your engineering teams, your Autogen setup is ready to go. Each developer can now use their designated API keys with appropriate access levels and budget controls. Apply your governance setup using the integration steps from earlier sections Monitor usage in Strata Cloud Manager: - Cost tracking by engineering team @@ -537,7 +487,7 @@ Monitor usage in Strata Cloud Manager: -### Enterprise Features Now Available +### What's Available **Autogen agents now have:** - Departmental budget controls @@ -552,7 +502,7 @@ Monitor usage in Strata Cloud Manager: -The AI Gateway adds production-grade features to Autogen agents including comprehensive observability (traces, logs, analytics), reliability (fallbacks, retries, load balancing), access to 3,000+ LLMs, cost management, and enterprise governance - all without changing your agent logic. +The AI Gateway adds production-grade features to Autogen agents including comprehensive observability (traces, logs, analytics), reliability (fallbacks, retries, load balancing), access to 3,000+ LLMs, cost management, and governance - all without changing your agent logic. diff --git a/aigw/integrations/agents/crewai.mdx b/aigw/integrations/agents/crewai.mdx index 55a53ccb..edb2bc90 100644 --- a/aigw/integrations/agents/crewai.mdx +++ b/aigw/integrations/agents/crewai.mdx @@ -18,7 +18,6 @@ The AI Gateway enhances CrewAI with production-readiness features, turning your - **Cost tracking and optimization** to manage your AI spend - **Access to 3,000+ LLMs** through a single integration - **Guardrails** to keep agent behaviour safe and compliant -- **Version-controlled prompts** for consistent agent performance Learn more about CrewAI's core concepts and features @@ -152,76 +151,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in CrewAI - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your CrewAI agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It's where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - -1. Iteratively develop prompts before using them in your agents -2. Test prompts with different variables and models -3. Compare outputs between different prompt versions -4. Collaborate with team members on prompt development - -This visual environment makes it easier to craft effective prompts for each step in your CrewAI agents' workflow. - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your code: - -```python -# Use a specific prompt version -prompt_data = portkey_admin.prompts.render( - prompt_id="YOUR_PROMPT_ID@version_number", - variables={ - "agent_role": "Senior Research Scientist", - "agent_goal": "Discover groundbreaking insights" - } -) -``` - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are a {{agent_role}} with expertise in {{domain}}. - -Your mission is to {{agent_goal}} by leveraging your knowledge -and experience in the field. - -Always maintain a {{tone}} tone and focus on providing {{focus_area}}. -``` - -When rendering, simply pass the variables: - -```python -prompt_data = portkey_admin.prompts.render( - prompt_id="YOUR_PROMPT_ID", - variables={ - "agent_role": "Senior Research Scientist", - "domain": "artificial intelligence", - "agent_goal": "discover groundbreaking insights", - "tone": "professional", - "focus_area": "practical applications" - } -) -``` - - - -### 4. Guardrails for Safe Crews +### 3. Guardrails for Safe Crews Guardrails ensure your CrewAI agents operate safely and respond appropriately in all situations. @@ -249,7 +179,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your CrewAI agents using the AI Gateway's metadata system. @@ -276,7 +206,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Crews +### 5. Caching for Efficient Crews Implement caching to make your CrewAI agents more efficient and cost-effective: @@ -292,7 +222,7 @@ Semantic caching considers the contextual similarity between input requests, cac -### 7. Model Interoperability +### 6. Model Interoperability CrewAI supports multiple LLM providers, and the AI Gateway extends this capability by providing access to over 200 LLMs through a unified interface. You can easily switch between different models without changing your core agent logic: @@ -311,17 +241,17 @@ The AI Gateway provides access to LLMs from providers including: See the full list of LLM providers supported by the AI Gateway -## Set Up Enterprise Governance for CrewAI +## Set Up Governance for CrewAI -**Why Enterprise Governance?** +**Why Governance?** If you are using CrewAI inside your organisation, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. @@ -460,7 +390,7 @@ Here's a basic configuration to route requests to OpenAI, specifically using GPT ### Step 4: Deploy & Monitor - After distributing API keys to your team members, your enterprise-ready CrewAI setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. + After distributing API keys to your team members, your CrewAI setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -472,7 +402,7 @@ Here's a basic configuration to route requests to OpenAI, specifically using GPT -### Enterprise Features Now Available +### What's Available **Your CrewAI integration now has:** - Departmental budget controls - Model access governance diff --git a/aigw/integrations/agents/langgraph.mdx b/aigw/integrations/agents/langgraph.mdx index 5b414db5..4515d874 100644 --- a/aigw/integrations/agents/langgraph.mdx +++ b/aigw/integrations/agents/langgraph.mdx @@ -18,7 +18,6 @@ The AI Gateway enhances LangGraph with production-readiness features, turning yo - **Cost tracking and optimization** to manage your AI spend - **Access to 3,000+ LLMs** through a single integration - **Guardrails** to keep agent behaviour safe and compliant -- **Version-controlled prompts** for consistent agent performance Learn more about LangGraph's core concepts and features @@ -416,73 +415,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in LangGraph - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your LangGraph agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It's where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - -1. Iteratively develop prompts before using them in your agents -2. Test prompts with different variables and models -3. Compare outputs between different prompt versions -4. Collaborate with team members on prompt development - -This visual environment makes it easier to craft effective prompts for each step in your LangGraph agent's workflow. - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your code: - -```python -# Use a specific prompt version -prompt_data = portkey_admin.prompts.render( - prompt_id="YOUR_PROMPT_ID@version_number", - variables={ - "user_input": "Tell me about quantum computing" - } -) -``` - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are an AI assistant specialized in {{agent_role}}. - -User question: {{user_input}} - -Please respond in a {{tone}} tone and include {{required_elements}}. -``` - -When rendering, simply pass the variables: - -```python -prompt_data = portkey_admin.prompts.render( - prompt_id="YOUR_PROMPT_ID", - variables={ - "agent_role": "search navigator", - "user_input": "Find information about climate change", - "tone": "informative", - "required_elements": "recent scientific findings" - } -) -``` - - - -### 4. Guardrails for Safe Agents +### 3. Guardrails for Safe Agents Guardrails ensure your LangGraph agents operate safely and respond appropriately in all situations. @@ -510,7 +443,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your LangGraph agents using the AI Gateway's metadata system. @@ -537,7 +470,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Agents +### 5. Caching for Efficient Agents Implement caching to make your LangGraph agents more efficient and cost-effective: @@ -553,7 +486,7 @@ Semantic caching considers the contextual similarity between input requests, cac -### 7. Model Interoperability +### 6. Model Interoperability LangGraph works with multiple LLM providers, and the AI Gateway extends this capability by providing access to over 200 LLMs through a unified interface. You can easily switch between different models without changing your core agent logic: @@ -572,17 +505,17 @@ The AI Gateway provides access to LLMs from providers including: See the full list of LLM providers supported by the AI Gateway -## Set Up Enterprise Governance for LangGraph +## Set Up Governance for LangGraph -**Why Enterprise Governance?** +**Why Governance?** If you are using LangGraph inside your organisation, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. @@ -714,7 +647,7 @@ Here's a basic configuration to route requests to OpenAI, specifically using GPT ### Step 4: Deploy & Monitor - After distributing API keys to your team members, your enterprise-ready LangGraph setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. + After distributing API keys to your team members, your LangGraph setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -726,7 +659,7 @@ Here's a basic configuration to route requests to OpenAI, specifically using GPT -### Enterprise Features Now Available +### What's Available **Your LangGraph integration now has:** - Departmental budget controls - Model access governance diff --git a/aigw/integrations/agents/livekit.mdx b/aigw/integrations/agents/livekit.mdx index a7428049..02dbeffa 100644 --- a/aigw/integrations/agents/livekit.mdx +++ b/aigw/integrations/agents/livekit.mdx @@ -1,6 +1,6 @@ --- title: 'LiveKit' -description: "Build production-ready voice AI agents with Prisma AIRS AI Gateway's enterprise features" +description: "Build production-ready voice AI agents with Prisma AIRS AI Gateway's observability, governance, and reliability features" --- import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; @@ -8,20 +8,20 @@ import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; -**Realtime API support is coming soon!** Reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) to be the first to know when LiveKit's realtime model integration with the AI Gateway is available. +**Realtime API support is coming soon!** LiveKit's realtime model integration with the AI Gateway is not yet available. -LiveKit is a powerful platform for building real-time voice and video applications. When combined with the AI Gateway, you get enterprise-grade features that make your LiveKit voice agents production-ready: +LiveKit is a powerful platform for building real-time voice and video applications. When combined with the AI Gateway, you get features that make your LiveKit voice agents production-ready: - **Unified AI Gateway** - Single interface for 3,000+ LLMs with API key management - **Centralized AI observability**: Real-time usage tracking for 40+ key metrics and logs for every request - **Governance** - Real-time spend tracking, set budget limits and RBAC in your LiveKit agents - **Security Guardrails** - PII detection, content filtering, and compliance controls -This guide will walk you through integrating the AI Gateway with LiveKit's STT-LLM-TTS pipeline to build enterprise-ready voice AI agents. +This guide will walk you through integrating the AI Gateway with LiveKit's STT-LLM-TTS pipeline to build production-ready voice AI agents. - If you are an enterprise looking to deploy LiveKit agents in production, [check out this section](#3-set-up-enterprise-governance-for-livekit). + If you are looking to deploy LiveKit agents in production across your organisation, [check out this section](#3-set-up-governance-for-livekit). # 1. Setting up the AI Gateway @@ -189,19 +189,19 @@ Build a simple voice assistant with Python in less than 10 minutes. -# 3. Set Up Enterprise Governance for Livekit +# 3. Set Up Governance for Livekit -**Why Enterprise Governance?** +**Why Governance?** If you are using Livekit inside your orgnaization, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. -**Enterprise Implementation Guide** +**Implementation Guide** @@ -264,7 +264,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw ### Step 4: Deploy & Monitor -After distributing API keys to your team members, your enterprise-ready Livekit setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. +After distributing API keys to your team members, your Livekit setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. Apply your governance setup using the integration steps from earlier sections Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -276,7 +276,7 @@ Monitor usage in Strata Cloud Manager: -### Enterprise Features Now Available +### What's Available **Livekit now has:** - Departmental budget controls @@ -288,7 +288,7 @@ Monitor usage in Strata Cloud Manager: # AI Gateway Features -Now that you have enterprise-grade Livekit setup, let's explore the comprehensive features the AI Gateway provides to ensure secure, efficient, and cost-effective AI operations. +Now that you have your Livekit setup, let's explore the comprehensive features the AI Gateway provides to ensure secure, efficient, and cost-effective AI operations. ### 1. Comprehensive Metrics Using the AI Gateway you can track 40+ key metrics including cost, token usage, response time, and performance across all your LLM providers in real time. You can also filter these metrics based on custom metadata that you can set in your configs. Learn more about metadata here. @@ -313,7 +313,7 @@ Using the AI Gateway, you can add custom metadata to your LLM requests for detai -### 5. Enterprise Access Management +### 5. Access Management @@ -321,11 +321,11 @@ Set and manage spending limits across teams and departments. Control costs with -Enterprise-grade SSO integration with support for SAML 2.0, Okta, Azure AD, and custom providers for secure authentication. +SSO integration with support for SAML 2.0, Okta, Azure AD, and custom providers for secure authentication. -Hierarchical organisation structure with workspaces, teams, and role-based access control for enterprise-scale deployments. +Hierarchical organisation structure with workspaces, teams, and role-based access control for large-scale deployments. @@ -410,6 +410,3 @@ Implement real-time protection for your LLM interactions with automatic detectio - [Read about advanced configs](/aigw/product/ai-gateway/configs) - [Learn about guardrails](/aigw/product/guardrails) - -For enterprise support and custom features for your LiveKit deployment, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/integrations/agents/mastra-agents.mdx b/aigw/integrations/agents/mastra-agents.mdx index 021864c2..6e0c35a3 100644 --- a/aigw/integrations/agents/mastra-agents.mdx +++ b/aigw/integrations/agents/mastra-agents.mdx @@ -18,7 +18,6 @@ The AI Gateway turns your experimental Mastra agents into production-ready syste - **Cost tracking and optimization** to manage your AI spend - **Access to 3,000+ LLMs** through a single integration - **Guardrails** to keep agent behaviour safe and compliant -- **Version-controlled prompts** for consistent agent performance Learn more about Mastra's core concepts and features @@ -234,57 +233,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in Mastra Agents - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your Mastra agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It's where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - -1. Iteratively develop prompts before using them in your agents -2. Test prompts with different variables and models -3. Compare outputs between different prompt versions -4. Collaborate with team members on prompt development - -This visual environment makes it easier to craft effective prompts for each step in your Mastra agent's workflow. - - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your request - - - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are an AI assistant helping with {{task_type}}. - -User question: {{user_input}} - -Please respond in a {{tone}} tone and include {{required_elements}}. -``` - -When rendering, simply pass the variables: - - - - -### 4. Guardrails for Safe Agents +### 3. Guardrails for Safe Agents Guardrails ensure your Mastra agents operate safely and respond appropriately in all situations. @@ -336,7 +285,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your Mastra agents using the AI Gateway's metadata system. @@ -384,7 +333,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Agents +### 5. Caching for Efficient Agents Implement caching to make your Mastra agents more efficient and cost-effective: @@ -446,7 +395,7 @@ Semantic caching considers the contextual similarity between input requests, cac -### 7. Model Interoperability +### 6. Model Interoperability With the AI Gateway, you can easily switch between different LLMs in your Mastra agents without changing your core agent logic. @@ -501,20 +450,20 @@ The AI Gateway provides access to over 200 LLMs through a unified interface, inc See the full list of LLM providers supported by the AI Gateway -## Set Up Enterprise Governance for Mastra Agents +## Set Up Governance for Mastra Agents -**Why Enterprise Governance?** +**Why Governance?** If you are using Mastra agents inside your organisation, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. -**Enterprise Implementation Guide** +**Implementation Guide** The AI Gateway allows you to use 3,000+ LLMs with your Mastra agents setup, with minimal configuration required. Let's set up the core components in the AI Gateway that you'll need for integration. @@ -659,7 +608,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw ### Step 4: Deploy & Monitor - After distributing API keys to your team members, your enterprise-ready Mastra setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. + After distributing API keys to your team members, your Mastra setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. Apply your governance setup using the integration steps from earlier sections. Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -670,7 +619,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw -### Enterprise Features Now Available +### What's Available **Mastra agents now has:** diff --git a/aigw/integrations/agents/openai-agents-ts.mdx b/aigw/integrations/agents/openai-agents-ts.mdx index 2e5ac3b9..99030349 100644 --- a/aigw/integrations/agents/openai-agents-ts.mdx +++ b/aigw/integrations/agents/openai-agents-ts.mdx @@ -17,7 +17,6 @@ The AI Gateway turns your experimental OpenAI Agents into production-ready syste - **Cost tracking and optimization** to manage your AI spend - **Access to 3,000+ LLMs** through a single integration - **Guardrails** to keep agent behaviour safe and compliant -- **Version-controlled prompts** for consistent agent performance Learn more about OpenAI Agents SDK's core concepts @@ -307,77 +306,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in OpenAI Agents - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your OpenAI Agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It's where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - -1. Iteratively develop prompts before using them in your agents -2. Test prompts with different variables and models -3. Compare outputs between different prompt versions -4. Collaborate with team members on prompt development - -This visual environment makes it easier to craft effective prompts for each step in your OpenAI Agents agent's workflow. - - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your code: - -```typescript -// Use a specific prompt version -const promptData = await gatewayClient.prompts.render({ - promptId: "YOUR_PROMPT_ID@version_number", - variables: { - user_input: "Tell me about quantum computing" - } -}); -``` - - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are an AI assistant helping with {{task_type}}. - -User question: {{user_input}} - -Please respond in a {{tone}} tone and include {{required_elements}}. -``` - -When rendering, simply pass the variables: - -```typescript -const promptData = await gatewayClient.prompts.render({ - promptId: "YOUR_PROMPT_ID", - variables: { - task_type: "research", - user_input: "Tell me about quantum computing", - tone: "professional", - required_elements: "recent academic references" - } -}); -``` - - - -### 4. Guardrails for Safe Agents +### 3. Guardrails for Safe Agents Guardrails ensure your OpenAI Agents operate safely and respond appropriately in all situations. @@ -424,7 +353,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your OpenAI Agents using the AI Gateway's metadata system. @@ -471,7 +400,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Agents +### 5. Caching for Efficient Agents Implement caching to make your OpenAI Agents agents more efficient and cost-effective: @@ -519,7 +448,7 @@ Semantic caching considers the contextual similarity between input requests, cac -### 7. Model Interoperability +### 6. Model Interoperability With the AI Gateway, you can easily switch between different LLMs in your OpenAI Agents without changing your core agent logic. @@ -578,19 +507,19 @@ The AI Gateway provides access to over 200 LLMs through a unified interface, inc See the full list of LLM providers supported by the AI Gateway -## Set Up Enterprise Governance for OpenAI Agents +## Set Up Governance for OpenAI Agents -**Why Enterprise Governance?** +**Why Governance?** If you are using OpenAI Agents inside your orgnaization, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. -**Enterprise Implementation Guide** +**Implementation Guide** The AI Gateway allows you to use 3,000+ LLMs with your OpenAI Agents setup, with minimal configuration required. Let's set up the core components in the AI Gateway that you'll need for integration. @@ -739,7 +668,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw ### Step 4: Deploy & Monitor -After distributing API keys to your team members, your enterprise-ready OpenAI Agents setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. +After distributing API keys to your team members, your OpenAI Agents setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. Apply your governance setup using the integration steps from earlier sections Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -751,7 +680,7 @@ Monitor usage in Strata Cloud Manager: -### Enterprise Features Now Available +### What's Available **OpenAI Agents now has:** - Departmental budget controls diff --git a/aigw/integrations/agents/openai-agents.mdx b/aigw/integrations/agents/openai-agents.mdx index 714d77af..74e30220 100644 --- a/aigw/integrations/agents/openai-agents.mdx +++ b/aigw/integrations/agents/openai-agents.mdx @@ -17,7 +17,6 @@ The AI Gateway turns your experimental OpenAI Agents into production-ready syste - **Cost tracking and optimization** to manage your AI spend - **Access to 3,000+ LLMs** through a single integration - **Guardrails** to keep agent behaviour safe and compliant -- **Version-controlled prompts** for consistent agent performance Learn more about OpenAI Agents SDK's core concepts @@ -445,58 +444,7 @@ This configuration will automatically try Claude if the GPT-4o request fails, en -### 3. Prompting in OpenAI Agents - -The AI Gateway's Prompt Engineering Studio helps you create, manage, and optimize the prompts used in your OpenAI Agents. Instead of hardcoding prompts or instructions, use the AI Gateway's prompt rendering API to dynamically fetch and apply your versioned prompts. - - - - - -Prompt Playground is a place to compare, test and deploy perfect prompts for your AI application. It’s where you experiment with different models, test variables, compare outputs, and refine your prompt engineering strategy before deploying to production. It allows you to: - - 1. Iteratively develop prompts before using them in your agents - 2. Test prompts with different variables and models - 3. Compare outputs between different prompt versions - 4. Collaborate with team members on prompt development - - This visual environment makes it easier to craft effective prompts for each step in your OpenAI Agents agent's workflow. - - - - -The Prompt Render API retrieves your prompt templates with all parameters configured: - - - - - -You can: -- Create multiple versions of the same prompt -- Compare performance between versions -- Roll back to previous versions if needed -- Specify which version to use in your request - - - - - -AI Gateway prompts use Mustache-style templating for easy variable substitution: - -``` -You are an AI assistant helping with {{task_type}}. - -User question: {{user_input}} - -Please respond in a {{tone}} tone and include {{required_elements}}. -``` - -When rendering, simply pass the variables: - - - - -### 4. Guardrails for Safe Agents +### 3. Guardrails for Safe Agents Guardrails ensure your OpenAI Agents operate safely and respond appropriately in all situations. @@ -544,7 +492,7 @@ The AI Gateway's guardrails can: Explore the AI Gateway's guardrail features to enhance agent safety -### 5. User Tracking with Metadata +### 4. User Tracking with Metadata Track individual users through your OpenAI Agents using the AI Gateway's metadata system. @@ -571,7 +519,7 @@ This enables: Explore how to use custom metadata to enhance your analytics -### 6. Caching for Efficient Agents +### 5. Caching for Efficient Agents Implement caching to make your OpenAI Agents agents more efficient and cost-effective: @@ -621,7 +569,7 @@ Implement caching to make your OpenAI Agents agents more efficient and cost-effe -### 7. Model Interoperability +### 6. Model Interoperability With the AI Gateway, you can easily switch between different LLMs in your OpenAI Agents without changing your core agent logic. @@ -681,7 +629,7 @@ The AI Gateway provides access to over 200 LLMs through a unified interface, inc See the full list of LLM providers supported by the AI Gateway -### 8. Tracing +### 7. Tracing The AI Gateway provides an opentelemetry compatible backend to store and query your traces. @@ -732,19 +680,19 @@ if __name__ == "__main__": OpenAI Agents SDK natively supports tools that enable your agents to interact with external systems and APIs. The AI Gateway provides full observability for tool usage in your agents: -## Set Up Enterprise Governance for OpenAI Agents +## Set Up Governance for OpenAI Agents -**Why Enterprise Governance?** +**Why Governance?** If you are using OpenAI Agents inside your orgnaization, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. -**Enterprise Implementation Guide** +**Implementation Guide** The AI Gateway allows you to use 3,000+ LLMs with your OpenAI Agents setup, with minimal configuration required. Let's set up the core components in the AI Gateway that you'll need for integration. @@ -893,7 +841,7 @@ For detailed key management instructions, see our [API Keys documentation](/aigw ### Step 4: Deploy & Monitor -After distributing API keys to your team members, your enterprise-ready OpenAI Agents setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. +After distributing API keys to your team members, your OpenAI Agents setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and budget controls. Apply your governance setup using the integration steps from earlier sections Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -905,7 +853,7 @@ Monitor usage in Strata Cloud Manager: -### Enterprise Features Now Available +### What's Available **OpenAI Agents now has:** - Departmental budget controls diff --git a/aigw/integrations/agents/openai-swarm.mdx b/aigw/integrations/agents/openai-swarm.mdx index bf1f5536..e654c809 100644 --- a/aigw/integrations/agents/openai-swarm.mdx +++ b/aigw/integrations/agents/openai-swarm.mdx @@ -74,7 +74,7 @@ By routing your OpenAI Swarm requests through the AI Gateway, you get access to

Access detailed logs of agent executions, function calls, and interactions. Debug and optimize your agents effectively.

- +

Implement budget limits, role-based access control, and audit trails for your agent operations.

@@ -155,9 +155,9 @@ Access a dedicated section to view records of agent executions, including parame -## 6. [Security & Compliance - Enterprise-Ready Controls](/aigw/product/enterprise-offering/security) +## 6. [Security & Compliance Controls](/aigw/product/enterprise-offering/security) -When deploying agents in production, security is crucial. The AI Gateway provides enterprise-grade security features: +When deploying agents in production, security is crucial. The AI Gateway provides these security features: diff --git a/aigw/integrations/agents/pydantic-ai.mdx b/aigw/integrations/agents/pydantic-ai.mdx index 753186bf..b96e50b9 100644 --- a/aigw/integrations/agents/pydantic-ai.mdx +++ b/aigw/integrations/agents/pydantic-ai.mdx @@ -861,17 +861,17 @@ The AI Gateway provides access to LLMs from providers including: See the full list of LLM providers supported by the AI Gateway -## Set Up Enterprise Governance for PydanticAI +## Set Up Governance for PydanticAI -**Why Enterprise Governance?** +**Why Governance?** If you are using PydanticAI inside your organisation, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users -The AI Gateway adds a comprehensive governance layer to address these enterprise needs. Let's implement these controls step by step. +The AI Gateway adds a comprehensive governance layer to address these needs. Let's implement these controls step by step. @@ -1011,7 +1011,7 @@ Here's a basic configuration to route requests to OpenAI, specifically using GPT ### Step 4: Deploy & Monitor - After distributing API keys to your team members, your enterprise-ready PydanticAI setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and governance controls. + After distributing API keys to your team members, your PydanticAI setup is ready to go. Each team member can now use their designated API keys with appropriate access levels and governance controls. Monitor usage in Strata Cloud Manager: - Cost tracking by department @@ -1023,7 +1023,7 @@ Here's a basic configuration to route requests to OpenAI, specifically using GPT -### Enterprise Features Now Available +### What's Available **Your PydanticAI integration now has:** - Team-based model access controls - Usage tracking & attribution diff --git a/aigw/integrations/agents/strands.mdx b/aigw/integrations/agents/strands.mdx index 98a3776f..6d2ea551 100644 --- a/aigw/integrations/agents/strands.mdx +++ b/aigw/integrations/agents/strands.mdx @@ -322,13 +322,13 @@ Override configuration for specific requests without changing the model: --- -## Enterprise Governance +## Governance If you are using Strands inside your organisation, you need to consider several governance aspects: - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing which teams can use specific models - **Usage Analytics**: Understanding how AI is being used across the organisation -- **Security & Compliance**: Maintaining enterprise security standards +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users ### Centralized Key Management @@ -365,19 +365,6 @@ All of this works automatically with your existing Strands agents—no code chan --- -## Contact & Support - - - - Get dedicated SLA-backed support. - - - Join our forums and Slack channel. - - - ---- - ## Resources diff --git a/aigw/integrations/cloud/azure.mdx b/aigw/integrations/cloud/azure.mdx index eba7a9b1..78de9fd0 100644 --- a/aigw/integrations/cloud/azure.mdx +++ b/aigw/integrations/cloud/azure.mdx @@ -3,7 +3,7 @@ title: "Microsoft Azure" description: "Discover how you can build your Gen AI platform on Azure using Prisma AIRS AI Gateway" --- -The AI Gateway supercharges your Azure AI infrastructure with an enterprise-ready production stack. Teams building on Azure achieve **75% faster time-to-market**, **significant cost optimization**, and **4× faster deployments** while maintaining seamless integration with existing Azure investments. +The AI Gateway supercharges your Azure AI infrastructure with a production stack while maintaining seamless integration with existing Azure investments. Our solution empowers you to build robust, scalable AI applications that leverage the full security and compliance capabilities of Azure. @@ -14,24 +14,22 @@ Our solution empowers you to build robust, scalable AI applications that leverag - Simplified deployment and scaling of production AI applications - Full compatibility with Azure's native security model -Schedule a 30-min strategy call - ## What You Can Do Today with the AI Gateway on Azure - -Spin up a managed AI Gateway instance directly inside your Azure subscription for maximum security and control. + +Run the AI Gateway data plane inside your own Azure subscription on AKS or Azure Container Apps, managed from Strata Cloud Manager. -Integrate seamlessly with your existing identity provider for enterprise-grade access control and automated user provisioning. +Integrate seamlessly with your existing identity provider for access control and automated user provisioning. - + Route all your Azure OpenAI requests (chat, vision, function-calling, DALL-E images) through the AI Gateway while retaining full Azure compliance and gaining enhanced observability. Bring any model deployed via Azure AI Studio (formerly AI Foundry) under the AI Gateway for consistent caching, retries, fallbacks, and observability. - + Apply robust, configurable content moderation and PII detection to every AI request—no code changes required in your applications. @@ -44,7 +42,7 @@ Expose the AI Gateway configurations as APIM endpoints, ensuring consistent gove ## Get Started in Minutes -1. **Deploy from the Azure Marketplace**: Launch the AI Gateway directly within your Azure subscription. +1. **Deploy the gateway in your Azure subscription** (optional): Follow the [AKS](/aigw/self-hosting/hybrid-deployments/azure/aks) or [Azure Container Apps](/aigw/self-hosting/hybrid-deployments/azure/aca) hybrid deployment guide. 2. **Connect Microsoft Entra ID for SSO & SCIM**: Follow our [SSO guide](/aigw/product/enterprise-offering/org-management/sso) and [Azure SCIM setup](/aigw/product/enterprise-offering/org-management/scim/azure-ad). 3. **Integrate your LLM**: Link your Azure OpenAI resources or Azure AI Foundry model deployments. 4. **Enable Azure Content Safety Guardrails**: Configure content filters within your AI Gateway Configs. @@ -153,7 +151,7 @@ public class ExampleAzureIntegration ## Why the AI Gateway + Azure? The Strategic Advantages -Enterprises choose to combine the AI Gateway with Microsoft Azure to gain a competitive edge in their AI development. This powerful synergy offers: +Teams choose to combine the AI Gateway with Microsoft Azure to gain a competitive edge in their AI development. This powerful synergy offers: ```mermaid flowchart LR @@ -170,26 +168,16 @@ flowchart LR 1. **Complete Cost Visibility & Control on Azure Spend**: Go beyond standard Azure billing. The AI Gateway allows you to tag every AI request (by application, environment, user, or custom dimension) and export granular metrics to Azure Monitor or your data warehouse for precise cost attribution and budget management. 2. **Unified Governance & Security Across Your Azure AI Landscape**: Enforce organisation-wide policies consistently. From robust content moderation with Azure Content Safety to PII detection, rate limits, and SSO via Microsoft Entra ID, the AI Gateway centralizes governance across all your Azure AI services and AI Gateway workspaces. -3. **Fortified Security with Azure-Native Integration**: Leverage Azure's robust security model. With the AI Gateway's Azure Marketplace deployment or Private Cloud options, secrets remain securely in Azure Key Vault, and AI traffic can be configured to never leave Azure’s backbone, ensuring compliance and data integrity. +3. **Fortified Security with Azure-Native Integration**: Leverage Azure's robust security model. With a hybrid deployment of the AI Gateway in your Azure subscription, secrets remain securely in Azure Key Vault, and AI traffic can be configured to never leave Azure’s backbone, ensuring compliance and data integrity. 4. **Boosted Developer Velocity & Efficiency**: Eliminate redundant setups and streamline AI development. The AI Gateway provides a single, consistent gateway layer, removing the need for per-team Azure OpenAI subscriptions or duplicated infrastructure, allowing your teams to build faster. 5. **Seamless Integration with the Azure Ecosystem**: The AI Gateway is built for Azure. Enjoy native support for Azure OpenAI (chat, vision, function-calling, images), Azure AI Foundry Models, and even familiar Microsoft tooling like the OpenAI C# SDK and Semantic Kernel, all while the AI Gateway handles the complexities of routing, cost attribution, and guardrails. -## Book an Enterprise Demo - - - --- ### Additional Resources -- [Azure OpenAI integration](/integrations/llms/azure-openai) +- [Azure OpenAI integration](/aigw/integrations/llms/azure-openai/azure-openai) - [Azure AI Foundry integration](/aigw/integrations/llms/azure-foundry) -- [Azure Content Safety guardrails](/product/guardrails/azure-guardrails) -- [Azure Private Cloud deployment guide](/product/enterprise-offering/private-cloud-deployments/azure) +- [Azure Content Safety guardrails](/aigw/integrations/guardrails/azure-guardrails) +- [AKS hybrid deployment guide](/aigw/self-hosting/hybrid-deployments/azure/aks) +- [Azure Container Apps hybrid deployment guide](/aigw/self-hosting/hybrid-deployments/azure/aca) diff --git a/aigw/integrations/guardrails/azure-guardrails.mdx b/aigw/integrations/guardrails/azure-guardrails.mdx index 2cfcef36..0e2a4bdb 100644 --- a/aigw/integrations/guardrails/azure-guardrails.mdx +++ b/aigw/integrations/guardrails/azure-guardrails.mdx @@ -154,7 +154,7 @@ Protected Material scans the LLM's response text and checks it against Azure's d **Use Cases** - Compliance with copyright and intellectual property laws - Preventing reproduction of licensed content -- Enterprise content governance +- Content governance For more information on Azure Protected Material Detection, visit the [official documentation](https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/protected-material). diff --git a/aigw/integrations/guardrails/bring-your-own-guardrails.mdx b/aigw/integrations/guardrails/bring-your-own-guardrails.mdx index bc3fc92a..a2d6673a 100644 --- a/aigw/integrations/guardrails/bring-your-own-guardrails.mdx +++ b/aigw/integrations/guardrails/bring-your-own-guardrails.mdx @@ -510,12 +510,3 @@ When triggered on a proxy request, the webhook payload includes `method`, `path` 3. **Default Behaviour**: If your webhook fails to respond within the timeout period, the AI Gateway will default to `verdict: true`. 4. **Event Type Awareness**: When implementing transformations, ensure your webhook checks the `eventType` field to determine whether it's being called before or after the LLM request. - -## Example Implementation - -Check out our Guardrail Webhook implementation on GitHub: - - -## Get Help - -Building custom webhooks? Reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) for support and to share your implementation experiences! diff --git a/aigw/integrations/guardrails/palo-alto-panw-prisma.mdx b/aigw/integrations/guardrails/palo-alto-panw-prisma.mdx index f2a0bc12..56a239b7 100644 --- a/aigw/integrations/guardrails/palo-alto-panw-prisma.mdx +++ b/aigw/integrations/guardrails/palo-alto-panw-prisma.mdx @@ -167,7 +167,7 @@ Prisma AIRS provides multi-layered protection against various AI-specific threat 4. **Policy Updates**: Keep your Prisma AIRS security policies updated based on emerging threats ## Get Support -- Palo Alto Networks support through your enterprise account +- Palo Alto Networks support through your Palo Alto Networks account ## Learn More diff --git a/aigw/integrations/libraries/android-studio.mdx b/aigw/integrations/libraries/android-studio.mdx index 8a3dcdd6..50cef025 100644 --- a/aigw/integrations/libraries/android-studio.mdx +++ b/aigw/integrations/libraries/android-studio.mdx @@ -13,7 +13,7 @@ description: 'Add observability, governance, and reliability to Android Studio w This guide shows how to add the AI Gateway as a **Third-Party Remote Provider** in Android Studio in a few minutes. -For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). +For deployments across teams, see [Set Up Governance](#3-set-up-governance). ## 1. Setup @@ -75,7 +75,7 @@ After adding the AI Gateway as a provider: The models shown are the ones you added in the AI Gateway (e.g. `@openai-prod/gpt-4o`, `@anthropic-prod/claude-3-5-sonnet-20241022`). All traffic is routed through the AI Gateway automatically. -**Fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and use the config’s virtual model in Android Studio. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and use the config’s virtual model in Android Studio. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/anthropic-computer-use.mdx b/aigw/integrations/libraries/anthropic-computer-use.mdx index 3dc6b9d6..dc1b5924 100644 --- a/aigw/integrations/libraries/anthropic-computer-use.mdx +++ b/aigw/integrations/libraries/anthropic-computer-use.mdx @@ -60,7 +60,7 @@ For more information on the computer use tool, please refer to the [Anthropic do # AI Gateway Features -Now that you have enterprise-grade Anthropic Computer Use setup, let's explore the comprehensive features the AI Gateway provides to ensure secure, efficient, and cost-effective AI-assisted development. +Now that you have Anthropic Computer Use set up, let's explore the features the AI Gateway provides to ensure secure, efficient, and cost-effective AI-assisted development. ### 1. Comprehensive Metrics Using the AI Gateway you can track 40+ key metrics including cost, token usage, response time, and performance across all your LLM providers in real time. Filter these metrics by developer, team, or project using custom metadata. @@ -90,7 +90,7 @@ Track coding patterns and productivity metrics with custom metadata: -### 5. Enterprise Access Management +### 5. Access Management @@ -98,7 +98,7 @@ Set and manage spending limits per developer or team. Prevent budget overruns wi -Enterprise-grade SSO integration for seamless developer onboarding and offboarding. +SSO integration for seamless developer onboarding and offboarding. @@ -181,7 +181,3 @@ The AI Gateway provides multiple security layers: **Join our Community** - [GitHub Repository](https://github.com/Portkey-AI) - - -For enterprise support and custom features for your development teams, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/integrations/libraries/anythingllm.mdx b/aigw/integrations/libraries/anythingllm.mdx index e8a36b00..8ef361bb 100644 --- a/aigw/integrations/libraries/anythingllm.mdx +++ b/aigw/integrations/libraries/anythingllm.mdx @@ -13,7 +13,7 @@ AnythingLLM is an all-in-one Desktop & Docker AI application with built-in RAG a This guide shows how to configure AnythingLLM with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). # 1. Setup @@ -71,7 +71,7 @@ Change models by updating the **Chat Model** field: All requests route through the AI Gateway automatically. -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Chat Model to `dummy`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Chat Model to `dummy`. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/claude-code.mdx b/aigw/integrations/libraries/claude-code.mdx index c10b9415..21311b75 100644 --- a/aigw/integrations/libraries/claude-code.mdx +++ b/aigw/integrations/libraries/claude-code.mdx @@ -1,6 +1,6 @@ --- title: 'Claude Code' -description: 'Integrate Prisma AIRS AI Gateway with Claude Code for enterprise-grade AI coding assistance with observability, reliability, and governance' +description: 'Integrate Prisma AIRS AI Gateway with Claude Code for AI coding assistance with observability, reliability, and governance' --- ### Universal provider (any model like gpt5.2, grok, kimi, etc) @@ -40,7 +40,7 @@ Use the same setup as Anthropic, then swap: Claude Code can run up bills fast with agentic loops. See exactly what each team, project, or developer is spending—and set hard budget limits. - Isolate teams with separate workspaces. Each gets its own budget, rate limits, and access controls. Perfect for enterprise multi-team setups. + Isolate teams with separate workspaces. Each gets its own budget, rate limits, and access controls. Suited to multi-team setups. Route through Anthropic, Bedrock, or Vertex AI. Switch providers with a single config change—no developer workflow changes needed. diff --git a/aigw/integrations/libraries/cline.mdx b/aigw/integrations/libraries/cline.mdx index 7fa2b406..8c8741e4 100644 --- a/aigw/integrations/libraries/cline.mdx +++ b/aigw/integrations/libraries/cline.mdx @@ -1,6 +1,6 @@ --- title: 'Cline' -description: 'Add enterprise-grade observability, cost tracking, and governance to your Cline AI coding assistant' +description: 'Add observability, cost tracking, and governance to your Cline AI coding assistant' --- Cline is an AI coding assistant for VS Code. Add Prisma AIRS AI Gateway to get: @@ -13,7 +13,7 @@ Cline is an AI coding assistant for VS Code. Add Prisma AIRS AI Gateway to get: This guide shows how to configure Cline with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). # 1. Setup @@ -74,7 +74,7 @@ Change models by updating the **Model ID** field: All requests route through the AI Gateway automatically. -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model ID to `dummy`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model ID to `dummy`. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/codex.mdx b/aigw/integrations/libraries/codex.mdx index 741e774e..0e425cbc 100644 --- a/aigw/integrations/libraries/codex.mdx +++ b/aigw/integrations/libraries/codex.mdx @@ -13,7 +13,7 @@ description: 'Add usage tracking, cost controls, and security guardrails to Code Configure Codex with the AI Gateway in a few minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). ## Quick setup with CLI @@ -179,7 +179,7 @@ wire_api = "chat" ``` -**Fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set `model` to the config’s virtual model. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set `model` to the config’s virtual model. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/conductor.mdx b/aigw/integrations/libraries/conductor.mdx index ac1d27bd..4d1e3578 100644 --- a/aigw/integrations/libraries/conductor.mdx +++ b/aigw/integrations/libraries/conductor.mdx @@ -11,7 +11,7 @@ description: 'Add usage tracking, cost controls, and security guardrails to Cond - **Governance** — budget limits, usage tracking, and team access controls - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). ## 1. Setup diff --git a/aigw/integrations/libraries/continue-dev.mdx b/aigw/integrations/libraries/continue-dev.mdx index 5ea637ce..3bae26db 100644 --- a/aigw/integrations/libraries/continue-dev.mdx +++ b/aigw/integrations/libraries/continue-dev.mdx @@ -1,6 +1,6 @@ --- title: 'Continue.dev' -description: 'Add enterprise-grade observability, cost tracking, and governance to your Continue.dev AI coding assistant' +description: 'Add observability, cost tracking, and governance to your Continue.dev AI coding assistant' --- Continue.dev is an open-source AI coding assistant for VS Code and JetBrains. Add Prisma AIRS AI Gateway to get: @@ -13,7 +13,7 @@ Continue.dev is an open-source AI coding assistant for VS Code and JetBrains. Ad This guide shows how to configure Continue.dev with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). --- @@ -254,7 +254,7 @@ models: ``` -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set `x-portkey-config` to your Config ID in `requestOptions.headers`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set `x-portkey-config` to your Config ID in `requestOptions.headers`. See [Set Up Governance](#3-set-up-governance) for examples. --- diff --git a/aigw/integrations/libraries/cursor.mdx b/aigw/integrations/libraries/cursor.mdx index 9dd6a74b..629ed0d7 100644 --- a/aigw/integrations/libraries/cursor.mdx +++ b/aigw/integrations/libraries/cursor.mdx @@ -15,7 +15,7 @@ description: "Add observability, governance, and reliability to Cursor with Pris - For enterprise governance, see [Enterprise Governance](#enterprise-governance). + For governance, see [Set Up Governance](#3-set-up-governance). ## 1. Setup diff --git a/aigw/integrations/libraries/goose.mdx b/aigw/integrations/libraries/goose.mdx index 7c7134f5..1ed651e4 100644 --- a/aigw/integrations/libraries/goose.mdx +++ b/aigw/integrations/libraries/goose.mdx @@ -13,7 +13,7 @@ Goose is an open source AI agent that automates engineering tasks. Add Prisma AI This guide shows how to configure Goose with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). # 1. Setup @@ -70,7 +70,7 @@ Change models by updating the model field: All requests route through the AI Gateway automatically. -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model to `dummy`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model to `dummy`. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/janhq.mdx b/aigw/integrations/libraries/janhq.mdx index b4a67cd0..7a0f714e 100644 --- a/aigw/integrations/libraries/janhq.mdx +++ b/aigw/integrations/libraries/janhq.mdx @@ -13,7 +13,7 @@ Jan is an open source alternative to ChatGPT that runs 100% offline on your comp This guide shows how to configure Jan with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). # 1. Setup @@ -67,7 +67,7 @@ Change models by updating the model selection: All requests route through the AI Gateway automatically. -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model to `dummy`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model to `dummy`. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/langchain-js.mdx b/aigw/integrations/libraries/langchain-js.mdx index 3d87d332..24535904 100644 --- a/aigw/integrations/libraries/langchain-js.mdx +++ b/aigw/integrations/libraries/langchain-js.mdx @@ -1,6 +1,6 @@ --- title: "Langchain (JS/TS)" -description: "Add Prisma AIRS AI Gateway's enterprise features to any Langchain app—observability, reliability, caching, and cost control." +description: "Add Prisma AIRS AI Gateway's observability, reliability, caching, and cost control to any Langchain app." --- import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; @@ -47,7 +47,7 @@ That's it! You now get: Langchain handles application orchestration. The AI Gateway adds production features: - + Every request logged with costs, latency, tokens. Team-level analytics and debugging. diff --git a/aigw/integrations/libraries/langchain-python.mdx b/aigw/integrations/libraries/langchain-python.mdx index 024b2d91..91dd5889 100644 --- a/aigw/integrations/libraries/langchain-python.mdx +++ b/aigw/integrations/libraries/langchain-python.mdx @@ -1,6 +1,6 @@ --- title: "Langchain (Python)" -description: "Add Prisma AIRS AI Gateway's enterprise features to any Langchain app—observability, reliability, caching, and cost control." +description: "Add Prisma AIRS AI Gateway's observability, reliability, caching, and cost control to any Langchain app." --- import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; @@ -45,7 +45,7 @@ That's it! You now get: Langchain handles application orchestration. The AI Gateway adds production features: - + Every request logged with costs, latency, tokens. Team-level analytics and debugging. @@ -420,11 +420,6 @@ vectors = embeddings.embed_documents(["Hello world", "Goodbye world"]) The AI Gateway supports OpenAI embeddings via `OpenAIEmbeddings`. For other providers (Cohere, Voyage), call the [embeddings endpoint](/aigw/api-reference/embeddings/create-embedding) directly. -## Prompt Management - -Use prompts from the AI Gateway's Prompt Library: - - ## Migration from Direct OpenAI Already using Langchain with OpenAI? Just update 3 parameters: diff --git a/aigw/integrations/libraries/langflow.mdx b/aigw/integrations/libraries/langflow.mdx index 3edbdfa4..d64c0131 100644 --- a/aigw/integrations/libraries/langflow.mdx +++ b/aigw/integrations/libraries/langflow.mdx @@ -1,9 +1,9 @@ --- title: 'Langflow' -description: 'Add enterprise-grade features to your Langflow AI workflows with Prisma AIRS AI Gateway' +description: 'Add observability, governance, and guardrails to your Langflow AI workflows with Prisma AIRS AI Gateway' --- -Langflow is an open-source visual framework for building multi-agent and RAG applications. The AI Gateway adds enterprise controls: +Langflow is an open-source visual framework for building multi-agent and RAG applications. The AI Gateway adds: - **3,000+ LLMs** — Single interface for all providers, not just OpenAI - **Observability** — Real-time tracking for 40+ metrics and logs @@ -11,7 +11,7 @@ Langflow is an open-source visual framework for building multi-agent and RAG app - **Guardrails** — PII detection, content filtering, compliance controls -For enterprise governance setup, see [Enterprise Governance](#enterprise-governance). +For governance setup, see [Set Up Governance](#3-set-up-governance). ## Quick Start diff --git a/aigw/integrations/libraries/librechat.mdx b/aigw/integrations/libraries/librechat.mdx index e1555c72..aabd9a47 100644 --- a/aigw/integrations/libraries/librechat.mdx +++ b/aigw/integrations/libraries/librechat.mdx @@ -10,11 +10,11 @@ import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; Add Prisma AIRS AI Gateway to LibreChat to get: - **Unified access to 3,000+ LLMs** through a single API - **Real-time observability** with 40+ metrics and detailed logs -- **Enterprise governance** with budget limits and RBAC +- **Governance** with budget limits and RBAC - **Security guardrails** for PII detection and content filtering -For enterprise governance setup, see [Enterprise Governance](#3-enterprise-governance). +For governance setup, see [Set Up Governance](#3-set-up-governance). ## 1. Setup the AI Gateway diff --git a/aigw/integrations/libraries/llama-index-python.mdx b/aigw/integrations/libraries/llama-index-python.mdx index cb12366e..7b5b6fe8 100644 --- a/aigw/integrations/libraries/llama-index-python.mdx +++ b/aigw/integrations/libraries/llama-index-python.mdx @@ -1,6 +1,6 @@ --- title: "LlamaIndex (Python)" -description: "Add Prisma AIRS AI Gateway's enterprise features to any LlamaIndex app—observability, reliability, caching, and cost control." +description: "Add Prisma AIRS AI Gateway's observability, reliability, caching, and cost control to any LlamaIndex app." --- import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; @@ -41,7 +41,7 @@ That's it! You now get: LlamaIndex handles data indexing and querying. The AI Gateway adds production features: - + Every request logged with costs, latency, tokens. Team-level analytics and debugging. @@ -339,11 +339,6 @@ Filter and analyze logs by metadata in Strata Cloud Manager. Track costs, performance, and debug issues -## Prompt Management - -Use prompts from the AI Gateway's Prompt Library: - - ## Migration from Direct OpenAI Already using LlamaIndex with OpenAI? Just update 3 parameters: diff --git a/aigw/integrations/libraries/mongodb.mdx b/aigw/integrations/libraries/mongodb.mdx index 9db176f3..1c038f33 100644 --- a/aigw/integrations/libraries/mongodb.mdx +++ b/aigw/integrations/libraries/mongodb.mdx @@ -1,13 +1,13 @@ --- title: "MongoDB" -description: "Store Prisma AIRS AI Gateway logs in MongoDB for enterprise deployments" +description: "Store Prisma AIRS AI Gateway logs in MongoDB for hybrid deployments" --- -Available for AI Gateway Enterprise users. +Available on hybrid deployments. -AI Gateway Enterprise lets you store all LLM logs in MongoDB for scalable, high-performance logging in production AI apps. +A hybrid AI Gateway deployment can store all LLM logs in MongoDB for scalable, high-performance logging in production AI apps. The AI Gateway is part of the [MongoDB partner ecosystem](https://cloud.mongodb.com/ecosystem/portkey-ai). @@ -15,7 +15,7 @@ The AI Gateway is part of the [MongoDB partner ecosystem](https://cloud.mongodb. ## Prerequisites -- AI Gateway Enterprise account +- Hybrid AI Gateway deployment - MongoDB instance - Kubernetes cluster @@ -80,9 +80,9 @@ mongodb://:@?tls=true&tlsCAFile=/etc/shared/document_db.pe Deploy the AI Gateway with MongoDB on: - - - + + + [Explore the AI Gateway’s features →](/aigw/introduction/feature-overview) diff --git a/aigw/integrations/libraries/n8n.mdx b/aigw/integrations/libraries/n8n.mdx index 7ee45b8a..63b4f99e 100644 --- a/aigw/integrations/libraries/n8n.mdx +++ b/aigw/integrations/libraries/n8n.mdx @@ -3,7 +3,7 @@ title: 'n8n' description: 'Add observability, cost controls, and security guardrails to your n8n workflows' --- -n8n is a workflow automation platform. The Prisma AIRS AI Gateway adds enterprise controls for production deployments: +n8n is a workflow automation platform. The Prisma AIRS AI Gateway adds controls for production deployments: - **3,000+ LLMs** — Single interface for all providers, not just OpenAI & Anthropic - **Observability** — Real-time tracking for 40+ metrics and logs @@ -11,7 +11,7 @@ n8n is a workflow automation platform. The Prisma AIRS AI Gateway adds enterpris - **Guardrails** — PII detection, content filtering, compliance controls -For enterprise governance setup, see [Enterprise Governance](#enterprise-governance). +For governance setup, see [Set Up Governance](#3-set-up-governance). ## Quick Start diff --git a/aigw/integrations/libraries/openai-agent-builder-python.mdx b/aigw/integrations/libraries/openai-agent-builder-python.mdx index 6d88fbae..c7fb1793 100644 --- a/aigw/integrations/libraries/openai-agent-builder-python.mdx +++ b/aigw/integrations/libraries/openai-agent-builder-python.mdx @@ -180,11 +180,6 @@ default_headers={ } ``` -### Prompt Templates - -Use the AI Gateway's prompt management for versioned prompts: - - ## Switching Providers Use any of 3,000+ models: @@ -200,7 +195,7 @@ model="@google-prod/gemini-2.0-flash" See all 3,000+ supported models -## Enterprise Governance +## Governance Set up centralized control for your workflows. diff --git a/aigw/integrations/libraries/openai-agent-builder.mdx b/aigw/integrations/libraries/openai-agent-builder.mdx index 31279616..f1d82181 100644 --- a/aigw/integrations/libraries/openai-agent-builder.mdx +++ b/aigw/integrations/libraries/openai-agent-builder.mdx @@ -176,11 +176,6 @@ defaultHeaders: { } ``` -### Prompt Templates - -Use the AI Gateway's prompt management for versioned prompts: - - ## Switching Providers Use any of 3,000+ models: @@ -196,7 +191,7 @@ model: '@google-prod/gemini-2.0-flash' See all 3,000+ supported models -## Enterprise Governance +## Governance Set up centralized control for your workflows. diff --git a/aigw/integrations/libraries/openai-compatible.mdx b/aigw/integrations/libraries/openai-compatible.mdx index fc56bc59..58191b09 100644 --- a/aigw/integrations/libraries/openai-compatible.mdx +++ b/aigw/integrations/libraries/openai-compatible.mdx @@ -7,7 +7,7 @@ import LegacyScreenshotNote from "/snippets/aigw/legacy-screenshot-note.mdx"; -The AI Gateway works with any tool or application that supports OpenAI-compatible APIs. Add enterprise features—observability, reliability, cost controls—with just 2 configuration changes. +The AI Gateway works with any tool or application that supports OpenAI-compatible APIs. Add observability, reliability, and cost controls with just 2 configuration changes. ## Quick Start @@ -31,7 +31,7 @@ You now get: ## Why Add the AI Gateway? - + Every request logged with costs, latency, tokens. Track usage across teams. diff --git a/aigw/integrations/libraries/openclaw.mdx b/aigw/integrations/libraries/openclaw.mdx index dc60e13a..0e11d9a0 100644 --- a/aigw/integrations/libraries/openclaw.mdx +++ b/aigw/integrations/libraries/openclaw.mdx @@ -365,13 +365,12 @@ When deploying to a team, attach configs to API keys so developers get reliabili Developers use a simple config — all routing and reliability logic is handled by the attached config. When you update the config, changes apply immediately. -### Enterprise Options +### Deployment and Access Options - **SaaS**: Everything on AI Gateway cloud - **Hybrid**: Gateway on your infra, management plane on the AI Gateway - - **Air-gapped** *(legacy — existing deployments only)*: Everything on your infra In hybrid mode, the gateway has no runtime dependency on the management plane — routing continues even if the connection drops. diff --git a/aigw/integrations/libraries/openwebui.mdx b/aigw/integrations/libraries/openwebui.mdx index 6c181e17..e4161832 100644 --- a/aigw/integrations/libraries/openwebui.mdx +++ b/aigw/integrations/libraries/openwebui.mdx @@ -1,12 +1,12 @@ --- title: "Open WebUI" -description: "Enterprise-grade cost tracking, observability, and more for Open WebUI" +description: "Cost tracking, observability, and more for Open WebUI" --- Add Prisma AIRS AI Gateway to Open WebUI to get: - **Unified access to 3,000+ LLMs** through a single API - **Real-time cost tracking** and per-user attribution -- **Enterprise governance** with budget limits and access controls +- **Governance** with budget limits and access controls - **Reliability features** like fallbacks, caching, and retries @@ -20,7 +20,7 @@ Add Prisma AIRS AI Gateway to Open WebUI to get: | Path | Best For | |------|----------| | **Direct OpenAI-compatible connection** | Quick setup, using Model Catalog models in Open WebUI | -| **AI Gateway Manifold Pipe** | Enterprise deployments needing per-user attribution with shared API keys | +| **AI Gateway Manifold Pipe** | Team deployments needing per-user attribution with shared API keys | Individual users: complete the workspace setup and one integration option below. @@ -77,9 +77,9 @@ Individual users: complete the workspace setup and one integration option below. Monitor requests and costs in the [Strata Cloud Manager](https://stratacloudmanager.paloaltonetworks.com/). -### Option B: AI Gateway Manifold Pipe (Enterprise) +### Option B: AI Gateway Manifold Pipe -The Manifold Pipe solves a critical enterprise problem: **per-user attribution with shared API keys**. +The Manifold Pipe solves a critical problem for team deployments: **per-user attribution with shared API keys**. In typical deployments, a shared API key means all requests appear anonymous in logs. The Manifold Pipe automatically forwards Open WebUI user context (email, name, role) to the AI Gateway, enabling true per-user cost tracking and governance. @@ -297,7 +297,7 @@ Open WebUI User B → Pipe adds metadata → Portkey The pipe uses Open WebUI's `__user__` context object and formats it as [AI Gateway metadata](/aigw/product/observability/metadata). -## 3. Enterprise Governance +## 3. Governance @@ -402,7 +402,7 @@ For other providers (Gemini, Vertex AI), add parameters via `override_params` in -### Enterprise +### Access Management @@ -436,7 +436,3 @@ For other providers (Gemini, Vertex AI), add parameters via `override_params` in ## Next Steps - [GitHub Repository](https://github.com/Portkey-AI) - - -For enterprise support, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/integrations/libraries/roo-code.mdx b/aigw/integrations/libraries/roo-code.mdx index 7798f360..5470b1fc 100644 --- a/aigw/integrations/libraries/roo-code.mdx +++ b/aigw/integrations/libraries/roo-code.mdx @@ -1,6 +1,6 @@ --- title: 'Roo Code' -description: 'Add enterprise-grade observability, cost tracking, and governance to your Roo AI coding assistant' +description: 'Add observability, cost tracking, and governance to your Roo AI coding assistant' --- Roo is an AI coding assistant for VS Code. Add Prisma AIRS AI Gateway to get: @@ -13,7 +13,7 @@ Roo is an AI coding assistant for VS Code. Add Prisma AIRS AI Gateway to get: This guide shows how to configure Roo with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). # 1. Setup @@ -80,7 +80,7 @@ Change models by updating the **Model** field: All requests route through the AI Gateway automatically. -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model to `dummy`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set Model to `dummy`. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/libraries/vercel.mdx b/aigw/integrations/libraries/vercel.mdx index 69eaa2fd..bd606011 100644 --- a/aigw/integrations/libraries/vercel.mdx +++ b/aigw/integrations/libraries/vercel.mdx @@ -3,14 +3,13 @@ title: 'Vercel AI SDK' description: 'Use Prisma AIRS AI Gateway with Vercel AI SDK to build production-ready AI apps with full observability, reliability, and 3,000+ model support' --- -The AI Gateway seamlessly integrates with the Vercel AI SDK, enabling you to build production-ready AI applications with enterprise-grade reliability, observability, and governance. Simply point the OpenAI provider to the AI Gateway and unlock powerful features: +The AI Gateway seamlessly integrates with the Vercel AI SDK, enabling you to build production-ready AI applications with reliability, observability, and governance. Simply point the OpenAI provider to the AI Gateway and unlock powerful features: * **Full-stack observability** - Complete tracing and analytics for every request * **3,000+ LLMs** - Switch between OpenAI, Anthropic, Google, AWS Bedrock, and 3,000+ models -* **Enterprise reliability** - Fallbacks, load balancing, automatic retries, and circuit breakers +* **Reliability** - Fallbacks, load balancing, automatic retries, and circuit breakers * **Smart caching** - Reduce costs up to 80% with semantic and simple caching * **Production guardrails** - 50+ built-in checks for safety and quality -* **Prompt management** - Version, test, and deploy prompts from the AI Gateway's studio **Migrated from `@portkey-ai/vercel-provider`?** @@ -492,7 +491,7 @@ const openai = createOpenAI({ }); ``` -## Enterprise Features +## Production Features ### Observability & Analytics @@ -514,7 +513,7 @@ Every request through the AI Gateway is automatically logged with: ### AI Gateway Features -AI Gateway makes your AI applications production-ready with enterprise reliability: +AI Gateway makes your AI applications production-ready with built-in reliability: @@ -567,10 +566,6 @@ The AI Gateway offers 50+ built-in guardrails including: Set up real-time safety and compliance checks -### Prompt Management - -Manage, version, and deploy prompts from the AI Gateway's Prompt Studio: - ## Migration Guide ### From `@portkey-ai/vercel-provider` diff --git a/aigw/integrations/libraries/zed.mdx b/aigw/integrations/libraries/zed.mdx index dc485707..94e83dab 100644 --- a/aigw/integrations/libraries/zed.mdx +++ b/aigw/integrations/libraries/zed.mdx @@ -1,6 +1,6 @@ --- title: "Zed" -description: "Learn how to integrate Prisma AIRS AI Gateway's enterprise features with Zed for enhanced observability, reliability and governance." +description: "Learn how to integrate Prisma AIRS AI Gateway with Zed for enhanced observability, reliability and governance." --- Zed is a next-generation code editor built for AI-powered development. Add the AI Gateway to get: @@ -13,7 +13,7 @@ Zed is a next-generation code editor built for AI-powered development. Add the A This guide shows how to configure Zed with the AI Gateway in under 5 minutes. - For enterprise deployments across teams, see [Enterprise Governance](#3-enterprise-governance). + For deployments across teams, see [Set Up Governance](#3-set-up-governance). # 1. Setup @@ -85,7 +85,7 @@ Add more models to your `available_models` array: All requests route through the AI Gateway automatically. -**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set model name to `dummy`. See [Enterprise Governance](#3-enterprise-governance) for examples. +**Want fallbacks, load balancing, or caching?** Create a [AI Gateway Config](/aigw/product/ai-gateway/configs), attach it to your API key, and set model name to `dummy`. See [Set Up Governance](#3-set-up-governance) for examples. import AdvancedFeatures from '/snippets/aigw/portkey-advanced-features.mdx'; diff --git a/aigw/integrations/llms.mdx b/aigw/integrations/llms.mdx index 2674fd50..ca148b5b 100644 --- a/aigw/integrations/llms.mdx +++ b/aigw/integrations/llms.mdx @@ -184,5 +184,3 @@ The AI Gateway has native integrations with the following frameworks. Click to r - -Have a suggestion for an integration with the AI Gateway? Drop us a message via [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). diff --git a/aigw/integrations/llms/ai21.mdx b/aigw/integrations/llms/ai21.mdx index 5fd8fdf9..0e3c9dc0 100644 --- a/aigw/integrations/llms/ai21.mdx +++ b/aigw/integrations/llms/ai21.mdx @@ -5,7 +5,7 @@ description: "Integrate AI21 models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [AI21's models](https://ai21.com). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -83,12 +83,6 @@ console.log(response.choices[0].message.content) See all setup options, code examples, and detailed instructions -## Managing AI21 Prompts - -Manage all prompt templates to AI21 in the Prompt Library. All current AI21 models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Next Steps diff --git a/aigw/integrations/llms/anthropic.mdx b/aigw/integrations/llms/anthropic.mdx index 1fda93b9..985a72d5 100644 --- a/aigw/integrations/llms/anthropic.mdx +++ b/aigw/integrations/llms/anthropic.mdx @@ -5,7 +5,7 @@ description: "Integrate Anthropic's Claude models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Anthropic's Claude APIs](https://docs.anthropic.com/claude/reference/getting-started-with-the-api). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). @@ -466,12 +466,6 @@ x-portkey-anthropic-beta: "token-efficient-tools-2025-02-19" ``` -## Managing Anthropic Prompts - -Manage all prompt templates to Anthropic in the Prompt Library. All current Anthropic models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Next Steps diff --git a/aigw/integrations/llms/anthropic/batches.mdx b/aigw/integrations/llms/anthropic/batches.mdx index c84dc1e7..8e0227bd 100644 --- a/aigw/integrations/llms/anthropic/batches.mdx +++ b/aigw/integrations/llms/anthropic/batches.mdx @@ -4,7 +4,7 @@ description: Run asynchronous batch jobs on Anthropic's Claude models through Pr --- -Self-hosted deployments require **Enterprise Gateway v2.6.1+**. +Hybrid deployments require **Enterprise Gateway v2.6.1+**. The AI Gateway exposes Anthropic's [Message Batches API](https://docs.anthropic.com/en/docs/build-with-claude/batch-processing) through the unified OpenAI-compatible `/v1/batches` endpoints. All five batch operations — create, retrieve, cancel, list, and get output — are supported, so you can run large, asynchronous Claude jobs at 50% lower cost without leaving the AI Gateway. diff --git a/aigw/integrations/llms/anthropic/prompt-caching.mdx b/aigw/integrations/llms/anthropic/prompt-caching.mdx index e792a1b2..4261e46e 100644 --- a/aigw/integrations/llms/anthropic/prompt-caching.mdx +++ b/aigw/integrations/llms/anthropic/prompt-caching.mdx @@ -4,7 +4,7 @@ title: 'Prompt Caching' Prompt caching on Anthropic lets you cache individual messages in your request for repeat use. With caching, you can free up your tokens to include more context in your prompt, and also deliver responses significantly faster and cheaper. -You can use this feature on our OpenAI-compliant universal API as well as with our prompt templates. +You can use this feature on our OpenAI-compliant universal API. ## API Support @@ -64,10 +64,6 @@ console.log(chatCompletion.choices[0].message.content); ``` -## Prompt Templates Support - -Set any message in your prompt template to be cached by just toggling the `Cache Control` setting in the UI: - ## Cache TTL Options By default, the cache has a **5-minute** lifetime that refreshes each time cached content is used. You can optionally specify a **1-hour** TTL by adding the `ttl` field to `cache_control`: diff --git a/aigw/integrations/llms/azure-foundry.mdx b/aigw/integrations/llms/azure-foundry.mdx index 7517627e..9ffecfa0 100644 --- a/aigw/integrations/llms/azure-foundry.mdx +++ b/aigw/integrations/llms/azure-foundry.mdx @@ -3,7 +3,7 @@ title: "Azure AI Foundry" description: "Learn how to integrate Azure AI Foundry with Prisma AIRS AI Gateway to access a wide range of AI models with enhanced observability and reliability features." --- -Azure AI Foundry provides a unified platform for enterprise AI operations, model building, and application development. With the AI Gateway, you can seamlessly integrate with various models available on Azure AI Foundry and take advantage of features like observability, prompt management, fallbacks, and more. +Azure AI Foundry provides a unified platform for enterprise AI operations, model building, and application development. With the AI Gateway, you can seamlessly integrate with various models available on Azure AI Foundry and take advantage of features like observability, fallbacks, and more. ## Quick Start @@ -119,7 +119,7 @@ Required parameters: To use this authentication your azure application need to have the role of: `conginitive services user`. - Enterprise-level authentication with Azure Entra ID: + Authentication with Azure Entra ID: Required parameters: @@ -541,13 +541,6 @@ Route requests based on specific conditions like user type or content requiremen } ``` ---- - -## Managing Prompts - -You can manage all prompts to Azure AI Foundry in the Prompt Library. Once you've created and tested a prompt in the library, call `POST /v1/prompts/{promptId}/completions` to use it in your application. - - --- ## Rerank diff --git a/aigw/integrations/llms/azure-openai/authentication.mdx b/aigw/integrations/llms/azure-openai/authentication.mdx index ed30d9e5..5811f410 100644 --- a/aigw/integrations/llms/azure-openai/authentication.mdx +++ b/aigw/integrations/llms/azure-openai/authentication.mdx @@ -14,7 +14,7 @@ The AI Gateway supports four authentication modes for Azure OpenAI and Azure AI | AWS federated (Entra ID) | `entraFederated` | Gateway runs on AWS (e.g. EKS with IRSA) and needs keyless access to Azure. Requires **Enterprise Gateway v2.6.2+**. | -`managed`, `workload`, and `entraFederated` are only available on self-hosted Enterprise deployments. `entraFederated` requires Node.js runtime. +`managed`, `workload`, and `entraFederated` are only available on hybrid deployments. `entraFederated` requires Node.js runtime. ## API key diff --git a/aigw/integrations/llms/azure-openai/azure-openai.mdx b/aigw/integrations/llms/azure-openai/azure-openai.mdx index 351cadb1..f47bcf89 100644 --- a/aigw/integrations/llms/azure-openai/azure-openai.mdx +++ b/aigw/integrations/llms/azure-openai/azure-openai.mdx @@ -3,7 +3,7 @@ title: "Azure OpenAI" description: "Azure OpenAI is a great alternative to accessing the best models including GPT-4 and more in your private environments. The Prisma AIRS AI Gateway provides complete support for Azure OpenAI." --- -With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, prompt management, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). ## Using Azure OpenAI with the AI Gateway @@ -115,12 +115,6 @@ Set up the AI Gateway with your Azure Integration as part of the initialization Use the AI Gateway instance to send requests to your Azure deployments. You can also override the provider slug directly in the API call if needed. -## Managing Azure OpenAI Prompts - -You can manage all prompts to Azure OpenAI in the Prompt Library. All the current models of OpenAI are supported and you can easily start testing different prompts. - -Once you're ready with your prompt, call the `POST /v1/prompts/{promptId}/completions` endpoint to use it in your application. - ## Image Generation The AI Gateway supports multiple modalities for Azure OpenAI and you can make image generation requests through AI Gateway the same way as making completion calls. diff --git a/aigw/integrations/llms/bedrock/aws-bedrock.mdx b/aigw/integrations/llms/bedrock/aws-bedrock.mdx index 850c9e65..060de0c7 100644 --- a/aigw/integrations/llms/bedrock/aws-bedrock.mdx +++ b/aigw/integrations/llms/bedrock/aws-bedrock.mdx @@ -4,7 +4,7 @@ title: "AWS Bedrock" The Prisma AIRS AI Gateway provides a robust and secure gateway to facilitate the integration of various Large Language Models (LLMs) into your applications, including models hosted on AWS Bedrock. -With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, prompt management, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). ## Using AWS Bedrock with the AI Gateway @@ -493,12 +493,6 @@ If you require the model to [respond with certain fields](https://docs.aws.amazo "additionalModelResponseFieldPaths": [ "/stop_sequence" ] ``` -## Managing AWS Bedrock Prompts - -You can manage all prompts to AWS bedrock in the Prompt Library. All the current models of Anthropic are supported and you can easily start testing different prompts. - -Once you're ready with your prompt, call the `POST /v1/prompts/{promptId}/completions` endpoint to use it in your application. - ## Making Requests without using the AI Gateway's Model Catalog If you do not want to add your AWS details to the AI Gateway vault, you can also directly pass them while instantiating the AI Gateway client. @@ -536,9 +530,9 @@ curl https://aigw.portkey.ai/v1/chat/completions \ --- -## Using AWS PrivateLink for Bedrock [Self Hosted Enterprise] +## Using AWS PrivateLink for Bedrock [Hybrid Deployments] -Though using assumed role is in itself enough for enterprise security. You can additional configure AWS [PrivateLink](https://docs.aws.amazon.com/vpc/latest/privatelink/what-is-privatelink.html) for Bedrock to ensure that your requests are not traversed outside your VPC. +Using an assumed role is secure on its own. You can additionally configure AWS [PrivateLink](https://docs.aws.amazon.com/vpc/latest/privatelink/what-is-privatelink.html) for Bedrock to ensure that your requests are not traversed outside your VPC. - Create a private link between the VPC you've deployed the AI Gateway and AWS Bedrock (the endpoint is in most cases `https://bedrock.{your_region}.amazonaws.com`). - When configuring your integration, simply configure the `custom host` option to point to your VPC endpoint for the private link. diff --git a/aigw/integrations/llms/bedrock/bedrock-knowledgebase.mdx b/aigw/integrations/llms/bedrock/bedrock-knowledgebase.mdx index db16f8ae..fd01ed8c 100644 --- a/aigw/integrations/llms/bedrock/bedrock-knowledgebase.mdx +++ b/aigw/integrations/llms/bedrock/bedrock-knowledgebase.mdx @@ -5,7 +5,7 @@ description: 'Create, manage, and connect your LLMs to organisational data using AWS Bedrock Knowledge Bases enables you to give foundation models access to your company's private data sources, delivering more relevant, accurate, and customized responses through Retrieval Augmented Generation (RAG). -With the AI Gateway's integration, you can seamlessly create, connect to, and query AWS Bedrock Knowledge Bases while gaining enterprise features like observability, caching, and reliability - all through a unified API that simplifies authentication. +With the AI Gateway's integration, you can seamlessly create, connect to, and query AWS Bedrock Knowledge Bases while gaining features like observability, caching, and reliability - all through a unified API that simplifies authentication. ## What is AWS Bedrock Knowledge Bases? @@ -92,9 +92,4 @@ Notice that we use two different AI Gateway client initializations. The first on Now that you can create and build a full RAG application with your knowledge base, you can explore other AI Gateway features to enhance your application: - Explore [AI Gateway Observability](/aigw/product/observability) to monitor your RAG pipeline. - Set up [Guardrails](/aigw/product/guardrails) to ensure safe and compliant responses. -- Implement Prompt Management for your RAG prompts. - Configure [Fallbacks](/aigw/product/ai-gateway/fallbacks) for high availability. - - -For enterprise support with your AWS Bedrock integration, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/integrations/llms/bedrock/prompt-caching.mdx b/aigw/integrations/llms/bedrock/prompt-caching.mdx index 600ea0f3..be495207 100644 --- a/aigw/integrations/llms/bedrock/prompt-caching.mdx +++ b/aigw/integrations/llms/bedrock/prompt-caching.mdx @@ -4,7 +4,7 @@ title: 'Prompt Caching on Bedrock' Prompt caching on Amazon Bedrock lets you cache specific portions of your requests for repeated use. This feature significantly reduces inference response latency and input token costs by allowing the model to skip recomputation of previously processed content. -With Prisma AIRS AI Gateway, you can easily implement Amazon Bedrock's prompt caching through our OpenAI-compliant unified API and prompt templates. +With Prisma AIRS AI Gateway, you can easily implement Amazon Bedrock's prompt caching through our OpenAI-compliant unified API. ## Model Support @@ -28,10 +28,6 @@ Amazon Bedrock prompt caching is generally available with the following models: When using prompt caching, you define **cache checkpoints** - markers that indicate parts of your prompt to cache. These cached sections must be static between requests; any alterations will result in a cache miss. - - You can also use Bedrock Prompt Caching Feature with the AI Gateway's Prompt Templates. - - ## Implementation Examples Here's how to implement prompt caching with the AI Gateway: diff --git a/aigw/integrations/llms/byteplus.mdx b/aigw/integrations/llms/byteplus.mdx index ad51cda6..fc4c74c9 100644 --- a/aigw/integrations/llms/byteplus.mdx +++ b/aigw/integrations/llms/byteplus.mdx @@ -5,7 +5,7 @@ description: "Integrate BytePlus ModelArk models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [BytePlus ModelArk](https://www.byteplus.com/en/product/modelark) models. -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). Provider Slug: `byteplus` diff --git a/aigw/integrations/llms/cerebras.mdx b/aigw/integrations/llms/cerebras.mdx index 7bd0b4e4..710539e5 100644 --- a/aigw/integrations/llms/cerebras.mdx +++ b/aigw/integrations/llms/cerebras.mdx @@ -5,7 +5,7 @@ description: "Integrate Cerebras models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Cerebras Inference API](https://cerebras.ai/inference). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start diff --git a/aigw/integrations/llms/claude-platform-aws/batches.mdx b/aigw/integrations/llms/claude-platform-aws/batches.mdx index 6fbee829..ca7e9d6f 100644 --- a/aigw/integrations/llms/claude-platform-aws/batches.mdx +++ b/aigw/integrations/llms/claude-platform-aws/batches.mdx @@ -4,7 +4,7 @@ description: Run asynchronous batch jobs on Claude Platform on AWS through Prism --- -Self-hosted deployments require **Enterprise Gateway v2.6.1+**. +Hybrid deployments require **Enterprise Gateway v2.6.1+**. The AI Gateway exposes Anthropic's [Message Batches API](https://docs.anthropic.com/en/docs/build-with-claude/batch-processing) for Claude Platform on AWS through the unified OpenAI-compatible `/v1/batches` endpoints. All five batch operations -- create, retrieve, cancel, list, and get output -- are supported, enabling large asynchronous Claude jobs at 50% lower cost. diff --git a/aigw/integrations/llms/claude-platform-aws/setup-assumed-role.mdx b/aigw/integrations/llms/claude-platform-aws/setup-assumed-role.mdx index da946731..4549a4d2 100644 --- a/aigw/integrations/llms/claude-platform-aws/setup-assumed-role.mdx +++ b/aigw/integrations/llms/claude-platform-aws/setup-assumed-role.mdx @@ -115,7 +115,7 @@ arn:aws:iam::039293892788:role/AirsGwEnterpriseRole ``` -This ARN is for the [hosted Strata Cloud Manager](https://stratacloudmanager.paloaltonetworks.com/). For enterprise self-hosted deployments, refer to the [Helm documentation](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/docs/Bedrock.md). +This ARN is for the [hosted Strata Cloud Manager](https://stratacloudmanager.paloaltonetworks.com/). For hybrid deployments, refer to the [Helm documentation](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/docs/Bedrock.md). ### Trust policy without external ID diff --git a/aigw/integrations/llms/cohere.mdx b/aigw/integrations/llms/cohere.mdx index 2d43549a..d3b966ac 100644 --- a/aigw/integrations/llms/cohere.mdx +++ b/aigw/integrations/llms/cohere.mdx @@ -5,7 +5,7 @@ description: "Integrate Cohere models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including Cohere's generation, embedding, and reranking endpoints. -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -95,12 +95,6 @@ Embedding endpoints are natively supported within the AI Gateway: Use Cohere reranking by calling `POST /v1/rerank` through the gateway with the body expected by [Cohere's reranking API](https://docs.cohere.com/reference/rerank-1): -## Managing Cohere Prompts - -Manage all prompt templates to Cohere in the Prompt Library. All current Cohere models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Next Steps diff --git a/aigw/integrations/llms/databricks.mdx b/aigw/integrations/llms/databricks.mdx index b5114dc0..1c0bc9e4 100644 --- a/aigw/integrations/llms/databricks.mdx +++ b/aigw/integrations/llms/databricks.mdx @@ -5,7 +5,7 @@ description: "Integrate Databricks Model Serving with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Databricks Model Serving](https://docs.databricks.com/en/machine-learning/model-serving/index.html) endpoints. -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). Provider Slug: `databricks` diff --git a/aigw/integrations/llms/deepinfra.mdx b/aigw/integrations/llms/deepinfra.mdx index 57431bc5..76d85d69 100644 --- a/aigw/integrations/llms/deepinfra.mdx +++ b/aigw/integrations/llms/deepinfra.mdx @@ -5,7 +5,7 @@ description: "Integrate Deepinfra models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Deepinfra's hosted models](https://deepinfra.com/models/text-generation). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start diff --git a/aigw/integrations/llms/deepseek.mdx b/aigw/integrations/llms/deepseek.mdx index b9fa1cee..45986a0c 100644 --- a/aigw/integrations/llms/deepseek.mdx +++ b/aigw/integrations/llms/deepseek.mdx @@ -5,7 +5,7 @@ description: "Integrate DeepSeek models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including DeepSeek's models. -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -107,12 +107,6 @@ The `deepseek-reasoner` model supports chain-of-thought reasoning. Use the `reas When streaming, `reasoning_content` is included in the delta for `deepseek-reasoner` responses. -## Managing DeepSeek Prompts - -Manage all prompt templates to DeepSeek in the Prompt Library. All current DeepSeek models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Supported Endpoints - Chat Completions diff --git a/aigw/integrations/llms/gemini.mdx b/aigw/integrations/llms/gemini.mdx index 01e2b006..02bb7e36 100644 --- a/aigw/integrations/llms/gemini.mdx +++ b/aigw/integrations/llms/gemini.mdx @@ -4,7 +4,7 @@ title: "Google Gemini" The Prisma AIRS AI Gateway provides a robust and secure gateway to facilitate the integration of various Large Language Models (LLMs) into your applications, including [Google Gemini APIs](https://cloud.google.com/vertex-ai/docs/generative-ai/model-reference/gemini). -With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, prompt management, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -1270,7 +1270,7 @@ Note that you will have to set [`strict_open_ai_compliance=False`](/aigw/product -Gemini grounding mode may not work through the AI Gateway. [Contact support](https://support.portkey.ai/forms/customer-portal-ticket-form) for assistance. +Gemini grounding mode may not work through the AI Gateway. --- diff --git a/aigw/integrations/llms/mistral-ai.mdx b/aigw/integrations/llms/mistral-ai.mdx index b5469cbd..30c66fa5 100644 --- a/aigw/integrations/llms/mistral-ai.mdx +++ b/aigw/integrations/llms/mistral-ai.mdx @@ -5,7 +5,7 @@ description: "Integrate Mistral AI models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Mistral AI's models](https://docs.mistral.ai/api/). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -164,12 +164,6 @@ The native `/v1/ocr` endpoint requires gateway version **2.18.0** or higher for --- -## Managing Mistral AI Prompts - -Manage all prompt templates to Mistral AI in the Prompt Library. All current Mistral AI models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Next Steps diff --git a/aigw/integrations/llms/ollama.mdx b/aigw/integrations/llms/ollama.mdx index a19cd0a0..80f82d66 100644 --- a/aigw/integrations/llms/ollama.mdx +++ b/aigw/integrations/llms/ollama.mdx @@ -5,32 +5,11 @@ description: "Integrate Ollama-hosted models with Prisma AIRS AI Gateway for loc The AI Gateway provides a robust gateway to facilitate the integration of your **locally hosted models through Ollama**. - - - -First, install the Gateway locally: - - -```bash npx -npx @portkey-ai/gateway -``` - -```bash Docker -docker run -d -p 8787:8787 portkeyai/gateway:latest -``` - - -Then, connect to your local Ollama instance: - - - - - ## Integration Steps -Expose your Ollama API using a tunneling service like [ngrok](https://ngrok.com/) or make it publicly accessible. Skip this if you're self-hosting the Gateway. +Expose your Ollama API using a tunneling service like [ngrok](https://ngrok.com/) or make it publicly accessible. Skip this if you run a [hybrid deployment](/aigw/self-hosting/private-network-access) that can reach Ollama on your network. For using Ollama with ngrok, here's a [useful guide](https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-use-ollama-with-ngrok): diff --git a/aigw/integrations/llms/openai.mdx b/aigw/integrations/llms/openai.mdx index 4481921d..f54a665e 100644 --- a/aigw/integrations/llms/openai.mdx +++ b/aigw/integrations/llms/openai.mdx @@ -5,7 +5,7 @@ description: "Integrate OpenAI's GPT models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate [OpenAI's APIs](https://platform.openai.com/docs/api-reference/introduction) into your applications, including GPT-4o, o1, DALL·E, Whisper, and more. -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). @@ -252,8 +252,6 @@ console.log(response) Function calls within your OpenAI SDK operations remain standard. These logs will appear in the AI Gateway, highlighting the utilized functions and their outputs. -Additionally, you can define functions within your prompts and call `POST /v1/prompts/{promptId}/completions` as above. - #### Function Calling with the Responses API The Responses API also supports function calling with the same powerful capabilities: diff --git a/aigw/integrations/llms/openrouter.mdx b/aigw/integrations/llms/openrouter.mdx index 4a3d0330..ef5d5f4a 100644 --- a/aigw/integrations/llms/openrouter.mdx +++ b/aigw/integrations/llms/openrouter.mdx @@ -5,7 +5,7 @@ description: "Integrate OpenRouter models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [OpenRouter's unified API](https://openrouter.ai). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). OpenRouter also proxies typed-judgment requests to models like TypeSafe's Jev. Route these through `/v1/decisions` with `x-portkey-provider: @openrouter`. See [TypeSafe (Jev)](/aigw/integrations/llms/typesafe). diff --git a/aigw/integrations/llms/ovhcloud.mdx b/aigw/integrations/llms/ovhcloud.mdx index 090b459c..a35de248 100644 --- a/aigw/integrations/llms/ovhcloud.mdx +++ b/aigw/integrations/llms/ovhcloud.mdx @@ -4,7 +4,7 @@ title: "OVHcloud AI Endpoints" The Prisma AIRS AI Gateway provides a robust and secure gateway to facilitate the integration of various Large Language Models (LLMs) into your applications, including [OVHcloud AI Endpoints](https://www.ovhcloud.com/en/public-cloud/ai-endpoints/). -With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, prompt management, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, and more, all while ensuring the secure management of your LLM API keys through [Model Catalog](/aigw/product/model-catalog). Provider Slug. `ovhcloud` @@ -40,11 +40,7 @@ Navigate to the [OVHcloud control panel](https://ovh.com/manager), in `Public Cl Use the AI Gateway instance to send requests to AI Endpoints. -## Managing OVHcloud AI Endpoints Prompts - -You can manage all prompts to AI Endpoints in the Prompt Library. All the current models of AI Endpoints are supported and you can easily start testing different prompts. - -Once you're ready with your prompt, call the `POST /v1/prompts/{promptId}/completions` endpoint to use it in your application. +## AI Endpoints Capabilities ### AI Endpoints Tool Calling Tool calling feature lets models trigger external tools based on conversation context. You define available functions, the model chooses when to use them, and your application executes them and returns results. diff --git a/aigw/integrations/llms/reka-ai.mdx b/aigw/integrations/llms/reka-ai.mdx index 2f4d61b4..6382ed63 100644 --- a/aigw/integrations/llms/reka-ai.mdx +++ b/aigw/integrations/llms/reka-ai.mdx @@ -5,7 +5,7 @@ description: "Integrate Reka AI models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Reka AI's multimodal models](https://www.reka.ai/). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -83,12 +83,6 @@ console.log(response.choices[0].message.content) See all setup options, code examples, and detailed instructions -## Managing Reka AI Prompts - -Manage all prompt templates to Reka AI in the Prompt Library. All current Reka AI models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Next Steps diff --git a/aigw/integrations/llms/snowflake-cortex.mdx b/aigw/integrations/llms/snowflake-cortex.mdx index 307c0b74..4f7168fb 100644 --- a/aigw/integrations/llms/snowflake-cortex.mdx +++ b/aigw/integrations/llms/snowflake-cortex.mdx @@ -1,6 +1,6 @@ --- title: "Snowflake Cortex" -description: Use Snowflake Cortex AI models through Prisma AIRS AI Gateway for enterprise data cloud AI. +description: Use Snowflake Cortex AI models through Prisma AIRS AI Gateway. --- ## Quick Start diff --git a/aigw/integrations/llms/suggest-a-new-integration.mdx b/aigw/integrations/llms/suggest-a-new-integration.mdx index 56ee1c6c..22cc2832 100644 --- a/aigw/integrations/llms/suggest-a-new-integration.mdx +++ b/aigw/integrations/llms/suggest-a-new-integration.mdx @@ -2,4 +2,4 @@ title: "Suggest a new integration!" --- -Have a suggestion for an integration with Prisma AIRS AI Gateway? Drop us a message via [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). +Have a suggestion for an integration with Prisma AIRS AI Gateway? Contact the Palo Alto Networks team with the provider or tool you would like supported. diff --git a/aigw/integrations/llms/together-ai.mdx b/aigw/integrations/llms/together-ai.mdx index 3ac3bc8e..9555d8ad 100644 --- a/aigw/integrations/llms/together-ai.mdx +++ b/aigw/integrations/llms/together-ai.mdx @@ -5,7 +5,7 @@ description: "Integrate Together AI models with Prisma AIRS AI Gateway" The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including [Together AI's hosted models](https://docs.together.ai/reference/inference). -With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, take advantage of features like fast AI gateway access, observability, and more, while securely managing API keys through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -135,12 +135,6 @@ Video generation pricing is logged in Strata Cloud Manager for cost tracking. --- -## Managing Together AI Prompts - -Manage all prompt templates to Together AI in the Prompt Library. All current Together AI models are supported, and you can easily test different prompts. - -Call the `POST /v1/prompts/{promptId}/completions` endpoint to use the prompt in an application. - ## Next Steps diff --git a/aigw/integrations/llms/vertex-ai.mdx b/aigw/integrations/llms/vertex-ai.mdx index 5a8f9e0d..122d9da8 100644 --- a/aigw/integrations/llms/vertex-ai.mdx +++ b/aigw/integrations/llms/vertex-ai.mdx @@ -4,7 +4,7 @@ title: "Google Vertex AI" The Prisma AIRS AI Gateway provides a robust and secure gateway to facilitate the integration of various Large Language Models (LLMs), and embedding models into your apps, including [Google Vertex AI](https://cloud.google.com/vertex-ai?hl=en). -With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, prompt management, and more, all while ensuring the secure management of your Vertex auth through [Model Catalog](/aigw/product/model-catalog). +With the AI Gateway, you can take advantage of features like fast AI gateway access, observability, and more, all while ensuring the secure management of your Vertex auth through [Model Catalog](/aigw/product/model-catalog). ## Quick Start @@ -841,11 +841,6 @@ The AI Gateway supports function calling mode on Google's Gemini Models. Explore [Function Calling](/guides/getting-started/function-calling) -## Managing Vertex AI Prompts - -You can manage all prompts to Google Gemini in the Prompt Library. All the models in the model garden are supported and you can easily start testing different prompts. - -Once you're ready with your prompt, call the `POST /v1/prompts/{promptId}/completions` endpoint to use it in your application. ## Image Generation Models diff --git a/aigw/integrations/mcp-clients/claude-code.mdx b/aigw/integrations/mcp-clients/claude-code.mdx index 0562467f..4a2a565a 100644 --- a/aigw/integrations/mcp-clients/claude-code.mdx +++ b/aigw/integrations/mcp-clients/claude-code.mdx @@ -134,11 +134,3 @@ To revoke access and clear stored tokens: - If your browser doesn't open automatically during auth, copy the provided URL - OAuth authentication works with both SSE and HTTP transports - Use descriptive names when adding servers for easier management - ---- - -## Support - -Need help? We're here for you: - -- 🎫 Support: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) diff --git a/aigw/integrations/mcp-clients/claude.mdx b/aigw/integrations/mcp-clients/claude.mdx index a56cf024..2d6cde64 100644 --- a/aigw/integrations/mcp-clients/claude.mdx +++ b/aigw/integrations/mcp-clients/claude.mdx @@ -104,11 +104,3 @@ Repeat the connection process for each server you want to add: - Linear: `https://aigw.portkey.ai/m/ws-abc123/linear-def456/mcp` - GitHub: `https://aigw.portkey.ai/m/ws-abc123/github-ghi789/mcp` - Slack: `https://aigw.portkey.ai/m/ws-abc123/slack-jkl012/mcp` - ---- - -## Support - -Need help? We're here for you: - -- 🎫 Support: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) diff --git a/aigw/integrations/mcp-clients/cursor.mdx b/aigw/integrations/mcp-clients/cursor.mdx index 35bbda1d..3fbba3d8 100644 --- a/aigw/integrations/mcp-clients/cursor.mdx +++ b/aigw/integrations/mcp-clients/cursor.mdx @@ -151,8 +151,6 @@ The AI Gateway handles all authentication flows automatically: --- -## Support +## Resources -Need help? We're here for you: -- 🎫 Support: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) - 📚 Documentation: [Prisma AIRS AI Gateway docs](/aigw/introduction/welcome) diff --git a/aigw/integrations/mcp-clients/librechat.mdx b/aigw/integrations/mcp-clients/librechat.mdx index d4997864..4083862b 100644 --- a/aigw/integrations/mcp-clients/librechat.mdx +++ b/aigw/integrations/mcp-clients/librechat.mdx @@ -103,11 +103,3 @@ mcpServers: type: "streamable-http" url: "https://aigw.portkey.ai/m/ws-abc123/slack-jkl012/mcp" ``` - ---- - -## Support - -Need help? We're here for you: - -- 🎫 Support: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) diff --git a/aigw/integrations/mcp-clients/vs-code.mdx b/aigw/integrations/mcp-clients/vs-code.mdx index 3bb17e65..a893236c 100644 --- a/aigw/integrations/mcp-clients/vs-code.mdx +++ b/aigw/integrations/mcp-clients/vs-code.mdx @@ -127,11 +127,3 @@ The AI Gateway handles all authentication flows automatically: } } ``` - ---- - -## Support - -Need help? We're here for you: - -- 🎫 Support: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) diff --git a/aigw/integrations/overview.mdx b/aigw/integrations/overview.mdx index 217604de..0a99b5a2 100644 --- a/aigw/integrations/overview.mdx +++ b/aigw/integrations/overview.mdx @@ -66,14 +66,6 @@ OpenAI Agents SDK enables building powerful AI agents with tools, memory, and mu Open-source framework focused on simplifying internal tool development. - - Cloud computing services and marketplace - - - - Cloud platform and services marketplace - - AI-powered solutions tailored for business efficiency. diff --git a/aigw/integrations/plugins/tavily.mdx b/aigw/integrations/plugins/tavily.mdx index 03810975..5418f1bb 100644 --- a/aigw/integrations/plugins/tavily.mdx +++ b/aigw/integrations/plugins/tavily.mdx @@ -212,7 +212,7 @@ Tavily-enriched requests are visible in Strata Cloud Manager. You can inspect: ## Get Support -If you run into issues with the AI Gateway setup, reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). For Tavily account, billing, or API-specific questions, use the Tavily support paths in your [dashboard](https://app.tavily.com). +For Tavily account, billing, or API-specific questions, use the Tavily support paths in your [dashboard](https://app.tavily.com). ## Learn More diff --git a/aigw/integrations/tracing-providers/arize.mdx b/aigw/integrations/tracing-providers/arize.mdx index 2ad22764..340c3ff5 100644 --- a/aigw/integrations/tracing-providers/arize.mdx +++ b/aigw/integrations/tracing-providers/arize.mdx @@ -213,10 +213,6 @@ Block harmful, toxic, or inappropriate content in real-time based on custom poli Fine-grained RBAC with team management, user permissions, and audit logs. - -Enterprise-grade security with SOC2 Type II certification and GDPR compliance. - - Complete audit trail of all API usage, configuration changes, and user actions. @@ -226,7 +222,7 @@ Zero data retention options and deployment in your own VPC for maximum privacy. -### 🏢 Enterprise Features +### 🏢 Governance Features diff --git a/aigw/integrations/tracing-providers/future-agi.mdx b/aigw/integrations/tracing-providers/future-agi.mdx index 974d50e6..8289b1f6 100644 --- a/aigw/integrations/tracing-providers/future-agi.mdx +++ b/aigw/integrations/tracing-providers/future-agi.mdx @@ -243,5 +243,5 @@ Leverage this integration in your CI/CD pipelines for: 5. Integrate into your CI/CD pipeline for continuous model quality monitoring -For advanced configurations and custom evaluators, check out the [FutureAGI documentation](https://docs.futureagi.com) and reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) for support. +For advanced configurations and custom evaluators, check out the [FutureAGI documentation](https://docs.futureagi.com). diff --git a/aigw/integrations/tracing-providers/honeyhive.mdx b/aigw/integrations/tracing-providers/honeyhive.mdx index 27b0453e..fa0f049e 100644 --- a/aigw/integrations/tracing-providers/honeyhive.mdx +++ b/aigw/integrations/tracing-providers/honeyhive.mdx @@ -289,7 +289,3 @@ def call_openai(): - [HoneyHive Documentation](https://docs.honeyhive.ai) - [AI Gateway Guide](/aigw/product/ai-gateway) - - -For enterprise support and custom features, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/integrations/tracing-providers/langfuse.mdx b/aigw/integrations/tracing-providers/langfuse.mdx index 18520e24..368402d5 100644 --- a/aigw/integrations/tracing-providers/langfuse.mdx +++ b/aigw/integrations/tracing-providers/langfuse.mdx @@ -234,7 +234,3 @@ client = OpenAI( - [Langfuse Documentation](https://langfuse.com/docs) - [AI Gateway Guide](/aigw/product/ai-gateway) - - -For enterprise support and custom features, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/integrations/tracing-providers/langsmith.mdx b/aigw/integrations/tracing-providers/langsmith.mdx index 3b82d8d1..8acc4e77 100644 --- a/aigw/integrations/tracing-providers/langsmith.mdx +++ b/aigw/integrations/tracing-providers/langsmith.mdx @@ -182,7 +182,3 @@ If you're already using LangSmith with OpenAI, migrating to use the AI Gateway i * [LangSmith Documentation](https://docs.smith.langchain.com/) * [AI Gateway Guide](/aigw/product/ai-gateway) - - - For enterprise support and custom features, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - diff --git a/aigw/introduction/feature-overview.mdx b/aigw/introduction/feature-overview.mdx index f24fd4a6..7d7b5d7d 100644 --- a/aigw/introduction/feature-overview.mdx +++ b/aigw/introduction/feature-overview.mdx @@ -120,7 +120,7 @@ Natively integrate the AI Gateway's routing, guardrails, and observability suite ## Governance & Administration -Organise your teams and entities with the access control, encryption, and audit trails that an enterprise AI platform needs. +Organise your teams and entities with access control, encryption, and audit trails. diff --git a/aigw/introduction/welcome.mdx b/aigw/introduction/welcome.mdx index ad7f70f2..0a7ac39d 100644 --- a/aigw/introduction/welcome.mdx +++ b/aigw/introduction/welcome.mdx @@ -2,7 +2,6 @@ title: "Welcome" sidebarTitle: "Welcome" description: The Prisma AIRS AI Gateway is a comprehensive platform designed to streamline and enhance AI integration for developers and organisations. It serves as a unified interface for interacting with over 3,000 AI models, offering advanced tools for control, visibility, and security in your Generative AI apps. -mode: "wide" --- It takes 2 mins to integrate and with that, it starts monitoring all of your LLM requests and makes your app resilient, secure, performant, and more accurate at the same time. @@ -75,13 +74,13 @@ createChatCompletion(); - The AI Gateway is hosted on edge workers throughout the world, ensuring minimal latency. Our benchmarks estimate a total latency addition between 20-40ms compared to direct API calls. This slight increase is often offset by the benefits of our caching and routing optimizations. + Our benchmarks estimate a total latency addition between 20-40ms compared to direct API calls. This slight increase is often offset by the benefits of our caching and routing optimizations. All data is encrypted in transit and at rest using industry-standard AES-256 encryption, and the gateway follows best practices for service security, data storage, and retrieval. - The gateway is built on scalable infrastructure and can handle millions of requests per minute with very high concurrency. Its edge architecture and scaling capabilities accommodate sudden spikes in traffic without performance degradation. + The gateway is built on scalable infrastructure and can handle millions of requests per minute with very high concurrency. Its scaling capabilities accommodate sudden spikes in traffic without performance degradation. No explicit timeout is imposed on requests. While the gateway does not time out requests on our end, we recommend implementing client-side timeouts appropriate for your use case to handle potential network issues or upstream API delays. diff --git a/aigw/product/administration/configure-data-visibility-settings.mdx b/aigw/product/administration/configure-data-visibility-settings.mdx index 14f8630b..a6f5afd0 100644 --- a/aigw/product/administration/configure-data-visibility-settings.mdx +++ b/aigw/product/administration/configure-data-visibility-settings.mdx @@ -4,8 +4,7 @@ title: "Configure Data Visibility Settings for Workspaces" **Prerequisites & Availability** - - **Enterprise (Hybrid Gateway)**: Gateway **v2.2.4** or later is required. Only logs generated from this version onward will be considered for user-level filtering. - - **Airgapped Deployments**: Requires Gateway **v2.2.4**, backend **v1.11.0**, and frontend **v1.6.4** or later. Only logs generated after upgrading to these versions will be considered for user-level filtering. + - **Hybrid Deployments**: Gateway **v2.2.4** or later is required. Only logs generated from this version onward will be considered for user-level filtering. - **SaaS Users**: This setting will be applied to logs starting from **19th February 2026**. diff --git a/aigw/product/administration/configure-guardrail-access-permissions.mdx b/aigw/product/administration/configure-guardrail-access-permissions.mdx index 66c646eb..de43de5f 100644 --- a/aigw/product/administration/configure-guardrail-access-permissions.mdx +++ b/aigw/product/administration/configure-guardrail-access-permissions.mdx @@ -45,10 +45,6 @@ To configure workspace-level overrides: Learn about the AI Gateway's access control features including user roles and organisation hierarchy - - Configure who can view and manage prompts within workspaces - - Learn how to enforce organisation-level request guardrails across all workspaces diff --git a/aigw/product/administration/configure-integration-access-permissions.mdx b/aigw/product/administration/configure-integration-access-permissions.mdx index 33b27a33..6d09d5c9 100644 --- a/aigw/product/administration/configure-integration-access-permissions.mdx +++ b/aigw/product/administration/configure-integration-access-permissions.mdx @@ -45,7 +45,3 @@ To configure workspace-level overrides: Learn about the AI Gateway's access control features including user roles and organisation hierarchy - - - Configure who can view and edit prompts within workspaces - diff --git a/aigw/product/administration/configure-prompt-access-permissions.mdx b/aigw/product/administration/configure-prompt-access-permissions.mdx deleted file mode 100644 index fa707c3f..00000000 --- a/aigw/product/administration/configure-prompt-access-permissions.mdx +++ /dev/null @@ -1,58 +0,0 @@ ---- -title: "Configure Prompt Access Permissions for Workspaces" ---- - -## Overview - -Prompt Management in the AI Gateway enables Organisation administrators to control who can view and edit prompts within workspaces. This feature provides granular permissions to restrict prompt access for workspace managers and members while keeping full access for organisation owners and admins. - -## Accessing Prompt Management - -1. Navigate to **Admin Settings** in Strata Cloud Manager -2. Select the **Security** tab from the left sidebar -3. Locate the **Prompt Management** section - -## Permission Settings - -The Prompt Management section provides four permission options: - -| Permission | Description | -|------------|-------------| -| **Managers View Prompts** | Allow workspace managers to view prompts within their workspace | -| **Managers Write Prompts** | Allow workspace managers to create, update, and delete prompts within their workspace | -| **Members View Prompts** | Allow workspace members to view prompts within their workspace | -| **Members Write Prompts** | Allow workspace members to create, update, and delete prompts within their workspace | - -**Default settings:** -- Managers can view and write prompts -- Members can view prompts but cannot write - - -Organisation **Owners** and **Admins** always have full access to prompts regardless of these settings. - - -## Workspace-Level Overrides - -Prompt access permissions support **workspace-level overrides**, allowing you to configure different prompt access settings per workspace. - -When workspace-level overrides are enabled for prompts: -- Each workspace can have its own `membersViewPrompts`, `membersWritePrompts`, `managersViewPrompts`, and `managersWritePrompts` settings -- Workspace settings take precedence over organisation-level settings - -To configure workspace-level overrides: -1. Enable the **Allow workspace-level configuration** toggle for Prompts in the Security settings -2. Navigate to the specific workspace settings to configure per-workspace prompt access - -## Related Features - - - Learn about the AI Gateway's access control features including user roles and organisation hierarchy - - - - Configure who can view logs and log metadata within workspaces - - - - Configure who can view analytics within workspaces - diff --git a/aigw/product/administration/configuring-request-logging.mdx b/aigw/product/administration/configuring-request-logging.mdx index 3566338f..a28ef7aa 100644 --- a/aigw/product/administration/configuring-request-logging.mdx +++ b/aigw/product/administration/configuring-request-logging.mdx @@ -62,7 +62,7 @@ If workspace-level control is enabled, workspace managers can configure logging 4. Save your changes -Changing from Full Logging to Metrics Only will not retroactively remove previously logged data. [Contact support](https://support.portkey.ai/forms/customer-portal-ticket-form) if you need to purge historical logs. +Changing from Full Logging to Metrics Only will not retroactively remove previously logged data. ## Precedence Order diff --git a/aigw/product/administration/enforce-budget-and-rate-limit.mdx b/aigw/product/administration/enforce-budget-and-rate-limit.mdx index e7157d2f..13172bb9 100644 --- a/aigw/product/administration/enforce-budget-and-rate-limit.mdx +++ b/aigw/product/administration/enforce-budget-and-rate-limit.mdx @@ -12,7 +12,7 @@ Looking to set limits for an entire workspace instead of individual API keys? Se ## Overview -For enterprises deploying AI at scale, maintaining financial oversight and operational control is crucial. The Prisma AIRS AI Gateway's governance features for API keys provide finance teams, IT departments, and executives with the transparency and guardrails needed to confidently scale AI adoption across the organisation. +For organisations deploying AI at scale, maintaining financial oversight and operational control is crucial. The Prisma AIRS AI Gateway's governance features for API keys provide finance teams, IT departments, and executives with the transparency and guardrails needed to confidently scale AI adoption across the organisation. By implementing budget and rate limits on API keys at both organisation and workspace levels, you can: diff --git a/aigw/product/administration/enforce-default-config.mdx b/aigw/product/administration/enforce-default-config.mdx index 806b16dd..9e576854 100644 --- a/aigw/product/administration/enforce-default-config.mdx +++ b/aigw/product/administration/enforce-default-config.mdx @@ -7,10 +7,6 @@ description: "Learn how to attach default configs to API keys for enforcing gove The Prisma AIRS AI Gateway allows you to attach default configs to API keys, enabling you to enforce specific routing rules, security controls, and other governance measures across all API calls made with those keys. This feature provides a powerful way to implement organisation-wide policies without requiring changes to individual application code. - - This feature is available on all AI Gateway plans. - - ## How It Works When you attach a default config to an API key: @@ -161,7 +157,7 @@ This flexibility allows for centralized governance while still enabling exceptio ## Use Cases -### Enterprise AI Governance +### AI Governance For large organisations with multiple AI applications, attaching default configs to API keys enables centralized governance: diff --git a/aigw/product/administration/enforce-saved-only-config.mdx b/aigw/product/administration/enforce-saved-only-config.mdx index f77dab05..b3644bdd 100644 --- a/aigw/product/administration/enforce-saved-only-config.mdx +++ b/aigw/product/administration/enforce-saved-only-config.mdx @@ -5,7 +5,7 @@ description: "Lock the Gateway to admin-curated providers, configs, and integrat --- - Hybrid enterprise customers require **Gateway version `2.11.0` or later** to enforce saved-only mode. On earlier versions the setting has no effect and inline configuration is not blocked. + Hybrid deployments require **Gateway version `2.11.0` or later** to enforce saved-only mode. On earlier versions the setting has no effect and inline configuration is not blocked. ## Overview diff --git a/aigw/product/administration/enforce-workspace-budget-and-rate-limits.mdx b/aigw/product/administration/enforce-workspace-budget-and-rate-limits.mdx index 206ec1bf..9a41ff48 100644 --- a/aigw/product/administration/enforce-workspace-budget-and-rate-limits.mdx +++ b/aigw/product/administration/enforce-workspace-budget-and-rate-limits.mdx @@ -111,4 +111,4 @@ Use the Admin API to programmatically manage workspace budgets and rate limits. ## Availability -Workspace Budget Limits are available to AI Gateway Enterprise customers and select Pro users. To enable these features for your account, please contact [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). +Workspace Budget Limits must be enabled for your organisation. diff --git a/aigw/product/ai-gateway/automatic-retries.mdx b/aigw/product/ai-gateway/automatic-retries.mdx index c8a51b9c..c3deefbb 100644 --- a/aigw/product/ai-gateway/automatic-retries.mdx +++ b/aigw/product/ai-gateway/automatic-retries.mdx @@ -3,10 +3,6 @@ title: "Automatic Retries" description: Automatically retry failed LLM requests with exponential backoff. --- - -Available on all Prisma AIRS AI Gateway plans. - - - Up to **5 retry attempts** - Trigger on **specific error codes** - **Exponential backoff** to prevent overload diff --git a/aigw/product/ai-gateway/batches.mdx b/aigw/product/ai-gateway/batches.mdx index aaaca903..2dbffc00 100644 --- a/aigw/product/ai-gateway/batches.mdx +++ b/aigw/product/ai-gateway/batches.mdx @@ -22,7 +22,7 @@ AI Gateway lets you send a single request that fan‑outs to hundreds—or milli Have the following ready to start making batch requests: 1. **Strata Cloud Manager account & API key**. -2. [Data Service](/aigw/changelog/data-service) to be enabled — required for **AI Gateway Managed Batching** or when **cost-attribution** is needed. +2. [Data Service](/changelog/data-service) to be enabled — required for **AI Gateway Managed Batching** or when **cost-attribution** is needed. 3. **Provider credentials** for each downstream model (OpenAI key, Bedrock IAM role, etc.). 4. A **AI Gateway File** (`input_file_id`) - **required only when using the AI Gateway Batch API (Mode #2)**. See [Files](/aigw/product/ai-gateway/files) to upload one. 5. Optional: Familiarity with the [Create Batch OpenAPI spec](/aigw/api-reference/batch/create-batch). diff --git a/aigw/product/ai-gateway/canary-testing.mdx b/aigw/product/ai-gateway/canary-testing.mdx index 6bfdbaa1..319b0f86 100644 --- a/aigw/product/ai-gateway/canary-testing.mdx +++ b/aigw/product/ai-gateway/canary-testing.mdx @@ -2,9 +2,6 @@ title: "Canary Testing" description: "You can use Prisma AIRS AI Gateway to also canary test new models or prompts in different environments. " --- - -This feature is available on all AI Gateway plans. - This uses the same techniques as [load balancing](/aigw/product/ai-gateway/load-balancing) but to achieve a different outcome. ### Example: Test Llama2 on 5% of the traffic diff --git a/aigw/product/ai-gateway/chat-completions.mdx b/aigw/product/ai-gateway/chat-completions.mdx index 7617b95a..a7f6eeeb 100644 --- a/aigw/product/ai-gateway/chat-completions.mdx +++ b/aigw/product/ai-gateway/chat-completions.mdx @@ -3,10 +3,6 @@ title: "Chat Completions" description: "Use OpenAI-compatible Chat Completions with any LLM provider through Prisma AIRS AI Gateway." --- - -Available on all AI Gateway plans. - - The [Chat Completions](https://platform.openai.com/docs/api-reference/chat) API is the most widely adopted format for LLM interaction. The AI Gateway makes it work with **every provider** — send the same `POST /v1/chat/completions` request to OpenAI, Anthropic, Gemini, Bedrock, or any of the 3,000+ supported models. ## Quick Start diff --git a/aigw/product/ai-gateway/circuit-breaker.mdx b/aigw/product/ai-gateway/circuit-breaker.mdx index 63739848..b723695b 100644 --- a/aigw/product/ai-gateway/circuit-breaker.mdx +++ b/aigw/product/ai-gateway/circuit-breaker.mdx @@ -3,10 +3,6 @@ title: "Circuit Breaker" description: Automatically stop routing to unhealthy targets until they recover. --- - -Available on all Prisma AIRS AI Gateway plans. - - Circuit breakers prevent cascading failures by temporarily blocking requests to targets that are failing. ## Config Schema diff --git a/aigw/product/ai-gateway/conditional-routing.mdx b/aigw/product/ai-gateway/conditional-routing.mdx index 3b60b921..1118c45e 100644 --- a/aigw/product/ai-gateway/conditional-routing.mdx +++ b/aigw/product/ai-gateway/conditional-routing.mdx @@ -2,10 +2,6 @@ title: 'Conditional Routing' --- - -This feature is available on all Prisma AIRS AI Gateway plans. - - Using AI Gateway, you can route your requests to different provider targets based on custom conditions you define. These can be conditions like: * If this user is on the `paid plan`, route their request to a `custom fine-tuned model` @@ -910,7 +906,3 @@ Conditional routing targets are fully composable — each target can itself be a 5. **Multiple Parameter Types**: The AI Gateway supports routing based on any parameter that can be sent in LLM requests including `model`, `temperature`, `top_p`, `frequency_penalty`, `presence_penalty`, `max_tokens`, and many others. 6. **Supported Query Paths**: `metadata.`, `params.`, and `url.pathname`. Nested paths beyond these (for example, `metadata.a.b`) are not supported. - - -[Reach out to us](https://support.portkey.ai/forms/customer-portal-ticket-form) to share your thoughts on this feature or to get help with setting up advanced routing scenarios. - diff --git a/aigw/product/ai-gateway/configs.mdx b/aigw/product/ai-gateway/configs.mdx index 0bb24a78..f6b7625f 100644 --- a/aigw/product/ai-gateway/configs.mdx +++ b/aigw/product/ai-gateway/configs.mdx @@ -1,10 +1,7 @@ --- title: "Configs" -description: This feature is available on all Prisma AIRS AI Gateway plans. +description: Define fallbacks, load balancing, retries, caching, and other routing rules for your gateway requests as reusable JSON configs. --- - -Available on all AI Gateway plans. - Configs streamline your Gateway management, enabling you to programmatically control various aspects like fallbacks, load balancing, retries, caching, and more. A configuration is a JSON object that can be used to define routing rules for all the requests coming to your gateway. You can configure multiple configs and use them in your requests. diff --git a/aigw/product/ai-gateway/custom-hosts.mdx b/aigw/product/ai-gateway/custom-hosts.mdx index facd4e6a..053954d9 100644 --- a/aigw/product/ai-gateway/custom-hosts.mdx +++ b/aigw/product/ai-gateway/custom-hosts.mdx @@ -4,7 +4,7 @@ description: "Route requests to privately hosted or local models using custom ho --- -Custom hosts with private network routing — including the `TRUSTED_CUSTOM_HOSTS` allowlist — applies only to **hybrid and air-gapped enterprise deployments** of the AI Gateway. On AI Gateway SaaS, custom host URLs must be publicly reachable; routing to private or internal network IPs is not supported. +Custom hosts with private network routing — including the `TRUSTED_CUSTOM_HOSTS` allowlist — applies only to **hybrid deployments** of the AI Gateway. On AI Gateway SaaS, custom host URLs must be publicly reachable; routing to private or internal network IPs is not supported. The `custom_host` parameter routes AI Gateway requests to your own model endpoints — whether running locally, in a private cloud, or on custom infrastructure. The AI Gateway validates all custom host URLs to prevent [SSRF](https://owasp.org/www-community/attacks/Server_Side_Request_Forgery) attacks. @@ -216,7 +216,7 @@ HTTP and HTTPS service ports (such as `80`, `443`, `8000`, `8080`, `8443`) are a The AI Gateway maintains a trusted hosts allowlist via the `TRUSTED_CUSTOM_HOSTS` environment variable. Trusted hostnames bypass private/reserved IP range checks at request time, and their DNS-resolved addresses bypass private/reserved IP checks at connection time. -`TRUSTED_CUSTOM_HOSTS` is available only on [self-hosted](/aigw/self-hosting/hybrid-deployments/architecture) hybrid and air-gapped enterprise deployments of the AI Gateway. +`TRUSTED_CUSTOM_HOSTS` is available only on [hybrid deployments](/aigw/self-hosting/hybrid-deployments/architecture) of the AI Gateway. ### Default behaviour @@ -288,7 +288,7 @@ Even for trusted hosts, the AI Gateway enforces: |----------|-----------------|---------------| | **Local development** (e.g., Ollama on `localhost`) | Allowed in non-production — `localhost` and `127.0.0.1` are trusted by default | None in dev. In production, add to `TRUSTED_CUSTOM_HOSTS`. See the [Ollama integration guide](/aigw/integrations/llms/ollama). | | **Docker containers** (`host.docker.internal`) | Allowed in non-production — trusted by default | In production, add to `TRUSTED_CUSTOM_HOSTS` | -| **Private network IP** (e.g., `172.31.2.45:8008`) | Blocked — falls within `172.16.0.0/12` | Add the IP to `TRUSTED_CUSTOM_HOSTS` (hybrid/air-gapped only) | +| **Private network IP** (e.g., `172.31.2.45:8008`) | Blocked — falls within `172.16.0.0/12` | Add the IP to `TRUSTED_CUSTOM_HOSTS` (hybrid only) | | **Internal DNS name** (e.g., `llm.svc.internal`) | Blocked — `.internal` TLD | Add the hostname (or `*.svc.internal`) to `TRUSTED_CUSTOM_HOSTS` | | **Cloud metadata** (`169.254.169.254`) | Always blocked | Cannot be allowlisted for security reasons | | **DNS rebinding** (`127.0.0.1.nip.io`) | Always blocked | Cannot be allowlisted for security reasons | diff --git a/aigw/product/ai-gateway/fallbacks.mdx b/aigw/product/ai-gateway/fallbacks.mdx index 84bc7e84..a3c9e38f 100644 --- a/aigw/product/ai-gateway/fallbacks.mdx +++ b/aigw/product/ai-gateway/fallbacks.mdx @@ -3,10 +3,6 @@ title: "Fallbacks" description: Automatically switch to backup LLMs when the primary fails. --- - -Available on all Prisma AIRS AI Gateway plans. - - Specify a prioritized list of providers/models. If the primary LLM fails, the AI Gateway automatically falls back to the next in line. ## Examples diff --git a/aigw/product/ai-gateway/fine-tuning.mdx b/aigw/product/ai-gateway/fine-tuning.mdx index c1f3e45d..da7707aa 100644 --- a/aigw/product/ai-gateway/fine-tuning.mdx +++ b/aigw/product/ai-gateway/fine-tuning.mdx @@ -4,7 +4,7 @@ description: "Run your fine-tuning jobs with Prisma AIRS AI Gateway" --- - [Data Service](/aigw/changelog/data-service) must be enabled for **AI Gateway Managed Fine-tuning** or when **cost-attribution** is required. + [Data Service](/changelog/data-service) must be enabled for **AI Gateway Managed Fine-tuning** or when **cost-attribution** is required. AI Gateway supports fine-tuning in two ways: @@ -90,7 +90,7 @@ curl -X POST --header 'Authorization: Bearer ' \ 'https://aigw.portkey.ai/v1/fine_tuning/jobs//cancel' ``` -When you use the AI Gateway as a provider for fine-tuning, everything is managed by the AI Gateway hosted solution, this includes job submission with provider, monitoring job status and able to use the fine-tuned model with the AI Gateway's Prompt Playground. +When you use the AI Gateway as a provider for fine-tuning, everything is managed by the AI Gateway hosted solution, this includes job submission with provider, monitoring job status and using the fine-tuned model through the AI Gateway. Providers like OpenAI, Azure OpenAPI does provider a straight forward approach diff --git a/aigw/product/ai-gateway/load-balancing.mdx b/aigw/product/ai-gateway/load-balancing.mdx index 760d2a48..2c53e9b2 100644 --- a/aigw/product/ai-gateway/load-balancing.mdx +++ b/aigw/product/ai-gateway/load-balancing.mdx @@ -3,10 +3,6 @@ title: "Load Balancing" description: Distribute traffic across multiple LLMs for high availability and optimal performance. --- - -Available on all Prisma AIRS AI Gateway plans. - - Distribute traffic across multiple LLMs to prevent any single provider from becoming a bottleneck. ## Examples diff --git a/aigw/product/ai-gateway/messages-api.mdx b/aigw/product/ai-gateway/messages-api.mdx index 7766f713..3254f0ff 100644 --- a/aigw/product/ai-gateway/messages-api.mdx +++ b/aigw/product/ai-gateway/messages-api.mdx @@ -3,8 +3,6 @@ title: "Messages" description: "Use Anthropic's Messages format with any provider — 3,000+ models, one endpoint." --- -Available on all Prisma AIRS AI Gateway plans. - The AI Gateway's `/v1/messages` endpoint accepts the [Anthropic Messages API](https://docs.anthropic.com/en/api/messages) format and routes to any of [3,000+ models](/aigw/product/model-catalog) across all major providers. Tools built natively on the Messages format — like [Claude Code](/aigw/integrations/libraries/claude-code) and the [Claude Agent SDK](/aigw/integrations/agents/claude-agent-sdk) — work with any backend model through the AI Gateway without modification. ## Why Messages API diff --git a/aigw/product/ai-gateway/multimodal-capabilities.mdx b/aigw/product/ai-gateway/multimodal-capabilities.mdx index 435be062..fcacc884 100644 --- a/aigw/product/ai-gateway/multimodal-capabilities.mdx +++ b/aigw/product/ai-gateway/multimodal-capabilities.mdx @@ -1,10 +1,6 @@ --- title: "Multimodal Capabilities" --- - -This feature is available on all Prisma AIRS AI Gateway plans. - - The Gateway is your unified interface for **multimodal models**, along with chat, text, and embedding models. Using the Gateway, you can call `vision`, `audio (text-to-speech & speech-to-text)`, `image generation` and other multimodal models from multiple providers (like `OpenAI`, `Anthropic`, `Stability AI`, etc.) — all using the familiar OpenAI signature. diff --git a/aigw/product/ai-gateway/multimodal-capabilities/function-calling.mdx b/aigw/product/ai-gateway/multimodal-capabilities/function-calling.mdx index c10eafb6..0b4337e8 100644 --- a/aigw/product/ai-gateway/multimodal-capabilities/function-calling.mdx +++ b/aigw/product/ai-gateway/multimodal-capabilities/function-calling.mdx @@ -137,16 +137,10 @@ On completion, the request logs in Strata Cloud Manager where tools and function -## Managing Functions and Tools in Prompts - -Create prompt templates with function/tool definitions in the AI Gateway's Prompt Library. Set the `tool_choice` parameter, and the AI Gateway validates your tool definition on the fly, eliminating syntax errors. - ## Supported Providers and Models The AI Gateway provides native function calling support across all major AI providers. -If you discover a function-calling capable LLM that isn't working with the AI Gateway, please let us know via [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). - ## Cookbook [**See the detailed cookbook on function calling.**](/guides/getting-started/function-calling) diff --git a/aigw/product/ai-gateway/multimodal-capabilities/thinking-mode.mdx b/aigw/product/ai-gateway/multimodal-capabilities/thinking-mode.mdx index 1471cfbe..b42d41ba 100644 --- a/aigw/product/ai-gateway/multimodal-capabilities/thinking-mode.mdx +++ b/aigw/product/ai-gateway/multimodal-capabilities/thinking-mode.mdx @@ -10,8 +10,6 @@ These reasoning-optimized models are built to excel in tasks requiring complex a The AI Gateway currently supports thinking-enabled models from **Anthropic**, **Google Vertex AI**, **Amazon Bedrock**, **OpenAI**, **Together AI**, and other **OpenAI-compatible providers**. -*Note: If a specific model is not supported for thinking on the AI Gateway, please reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form).* - ## Using Thinking Mode 1. You must set `strict_open_ai_compliance=False` in your headers or client configuration diff --git a/aigw/product/ai-gateway/multimodal-capabilities/vision.mdx b/aigw/product/ai-gateway/multimodal-capabilities/vision.mdx index 95e4e41c..d4ea525a 100644 --- a/aigw/product/ai-gateway/multimodal-capabilities/vision.mdx +++ b/aigw/product/ai-gateway/multimodal-capabilities/vision.mdx @@ -121,10 +121,6 @@ On completion, the request will get logged in the logs UI where any image inputs -## Creating prompt templates for vision models - -The AI Gateway's prompt library supports creating templates with image inputs. If the same image will be used in all prompt calls, you can save it as part of the template's image URL itself. Or, if the image will be sent via the API as a variable, add a variable to the image link. - ## Supported Providers and Models The AI Gateway supports all vision models from its integrated providers as they become available. The table below shows some examples of supported vision models. Please raise a [request](/aigw/integrations/llms/suggest-a-new-integration) to add a provider to the AI gateway. diff --git a/aigw/product/ai-gateway/realtime-api.mdx b/aigw/product/ai-gateway/realtime-api.mdx index e9e59e40..a4e0bd83 100644 --- a/aigw/product/ai-gateway/realtime-api.mdx +++ b/aigw/product/ai-gateway/realtime-api.mdx @@ -3,10 +3,6 @@ title: "Realtime API" description: "Use OpenAI's Realtime API with logs, cost tracking, and more!" --- - - This feature is available on all Prisma AIRS AI Gateway plans. - - [OpenAI's Realtime API](https://platform.openai.com/docs/guides/realtime) while the fastest way to use multi-modal generation, presents its own set of problems around logging, cost tracking and guardrails. AI Gateway provides a solution to these problems with a seamless integration. Its logging is unique in that it captures the entire request and response, including the model's response, cost, and guardrail violations. diff --git a/aigw/product/ai-gateway/request-timeouts.mdx b/aigw/product/ai-gateway/request-timeouts.mdx index 9b3fed7d..f33bd4fc 100644 --- a/aigw/product/ai-gateway/request-timeouts.mdx +++ b/aigw/product/ai-gateway/request-timeouts.mdx @@ -3,10 +3,6 @@ title: "Request Timeouts" description: "Manage unpredictable LLM latencies effectively with Prisma AIRS AI Gateway's **Request Timeouts**." --- - - This feature is available on all AI Gateway plans. - - This feature allows automatic termination of requests that exceed a specified duration, letting you gracefully handle errors or make another, faster request. ## Enabling Request Timeouts diff --git a/aigw/product/ai-gateway/responses-api.mdx b/aigw/product/ai-gateway/responses-api.mdx index 752188e9..341a31f2 100644 --- a/aigw/product/ai-gateway/responses-api.mdx +++ b/aigw/product/ai-gateway/responses-api.mdx @@ -3,10 +3,6 @@ title: "Open Responses" description: "Use the Responses API with any LLM provider through Prisma AIRS AI Gateway — fully compliant with the Open Responses specification." --- - -Available on all AI Gateway plans. - - [Open Responses](https://www.openresponses.org/) is an open-source specification for multi-provider, interoperable LLM interfaces based on the OpenAI Responses API. It defines a shared schema for calling language models, streaming results, and composing agentic workflows — independent of provider. **the AI Gateway is fully Open Responses compliant.** The Responses API works with every provider and model in the AI Gateway's catalog — including Anthropic, Gemini, Bedrock, and 60+ other providers that don't natively support it. diff --git a/aigw/product/ai-gateway/universal-api.mdx b/aigw/product/ai-gateway/universal-api.mdx index 06c65025..5c80c0e2 100644 --- a/aigw/product/ai-gateway/universal-api.mdx +++ b/aigw/product/ai-gateway/universal-api.mdx @@ -4,10 +4,6 @@ sidebarTitle: "Overview" description: "One API for 3,000+ LLMs across every major provider. Use OpenAI's Chat Completions, Responses API, or Anthropic's Messages format -- Prisma AIRS AI Gateway translates between them all." --- - -Available on all AI Gateway plans. - - AI Gateway provides a single, unified API for 3,000+ models from every major provider. Write once in any format, switch providers by changing one parameter. ## Three API Formats, Any Provider diff --git a/aigw/product/ai-gateway/virtual-keys/bedrock-amazon-assumed-role.mdx b/aigw/product/ai-gateway/virtual-keys/bedrock-amazon-assumed-role.mdx index e11d0cf1..a9222907 100644 --- a/aigw/product/ai-gateway/virtual-keys/bedrock-amazon-assumed-role.mdx +++ b/aigw/product/ai-gateway/virtual-keys/bedrock-amazon-assumed-role.mdx @@ -2,10 +2,6 @@ title: "Connect Bedrock with Amazon Assumed Role" description: "How to create a new integration for Bedrock using Amazon Assumed Role Authentication" --- - -Available on all plans. - - Connecting to Bedrock using an AWS assumed role? Check out the documentation here. @@ -53,7 +49,7 @@ arn:aws:iam::039293892788:role/AirsGwEnterpriseRole The above ARN applies to the SaaS deployment of the AI Gateway, managed through [Strata Cloud Manager](https://stratacloudmanager.paloaltonetworks.com/).
-To enable **Assumed Role for AWS in your AI Gateway Enterprise deployment**, you can refer to [this guide](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/docs/Bedrock.md). If you face any issue, please reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). +To enable **Assumed Role for AWS in your hybrid deployment**, you can refer to [this guide](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/docs/Bedrock.md).
Paste the following JSON into the trust policy editor and click *Update Trust Policy*. diff --git a/aigw/product/enterprise-offering/access-control-management.mdx b/aigw/product/enterprise-offering/access-control-management.mdx index f16a115b..1b5155e7 100644 --- a/aigw/product/enterprise-offering/access-control-management.mdx +++ b/aigw/product/enterprise-offering/access-control-management.mdx @@ -3,11 +3,11 @@ title: "Access Control Management" description: "With customizable user roles, API key management, and comprehensive audit logs, Prisma AIRS AI Gateway provides the flexibility and control needed to ensure secure collaboration & maintain a strong security posture" --- -At the AI Gateway, we understand the critical importance of access control and data security for enterprise customers. Our platform provides a robust and flexible access control management system that enables you to safeguard your sensitive information while empowering your teams to collaborate effectively. +At the AI Gateway, we understand the critical importance of access control and data security. Our platform provides a robust and flexible access control management system that enables you to safeguard your sensitive information while empowering your teams to collaborate effectively. ## 1\. Isolated and Customizable Organisations -The AI Gateway's enterprise version allows you to create multiple `organizations`, each serving as a secure and isolated environment for your teams or projects. This multi-tenant architecture ensures that your data, logs, analytics, prompts, integrations, providers, configs, guardrails, and API keys are strictly confined within each `organization`, preventing unauthorized access and maintaining data confidentiality. +The AI Gateway allows you to create multiple `organizations`, each serving as a secure and isolated environment for your teams or projects. This multi-tenant architecture ensures that your data, logs, analytics, prompts, integrations, providers, configs, guardrails, and API keys are strictly confined within each `organization`, preventing unauthorized access and maintaining data confidentiality. With the ability to create and manage multiple organisations, you can tailor access control to match your company's structure and project requirements. Users can be assigned to specific organisations, and they can seamlessly switch between them using the AI Gateway's intuitive user interface. diff --git a/aigw/product/enterprise-offering/audit-logs.mdx b/aigw/product/enterprise-offering/audit-logs.mdx index 41ae21bb..d6204c3a 100644 --- a/aigw/product/enterprise-offering/audit-logs.mdx +++ b/aigw/product/enterprise-offering/audit-logs.mdx @@ -76,9 +76,9 @@ The AI Gateway provides powerful filtering options to help you find specific aud - **Country**: Filter by geographic location of requests - **Time Range**: Filter logs within a specific time period -## Enterprise Features +## Key Capabilities -The AI Gateway's Audit Logs include enterprise-grade capabilities: +The AI Gateway's Audit Logs provide: ### 1. Complete Visibility - Full user attribution for every action @@ -87,11 +87,10 @@ The AI Gateway's Audit Logs include enterprise-grade capabilities: - Searchable audit trail ### 2. Compliance & Security -- SOC 2, ISO 27001, GDPR, and HIPAA compliant - PII data protection - Indefinite log retention -### 3. Enterprise-Grade Features +### 3. Governance Controls - Role-based access control - Cross-organisation visibility - Custom retention policies diff --git a/aigw/product/enterprise-offering/budget-policies.mdx b/aigw/product/enterprise-offering/budget-policies.mdx index 77f34a65..0e180332 100644 --- a/aigw/product/enterprise-offering/budget-policies.mdx +++ b/aigw/product/enterprise-offering/budget-policies.mdx @@ -519,29 +519,7 @@ Limit requests using a specific gateway config to 200 RPM. --- -### Use Case 9: Prompt Specific Usage Budget - -Limit a specific prompt template to 1 million tokens per month. - -```json -{ - "type": "usage_limits", - "policy": { - "conditions": [ - { "key": "prompt", "value": "customer-support-v2" } - ], - "group_by": [ { "key": "prompt"}], - "credit_limit": 1000000, - "type": "tokens", - "periodic_reset": "monthly", - "status": "active" - } -} -``` - ---- - -### Use Case 10: Multiple Allowed Values (OR Logic) +### Use Case 9: Multiple Allowed Values (OR Logic) Allow only specific API keys to use expensive models, with a combined rate limit. @@ -564,7 +542,7 @@ Allow only specific API keys to use expensive models, with a combined rate limit --- -### Use Case 11: Exclude Specific Models +### Use Case 10: Exclude Specific Models Apply rate limit to all OpenAI models EXCEPT GPT-4o. @@ -586,7 +564,7 @@ Apply rate limit to all OpenAI models EXCEPT GPT-4o. --- -### Use Case 12: Per User, Per Model Budget +### Use Case 11: Per User, Per Model Budget Track spend separately for each user AND model combination. @@ -611,7 +589,7 @@ Track spend separately for each user AND model combination. --- -### Use Case 13: Team Based Provider Quota +### Use Case 12: Team Based Provider Quota Limit each team (from metadata) to specific token quotas per provider. @@ -636,7 +614,7 @@ Limit each team (from metadata) to specific token quotas per provider. --- -### Use Case 14: Exclude Internal API Keys from Limits +### Use Case 13: Exclude Internal API Keys from Limits Apply rate limits to all API keys except internal ones. @@ -660,7 +638,7 @@ Apply rate limits to all API keys except internal ones. --- -### Use Case 15: Combined Conditions - Premium Users on Specific Models +### Use Case 14: Combined Conditions - Premium Users on Specific Models Rate limit premium tier users only when using expensive models. @@ -685,7 +663,7 @@ Rate limit premium tier users only when using expensive models. --- -### Use Case 16: Global MCP Rate Limit +### Use Case 15: Global MCP Rate Limit Limit all MCP tool calls in a workspace to 500 requests per minute. @@ -706,7 +684,7 @@ Limit all MCP tool calls in a workspace to 500 requests per minute. --- -### Use Case 17: Per-User MCP Rate Limit +### Use Case 16: Per-User MCP Rate Limit Limit each user to 50 MCP tool calls per minute using metadata. @@ -731,7 +709,7 @@ Limit each user to 50 MCP tool calls per minute using metadata. --- -### Use Case 18: Per-Server Token Limit +### Use Case 17: Per-Server Token Limit Limit token consumption on the GitHub MCP server to 500,000 tokens per day. @@ -753,7 +731,7 @@ Limit token consumption on the GitHub MCP server to 500,000 tokens per day. --- -### Use Case 19: Per-API-Key Limits Grouped by MCP Server +### Use Case 18: Per-API-Key Limits Grouped by MCP Server Give each API key its own rate limit, tracked independently per MCP server. @@ -779,7 +757,7 @@ Give each API key its own rate limit, tracked independently per MCP server. --- -### Use Case 20: Exclude Specific Tools from a Server Limit +### Use Case 19: Exclude Specific Tools from a Server Limit Apply rate limit to all Slack MCP tools except `list_channels`. diff --git a/aigw/product/enterprise-offering/components.mdx b/aigw/product/enterprise-offering/components.mdx index 431d8cc0..6182b14d 100644 --- a/aigw/product/enterprise-offering/components.mdx +++ b/aigw/product/enterprise-offering/components.mdx @@ -1,8 +1,8 @@ --- -title: "Enterprise Components" +title: "Deployment Components" --- -The Prisma AIRS AI Gateway's Enterprise Components provide the core infrastructure needed for production deployments. Each component handles a specific function - analytics, logging, or caching - with multiple implementation options to match your requirements. +The Prisma AIRS AI Gateway's deployment components provide the core infrastructure needed for production deployments. Each component handles a specific function - analytics, logging, or caching - with multiple implementation options to match your requirements. --- @@ -22,7 +22,7 @@ The AI Gateway leverages Clickhouse as the primary Analytics Store for the Contr ## Log Store -The AI Gateway provides flexible options for storing and managing logs in your enterprise deployment. Choose from various storage solutions including MongoDB for document-based storage, AWS S3 for cloud-native object storage, or Wasabi for cost-effective cloud storage. Each option offers different benefits in terms of scalability, cost, and integration capabilities. +The AI Gateway provides flexible options for storing and managing logs in your hybrid deployment. Choose from various storage solutions including MongoDB for document-based storage, AWS S3 for cloud-native object storage, or Wasabi for cost-effective cloud storage. Each option offers different benefits in terms of scalability, cost, and integration capabilities. @@ -45,15 +45,11 @@ The `v2` format organizes logs by time hierarchy, making it easier to manage ret Changing `LOG_STORE_FILE_PATH_FORMAT` only affects newly written logs. Previously written logs retain their original path format and are not migrated.
- -The `v2` format is not supported for air-gapped deployments where `LOG_STORE` is set to `control_plane`. - - --- ## Cache Store -The AI Gateway supports robust caching solutions to optimize performance and reduce latency in your enterprise deployment. Choose between Redis for in-memory caching or AWS ElastiCache for a fully managed caching service. +The AI Gateway supports robust caching solutions to optimize performance and reduce latency in your hybrid deployment. Choose between Redis for in-memory caching or AWS ElastiCache for a fully managed caching service. diff --git a/aigw/product/enterprise-offering/kms.mdx b/aigw/product/enterprise-offering/kms.mdx index 38d6e0a3..4d823588 100644 --- a/aigw/product/enterprise-offering/kms.mdx +++ b/aigw/product/enterprise-offering/kms.mdx @@ -65,9 +65,7 @@ arn:aws:iam::299329113195:role/EnterpriseKMSPolicy ``` -The above ARN only works when management plane is hosted on [hosted app](https://stratacloudmanager.paloaltonetworks.com/).
- -To enable KMS for AWS in your AI Gateway Enterprise self hosted management plane deployment. Please reach out to your AI Gateway representative or contact [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). +The above ARN only works when management plane is hosted on [hosted app](https://stratacloudmanager.paloaltonetworks.com/).
## AWS KMS Key Creation Guide diff --git a/aigw/product/enterprise-offering/org-management.mdx b/aigw/product/enterprise-offering/org-management.mdx index bae93147..1c3459d2 100644 --- a/aigw/product/enterprise-offering/org-management.mdx +++ b/aigw/product/enterprise-offering/org-management.mdx @@ -3,7 +3,7 @@ title: "Org Management" description: "A high-level introduction to Prisma AIRS AI Gateway's organisation management structure and key concepts." --- -The AI Gateway's organisation management structure provides a hierarchical system for managing teams, resources, and access within your AI development environment. This structure is designed to offer flexibility and security for enterprises of various sizes. +The AI Gateway's organisation management structure provides a hierarchical system for managing teams, resources, and access within your AI development environment. This structure is designed to offer flexibility and security for organisations of various sizes. The account hierarchy in the AI Gateway is organized as follows: diff --git a/aigw/product/enterprise-offering/org-management/api-key-rotation.mdx b/aigw/product/enterprise-offering/org-management/api-key-rotation.mdx index 31ca558a..a172b0ee 100644 --- a/aigw/product/enterprise-offering/org-management/api-key-rotation.mdx +++ b/aigw/product/enterprise-offering/org-management/api-key-rotation.mdx @@ -6,9 +6,7 @@ description: "Periodically replace API key secrets without downtime using manual **Availability:** API key rotation is available on Prisma AIRS AI Gateway Cloud. -For self-hosted deployments: - - **Airgapped:** Backend `v1.14.0+` - - **Hybrid:** Gateway Enterprise `v2.5.0+` +For hybrid deployments, it requires Gateway `v2.5.0+`. diff --git a/aigw/product/enterprise-offering/org-management/jwt.mdx b/aigw/product/enterprise-offering/org-management/jwt.mdx index 9b539a5c..379ffe67 100644 --- a/aigw/product/enterprise-offering/org-management/jwt.mdx +++ b/aigw/product/enterprise-offering/org-management/jwt.mdx @@ -5,7 +5,7 @@ description: Configure JWT-based authentication for your organisation in Prisma The AI Gateway supports JWT-based authentication in addition to API key authentication. Clients send a JWT in `x-portkey-api-key` or `Authorization: Bearer`; the token is validated against JWKS configured on the organisation. -**AI Gateway Cloud** validates JWTs on the management plane. **[Hybrid](/aigw/self-hosting/hybrid-deployments/architecture) and [air-gapped](/self-hosting/airgapped/model-pricing) deployments** can optionally enable **gateway-local JWT authentication** so the AI Gateway validates tokens locally instead of delegating every request to the management plane. +**AI Gateway Cloud** validates JWTs on the management plane. **[Hybrid deployments](/aigw/self-hosting/hybrid-deployments/architecture)** can optionally enable **gateway-local JWT authentication** so the AI Gateway validates tokens locally instead of delegating every request to the management plane. Optionally validate JWTs inside a config's hook pipeline using the JWT guardrail plugin. This is separate from gateway-local JWT auth, which replaces API key validation at the gateway. @@ -197,7 +197,7 @@ If you prefer Python for signing, you can generate the RSA key pair using your p 3. If valid, the request is authenticated and user details are extracted for authorization and logging. 4. If invalid, the request is rejected. Common responses include **401 Unauthorized** (invalid or expired token), **403 Forbidden** (scope or workspace not allowed), **412** (usage limit exceeded), and **429** (rate limit exceeded). -On hybrid and air-gapped gateways with `JWT_ENABLED=ON`, validation runs locally on the gateway. Otherwise, JWT-shaped tokens are validated by the management plane. +On hybrid gateways with `JWT_ENABLED=ON`, validation runs locally on the gateway. Otherwise, JWT-shaped tokens are validated by the management plane. ## Authorization & Scopes @@ -469,9 +469,9 @@ main(); ## Gateway-Local JWT Authentication -Gateway-local JWT authentication is available only on **[hybrid](/aigw/self-hosting/hybrid-deployments/architecture)** and **air-gapped** deployments. +Gateway-local JWT authentication is available only on **[hybrid](/aigw/self-hosting/hybrid-deployments/architecture)** deployments. -It requires gateway **2.5.0** or higher (Backend **v1.13.0** or higher for air-gapped) with `JWT_ENABLED=ON`. +It requires gateway **2.5.0** or higher with `JWT_ENABLED=ON`. When enabled, your AI Gateway validates JWTs locally instead of sending every request to the management plane for authentication. This reduces latency and keeps authentication working even when the management plane is temporarily unreachable. @@ -705,7 +705,7 @@ sequenceDiagram - Set `JWT_ENABLED=ON`, `PORTKEY_CLIENT_AUTH`, and `ORGANISATIONS_TO_SYNC` on the gateway - Configure JWKS for your org in the AI Gateway admin UI -- Ensure the gateway cache is enabled (required for hybrid and air-gapped deployments) +- Ensure the gateway cache is enabled (required for hybrid deployments) **Mode A:** diff --git a/aigw/product/enterprise-offering/org-management/scim/azure-ad.mdx b/aigw/product/enterprise-offering/org-management/scim/azure-ad.mdx index 636ed2e3..aff20c47 100644 --- a/aigw/product/enterprise-offering/org-management/scim/azure-ad.mdx +++ b/aigw/product/enterprise-offering/org-management/scim/azure-ad.mdx @@ -32,7 +32,6 @@ First, create a new Azure Entra application to set up SCIM provisioning with the - [AI Gateway Settings Page](https://stratacloudmanager.paloaltonetworks.com/) 6. Fill in the values from Strata Cloud Manager in Entra's provisioning settings and click **`Test Connection`**. If successful, click **`Create`**. -> If the test connection returns any errors, please contact [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). --- diff --git a/aigw/product/enterprise-offering/org-management/scim/group-management.mdx b/aigw/product/enterprise-offering/org-management/scim/group-management.mdx index df7bbd30..3e8f2904 100644 --- a/aigw/product/enterprise-offering/org-management/scim/group-management.mdx +++ b/aigw/product/enterprise-offering/org-management/scim/group-management.mdx @@ -192,27 +192,6 @@ If you need to remove users from a workspace through the UI, remove them via wor --- -## User-Based Group Management Mode (AirGapped only) - -By default, group memberships are managed through the SCIM `/Groups` endpoint. For identity providers that manage group assignments through user attributes (e.g., Azure Entra), you can enable **User-Based Group Management Mode**. - -When this mode is enabled: -- Group memberships can be managed via the SCIM `/Users` endpoint (create, update, and patch operations) -- The `groups` attribute on SCIM user responses includes the user's current group memberships -- Group member operations on the `/Groups` PATCH endpoint are skipped to avoid conflicts - -### Enabling User-Based Group Management - -Set the following environment variable on your backend deployment: - -``` -SCIM_MEMBERSHIP_USER_MODE=ON -``` - -This mode is useful when your identity provider pushes group membership updates as part of user provisioning rather than group provisioning operations. - ---- - ## Benefits The flexible group management feature provides several advantages: @@ -223,6 +202,5 @@ The flexible group management feature provides several advantages: - **Simplified management** - Manage all mappings from Management Plane - **Role consistency** - All group members automatically receive the same role across every mapped workspace - **Custom naming format** - Configure prefix and separator to match your existing group naming conventions for automatic mapping -- **User-based management** - Optionally manage group memberships via the `/Users` endpoint for providers that require it --- diff --git a/aigw/product/enterprise-offering/org-management/scim/scim.mdx b/aigw/product/enterprise-offering/org-management/scim/scim.mdx index 467aa9d8..caeb5631 100644 --- a/aigw/product/enterprise-offering/org-management/scim/scim.mdx +++ b/aigw/product/enterprise-offering/org-management/scim/scim.mdx @@ -93,7 +93,3 @@ Currently, we support SCIM provisioning for the following identity providers: - **Invalid Token**: Ensure the bearer token is correctly generated and included in the request header. - **403 Forbidden**: Check if the provided SCIM Base URL and token are correct. - **User Not Provisioned**: Ensure the user attributes meet our platform's requirements. - ---- - -For further assistance, please contact our support team via [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). diff --git a/aigw/product/enterprise-offering/org-management/sso.mdx b/aigw/product/enterprise-offering/org-management/sso.mdx index 215b772b..4c6163fa 100644 --- a/aigw/product/enterprise-offering/org-management/sso.mdx +++ b/aigw/product/enterprise-offering/org-management/sso.mdx @@ -1,9 +1,9 @@ --- title: "SSO" -description: "SSO support for enterprises" +description: "Integrate your identity provider with Prisma AIRS AI Gateway using OIDC or SAML 2.0" --- -Prisma AIRS AI Gateway supports following authentication protocols for enterprise customers. +Prisma AIRS AI Gateway supports the following authentication protocols. 1. **OIDC** (OpenID Connect) 2. **SAML 2.0** (Security Assertion Markup Language) @@ -50,8 +50,6 @@ Self-hosted deployments running in single-organisation SSO mode configure the sa - profile - email - offline_access - - For airgapped deployments, you can configure additional custom OIDC scopes by setting the `OIDC_CUSTOM_SCOPES` environment variable (comma-separated list) on the backend service. #### General diff --git a/aigw/product/enterprise-offering/otel/analytics.mdx b/aigw/product/enterprise-offering/otel/analytics.mdx index 09e00927..1f29fcdb 100644 --- a/aigw/product/enterprise-offering/otel/analytics.mdx +++ b/aigw/product/enterprise-offering/otel/analytics.mdx @@ -8,7 +8,7 @@ The AI Gateway supports sending your Analytics data to OpenTelemetry (OTel) comp ## Overview -While the AI Gateway leverages Clickhouse as the primary Analytics Store for the Control Panel by default, enterprise customers can integrate the AI Gateway's analytics data with their existing data infrastructure through OpenTelemetry. +While the AI Gateway leverages Clickhouse as the primary Analytics Store for the Control Panel by default, you can integrate the AI Gateway's analytics data with their existing data infrastructure through OpenTelemetry. This export focuses on **aggregated metrics and analytics data** such as: @@ -73,7 +73,7 @@ EXPERIMENTAL_GEN_AI_OTEL_RESOURCE_ATTRIBUTES: ApplicationShortName=agent-test,As ## Integration Options -Enterprise customers commonly use these analytics exports with: +Common destinations for these analytics exports include: - **Datadog**: Monitor and analyze your AI operations alongside other application metrics - **AWS S3**: Store analytics data for long-term retention and analysis @@ -90,11 +90,3 @@ This feature enables: - Custom analytics dashboards in your preferred tools - Integration with existing alerting systems - Compliance and audit trail requirements - -## Getting Support - -For additional assistance with setting up analytics data export: - -- Reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) - -Our team can help you with best practices for configuring your OTel collectors and integrating with your existing systems. diff --git a/aigw/product/enterprise-offering/secret-references.mdx b/aigw/product/enterprise-offering/secret-references.mdx index ea0574ba..07ff4428 100644 --- a/aigw/product/enterprise-offering/secret-references.mdx +++ b/aigw/product/enterprise-offering/secret-references.mdx @@ -13,7 +13,6 @@ This keeps sensitive material in infrastructure you already control and audit. Requires: - Gateway version **2.2.4** or higher -- Backend version **1.12.0** or higher (for air-gapped deployments). ## Supported Secret Managers @@ -268,7 +267,7 @@ If you're storing provider API keys in your vault, map your secrets in LLM Integ MCP integration secret mappings require: -- Enterprise Gateway version **2.7.0** or higher +- Gateway version **2.7.0** or higher - Frontend version **1.8.1** or higher - Backend version **1.15.0** or higher diff --git a/aigw/product/enterprise-offering/security.mdx b/aigw/product/enterprise-offering/security.mdx index 0c95df3d..1de88240 100644 --- a/aigw/product/enterprise-offering/security.mdx +++ b/aigw/product/enterprise-offering/security.mdx @@ -26,19 +26,11 @@ These encryption standards are part of our commitment to maintaining secure data Our access control measures are designed with granularity to offer precise control over who can see and manage data. -For enterprise clients, we provide enhanced [Role-Based Access Control (RBAC)](/aigw/product/enterprise-offering/access-control-management#id-2.-fine-grained-user-roles-and-permissions), allowing for stringent governance suited to complex organisational needs. +We provide [Role-Based Access Control (RBAC)](/aigw/product/enterprise-offering/access-control-management#id-2.-fine-grained-user-roles-and-permissions), allowing for stringent governance suited to complex organisational needs. This system is pivotal for enforcing security policies and ensuring that only authorized personnel have access to sensitive operations and data. -## Compliance and Data Privacy - -### Compliance with Standards - -Our platform is compliant with leading security standards, including SOC 2, ISO 27001, GDPR, and HIPAA. - -The AI Gateway undergoes regular audits, compliance checks, and penetration testing conducted by third-party security experts to ensure continuous adherence to these standards. - -These certifications demonstrate our commitment to global security practices and our ability to meet diverse regulatory requirements. +## Data Privacy ### Privacy Protections @@ -54,8 +46,6 @@ We protect our systems with advanced firewall technologies and DDoS prevention m ### **Reliability and Availability**: -The AI Gateway offers an industry-leading 99.995% uptime. - Our failover mechanisms are sophisticated, designed to handle unexpected scenarios seamlessly and without service interruption. ## Incident Management and Continuous Improvement @@ -71,9 +61,3 @@ We maintain transparent communication with our clients throughout the incident m ### Updates and Continuous Improvement Security at the AI Gateway is dynamic; we continually refine our security measures and systems to address emerging threats and incorporate best practices. Our ongoing commitment to improvement helps us stay ahead of the curve in cybersecurity and operational performance. - -## Contact Information - -For more detailed information or specific inquiries regarding our security measures, please reach out to our support team: - -* **Support**: [Submit a ticket](https://support.portkey.ai/forms/customer-portal-ticket-form) diff --git a/aigw/product/guardrails.mdx b/aigw/product/guardrails.mdx index 17975004..08fe2871 100644 --- a/aigw/product/guardrails.mdx +++ b/aigw/product/guardrails.mdx @@ -3,13 +3,6 @@ title: "Guardrails" description: "Ship to production confidently with Prisma AIRS AI Gateway Guardrails on your requests & responses" --- - -This feature is available on all plans. -* **Developer**: Access to `BASIC` Guardrails -* **Production**: Access to `BASIC`, `PARTNER`, `PRO` Guardrails. -* **Enterprise**: Access to **all** Guardrails plus `custom` Guardrails. - - LLMs are brittle - not just in API uptimes or their inexplicable `400`/`500` errors, but also in their core behaviour. You can get a response with a `200` status code that completely errors out for your app's pipeline due to mismatched output. With the AI Gateway's Guardrails, we now help you enforce LLM behaviour in real-time with our _Guardrails on the Gateway_ pattern. Use the AI Gateway's Guardrails to verify your LLM inputs AND outputs, adhering to your specifed checks. Since Guardrails are built into the gateway itself, you can orchestrate your request - with actions ranging from _denying the request_, _logging the guardrail result_, _creating an evals dataset_, _falling back to another LLM or prompt_, _retrying the request_, and more. @@ -791,6 +784,4 @@ If you already have a custom guardrail pipeline where you send your inputs/outpu By appropriately configuring Guardrail Actions, you can maintain the integrity and reliability of your AI app, ensuring that only safe and compliant requests are processed. -To enable Guardrails for your org, reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). - --- diff --git a/aigw/product/guardrails/capabilities.mdx b/aigw/product/guardrails/capabilities.mdx index 1e513c9b..1a0232fa 100644 --- a/aigw/product/guardrails/capabilities.mdx +++ b/aigw/product/guardrails/capabilities.mdx @@ -16,7 +16,6 @@ Guardrails run at the Gateway layer — provider-agnostic, between your app and | `/v1/embeddings` | ✓ | — | | `/v1/messages` (Anthropic) | ✓ | ✓ | | `/v1/responses` | ✓ | ✓ | -| `/v1/prompts/{promptId}/completions` | ✓ | ✓ | | `/v1/decisions` | ✓ (state only) | — | @@ -43,7 +42,6 @@ The following Inference API endpoints **do not run guardrails**: - **Fine-tuning (all endpoints)**: `/v1/fine_tuning/*` - **Moderations**: `/v1/moderations` - **Models**: `/v1/models` -- **Prompt rendering (no model call)**: `/v1/prompts/{promptId}/render` - **Response retrieval (no model call)**: `GET /v1/responses/*` (guardrails apply only to `POST /v1/responses`) --- @@ -121,7 +119,6 @@ Tool call arguments (JSON) are treated as text — checks like Regex Match, Cont | `/embeddings` | ✓ (input only) | | `/messages` (Anthropic) | ✓ | | `/responses` | ✓ | -| `/prompts/{promptId}/completions` | ✓ | | `/decisions` | ✓ (input only, `state` field) | | All providers & models | ✓ | | Input guardrails | ✓ | diff --git a/aigw/product/guardrails/embedding-guardrails.mdx b/aigw/product/guardrails/embedding-guardrails.mdx index 82f0d5f1..9568cceb 100644 --- a/aigw/product/guardrails/embedding-guardrails.mdx +++ b/aigw/product/guardrails/embedding-guardrails.mdx @@ -15,7 +15,7 @@ Without proper guardrails, sensitive customer data can leak into vector database 3. **Cost optimization**: Avoid unnecessary API calls for data that doesn't meet your criteria 4. **Compliance**: Maintain regulatory compliance by filtering problematic content -By implementing guardrails at the embedding stage, you create a critical safety layer that protects your entire AI pipeline. For technical teams already building with embeddings, the AI Gateway's guardrails integrate seamlessly with existing workflows while providing the security measures that enterprise applications demand. +By implementing guardrails at the embedding stage, you create a critical safety layer that protects your entire AI pipeline. For technical teams already building with embeddings, the AI Gateway's guardrails integrate seamlessly with existing workflows while providing the security measures that production applications demand. ## How It Works @@ -145,10 +145,6 @@ All guardrail actions on embedding requests are logged in Strata Cloud Manager, -## Get Support - -If you're implementing guardrails for embeddings and need assistance, reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). - ## Learn More - [AI Gateway Guardrails Overview](/aigw/product/guardrails) diff --git a/aigw/product/guardrails/list-of-guardrail-checks.mdx b/aigw/product/guardrails/list-of-guardrail-checks.mdx index 0a5b59ac..98765aea 100644 --- a/aigw/product/guardrails/list-of-guardrail-checks.mdx +++ b/aigw/product/guardrails/list-of-guardrail-checks.mdx @@ -103,7 +103,7 @@ Each Guardrail Check has a specific purpose, it's own parameters, supported hook img="/images/product/Guardrails/Pillar.jpg" > * Scan prompts and responses comprehensively * Detect PII, toxicity, and - injection attacks * Enterprise security and compliance features + injection attacks * Security and compliance features -**Self-hosted MCP Gateway only.** OAuth in SCM currently works only with self-hosted (enterprise) MCP Gateway deployments. It is not available on the cloud-managed gateway. +**Self-hosted MCP Gateway only.** OAuth in SCM currently works only with self-hosted MCP Gateway deployments. It is not available on the cloud-managed gateway. --- @@ -46,9 +46,11 @@ sequenceDiagram Before CAS authentication works for your MCP Gateway: -1. **CIE Directory Sync configured** — Users must be provisioned into workspaces via [CIE Directory Sync](/aigw/product/enterprise-offering/org-management/directory-sync/cie-directory-sync). CAS authenticates users, but CIE is what provisions them into the system. Without CIE sync, authenticated users cannot be resolved. +1. **Gateway version 2.22.0 or later** — On 2.21.0 the consent screen at `GET /oauth//authorize` returns `500` with `{"status":"failure","message":"immutable"}`, so the flow stops after client registration and before the CAS login page. Check the running version with `GET /v1/health`. -2. **Authentication Profile selected** — An Auth Profile must be selected in the CIE Directory Sync configuration. This profile determines which identity provider is used for the CAS login flow. Auth Profiles are managed in the [CIE Authentication Profiles](https://docs.paloaltonetworks.com/identity/cloud-identity-engine/authenticate-users-with-the-cloud-identity-engine) console. +2. **CIE Directory Sync configured** — Users must be provisioned into workspaces via [CIE Directory Sync](/aigw/product/enterprise-offering/org-management/directory-sync/cie-directory-sync). CAS authenticates users, but CIE is what provisions them into the system. Without CIE sync, authenticated users cannot be resolved. + +3. **Authentication Profile selected** — An Auth Profile must be selected in the CIE Directory Sync configuration. This profile determines which identity provider is used for the CAS login flow. Auth Profiles are managed in the [CIE Authentication Profiles](https://docs.paloaltonetworks.com/identity/cloud-identity-engine/authenticate-users-with-the-cloud-identity-engine) console. CAS relies on CIE for user provisioning. If CIE Directory Sync is not configured, or no group-workspace mappings exist, users will authenticate successfully with CAS but fail to resolve — resulting in an authorization error. @@ -330,6 +332,8 @@ The browser redirects back to your MCP client and the connection succeeds. GitHu | Issue | Cause | Resolution | |-------|-------|------------| +| Client reports `Dynamic Client Registration rejected (HTTP 401)` with `Portkey Error: Invalid API Key. Error Code: 03` | The registration request reached the AI Gateway service instead of the MCP Gateway. The MCP Gateway's `/oauth/register` does not require an API key | Fix the MCP Gateway base URL or ingress routing. See [Client registration fails with "Invalid API Key"](/aigw/help-center/mcp-gateway-troubleshooting#client-registration-fails-with-invalid-api-key-error-code-03) | +| `500` `{"status":"failure","message":"immutable"}` from `/oauth//authorize` | Gateway 2.21.0 consent-screen bug | Upgrade the gateway to 2.22.0 or later | | User authenticates but gets "authorization failed" | User not provisioned via CIE | Configure [CIE Directory Sync](/aigw/product/enterprise-offering/org-management/directory-sync/cie-directory-sync) and verify group-workspace mappings | | "Email not found in claims" error | IdP not returning email attribute | Ensure the IdP includes the email claim. Verify the User Identity Attribute (UPN vs Mail) in CIE matches what the IdP returns | | Consent page shows but redirect fails | Browser blocking the redirect | Check for browser extensions or corporate policies blocking redirects to custom URI schemes (`cursor://`, `vscode://`) | diff --git a/aigw/product/mcp-gateway/authentication/jwt.mdx b/aigw/product/mcp-gateway/authentication/jwt.mdx index ef80dfa8..9b5fb7e3 100644 --- a/aigw/product/mcp-gateway/authentication/jwt.mdx +++ b/aigw/product/mcp-gateway/authentication/jwt.mdx @@ -82,7 +82,7 @@ Embed public keys directly in the configuration. Use for self-contained deployme Update keys manually when they rotate. -**Best for:** Air-gapped environments, testing, or when IdP doesn't expose JWKS. +**Best for:** Networks where the gateway can't reach your IdP, testing, or when your IdP doesn't expose JWKS. ### Token Introspection (RFC 7662) diff --git a/aigw/product/mcp-gateway/authentication/oauth-client-metadata.mdx b/aigw/product/mcp-gateway/authentication/oauth-client-metadata.mdx index 7c44b2db..39014b93 100644 --- a/aigw/product/mcp-gateway/authentication/oauth-client-metadata.mdx +++ b/aigw/product/mcp-gateway/authentication/oauth-client-metadata.mdx @@ -142,9 +142,9 @@ Any of these keys in `authorization_params` are ignored and logged as a warning. --- -## Example: Enterprise Compliance +## Example: Compliance Requirements -Your enterprise requires all OAuth registrations to include legal contact information and link to corporate policies: +Your organisation requires all OAuth registrations to include legal contact information and link to corporate policies: ```json { diff --git a/aigw/product/mcp-gateway/internal-mcp-servers.mdx b/aigw/product/mcp-gateway/internal-mcp-servers.mdx index 82e61497..43aab2b8 100644 --- a/aigw/product/mcp-gateway/internal-mcp-servers.mdx +++ b/aigw/product/mcp-gateway/internal-mcp-servers.mdx @@ -1,11 +1,11 @@ --- title: Add Internal MCP Servers -description: Add your internal MCP servers to Prisma AIRS AI Gateway. Get enterprise-grade auth, access control, and logging without building it. +description: Add your internal MCP servers to Prisma AIRS AI Gateway. Get auth, access control, and logging without building it. --- Your team has built MCP servers for internal docs, proprietary APIs, databases. Now you need authentication, access control, and logging for each one. Building that infrastructure is expensive. Maintaining it is harder. -**Add your internal servers to the AI Gateway.** Get enterprise-grade infrastructure without writing a line of auth code. +**Add your internal servers to the AI Gateway.** Get authentication, access control, and logging without writing a line of auth code. --- diff --git a/aigw/product/mcp-gateway/mcp-registry.mdx b/aigw/product/mcp-gateway/mcp-registry.mdx index 9fcd7194..9464325b 100644 --- a/aigw/product/mcp-gateway/mcp-registry.mdx +++ b/aigw/product/mcp-gateway/mcp-registry.mdx @@ -101,7 +101,7 @@ Once added, each server has a detail page with four tabs: **Overview**, **Capabi - Connect your own MCP servers with enterprise auth, access control, and logging. + Connect your own MCP servers with auth, access control, and logging. Setup guides for Claude Desktop, Cursor, Python SDK, TypeScript SDK, and more. diff --git a/aigw/product/mcp-gateway/registry-api.mdx b/aigw/product/mcp-gateway/registry-api.mdx index 29cb7e23..0339dbd7 100644 --- a/aigw/product/mcp-gateway/registry-api.mdx +++ b/aigw/product/mcp-gateway/registry-api.mdx @@ -21,30 +21,30 @@ This is the machine-readable counterpart to the [MCP Registry](/aigw/product/mcp ## Endpoint ```bash -GET https://aigw.portkey.ai/m/v0.1/servers +GET https://mcp-aigw.portkey.ai/v0.1/servers ``` -Authenticate with the `Authorization: Bearer $PORTKEY_API_KEY` header. The key needs the `mcp_servers.list` scope. +Authenticate with the `x-portkey-api-key` header. The key needs the `mcp_servers.list` scope. ```sh cURL -curl https://aigw.portkey.ai/m/v0.1/servers \ - -H "Authorization: Bearer $PORTKEY_API_KEY" +curl https://mcp-aigw.portkey.ai/v0.1/servers \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" ``` ```python Python import requests resp = requests.get( - "https://aigw.portkey.ai/m/v0.1/servers", - headers={"Authorization": f"Bearer {PORTKEY_API_KEY}"} + "https://mcp-aigw.portkey.ai/v0.1/servers", + headers={"x-portkey-api-key": PORTKEY_API_KEY} ) catalog = resp.json() ``` ```js JavaScript -const resp = await fetch("https://aigw.portkey.ai/m/v0.1/servers", { - headers: { "Authorization": `Bearer ${PORTKEY_API_KEY}` } +const resp = await fetch("https://mcp-aigw.portkey.ai/v0.1/servers", { + headers: { "x-portkey-api-key": PORTKEY_API_KEY } }) const catalog = await resp.json() ``` @@ -158,8 +158,8 @@ while True: params["cursor"] = cursor page = requests.get( - "https://aigw.portkey.ai/m/v0.1/servers", - headers={"Authorization": f"Bearer {PORTKEY_API_KEY}"}, + "https://mcp-aigw.portkey.ai/v0.1/servers", + headers={"x-portkey-api-key": PORTKEY_API_KEY}, params=params ).json() diff --git a/aigw/product/model-catalog.mdx b/aigw/product/model-catalog.mdx index e93102dd..fdee020b 100644 --- a/aigw/product/model-catalog.mdx +++ b/aigw/product/model-catalog.mdx @@ -306,7 +306,7 @@ Each custom model gets the same governance controls as standard models. ### Overriding Model Details (Custom Pricing) Override default model pricing for: -- **Negotiated rates**: If you have enterprise agreements with providers +- **Negotiated rates**: If you have custom pricing agreements with providers - **Internal chargebacks**: Set custom rates for internal cost allocation - **Free tier models**: Mark certain models as free for specific teams diff --git a/aigw/product/model-catalog/connect-bedrock-with-amazon-assumed-role.mdx b/aigw/product/model-catalog/connect-bedrock-with-amazon-assumed-role.mdx index 0bb5cf40..4ebf6e68 100644 --- a/aigw/product/model-catalog/connect-bedrock-with-amazon-assumed-role.mdx +++ b/aigw/product/model-catalog/connect-bedrock-with-amazon-assumed-role.mdx @@ -50,7 +50,7 @@ arn:aws:iam::039293892788:role/AirsGwEnterpriseRole The above ARN applies to the SaaS deployment of the AI Gateway, managed through [Strata Cloud Manager](https://stratacloudmanager.paloaltonetworks.com/).
-To enable **Assumed Role for AWS in your AI Gateway Enterprise deployment**, you can refer to [this guide](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/docs/Bedrock.md). If you face any issue, please reach out to [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). +To enable **Assumed Role for AWS in your hybrid deployment**, you can refer to [this guide](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/docs/Bedrock.md).
Paste the following JSON into the trust policy editor and click *Update Trust Policy*. diff --git a/aigw/product/model-catalog/gateway-models.mdx b/aigw/product/model-catalog/gateway-models.mdx index 647ee020..a03c58f2 100644 --- a/aigw/product/model-catalog/gateway-models.mdx +++ b/aigw/product/model-catalog/gateway-models.mdx @@ -640,7 +640,7 @@ The models database powers automatic cost tracking for all requests through the
-If you have negotiated enterprise rates, you can override the default pricing: +If you have negotiated rates with a provider, you can override the default pricing: Set custom input/output costs to match your contracts diff --git a/aigw/product/model-catalog/model-provisioning.mdx b/aigw/product/model-catalog/model-provisioning.mdx index 9d21f33d..58893f59 100644 --- a/aigw/product/model-catalog/model-provisioning.mdx +++ b/aigw/product/model-catalog/model-provisioning.mdx @@ -44,7 +44,7 @@ Each custom model gets the same governance controls as standard models. Per-mode Override default model pricing for: -- **Negotiated rates**: If you have enterprise agreements with providers +- **Negotiated rates**: If you have custom pricing agreements with providers - **Internal chargebacks**: Set custom rates for internal cost allocation - **Free tier models**: Mark certain models as free for specific teams diff --git a/aigw/product/model-catalog/pricing-adjustments.mdx b/aigw/product/model-catalog/pricing-adjustments.mdx index 7dfb247b..6e6598da 100644 --- a/aigw/product/model-catalog/pricing-adjustments.mdx +++ b/aigw/product/model-catalog/pricing-adjustments.mdx @@ -8,12 +8,10 @@ Pricing Adjustments let you apply a discount or markup to an Integration, so cos Requires: - Gateway version **2.7.0** or higher -- Backend version **1.15.0** or higher (for air-gapped deployments) -- Frontend version **1.8.1** or higher (for air-gapped deployments). This is especially useful for: -- **Negotiated Discounts**: Reflect enterprise contracts or committed-use rates from a provider on the corresponding Integration. +- **Negotiated Discounts**: Reflect contract pricing or committed-use rates from a provider on the corresponding Integration. - **Internal Cost Showback**: Apply a markup so the cost reported to internal teams or workspaces includes your platform overhead. - **Custom Per-Integration Rates**: Maintain different effective pricing across multiple Integrations of the same provider (e.g. a discounted production Integration alongside a standard-rate sandbox). diff --git a/aigw/product/observability/auto-instrumentation.mdx b/aigw/product/observability/auto-instrumentation.mdx index ed8bfe12..6ff2c05f 100644 --- a/aigw/product/observability/auto-instrumentation.mdx +++ b/aigw/product/observability/auto-instrumentation.mdx @@ -28,7 +28,3 @@ Strata Cloud Manager can be used to view the logs, traces, and metrics. We currently support auto-instrumentation for the following frameworks: - [CrewAI](/aigw/integrations/agents/crewai#auto-instrumentation) - [LangGraph](/aigw/integrations/agents/langgraph#auto-instrumentation) - - - To request support for another framework, please reach out to us via [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form). - diff --git a/aigw/product/observability/cost-management.mdx b/aigw/product/observability/cost-management.mdx index e9fbd560..1981d3f5 100644 --- a/aigw/product/observability/cost-management.mdx +++ b/aigw/product/observability/cost-management.mdx @@ -3,7 +3,7 @@ title: "Model Pricing and Cost Management" description: "Learn how Prisma AIRS AI Gateway handles pricing data, cost calculations, and pricing updates across different deployment modes." --- -The AI Gateway tracks pricing for all providers and models for which it has pricing support, providing real-time cost visibility and budget management capabilities. The platform automatically handles pricing updates and supports custom pricing configurations for enterprise customers. +The AI Gateway tracks pricing for all providers and models for which it has pricing support, providing real-time cost visibility and budget management capabilities. The platform automatically handles pricing updates and supports custom pricing configurations. ## How Pricing Data Works @@ -26,7 +26,7 @@ For token-based budgets, the AI Gateway tracks both input and output tokens acro ## Deployment-Specific Pricing Management -The AI Gateway supports two deployment modes for new customers, each with different pricing update mechanisms. Air-gapped is documented below for existing deployments only. +The AI Gateway supports two deployment modes, each with different pricing update mechanisms. ### 1. SaaS Deployment (AI Gateway Cloud) @@ -41,17 +41,6 @@ The AI Gateway supports two deployment modes for new customers, each with differ - **Maintenance**: Minimal - pricing JSON updates automatically from central management plane - **Manual Updates**: Only required when new pricing features are introduced -### 3. Air-Gapped Deployment (Legacy) - - -Air-gapped deployment is no longer offered for new customers. This section applies to existing air-gapped deployments, which remain supported. - - -- **Update Frequency**: Configurable — can be automatic (with outbound access) or manual -- **Management**: Customer-managed pricing sources -- **Process**: Pricing data can be fetched from the AI Gateway's hosted config service, loaded from a log store, or mounted as local files. See the [Air-Gapped Model Pricing guide](/self-hosting/airgapped/model-pricing) for setup options -- **Requirements**: Helm chart `app-1.5.0+`, Backend `v1.7.0+`, Enterprise Gateway `v2.0.0+` - ## Model Catalog and Custom Pricing ### Upgrading to Model Catalog @@ -124,7 +113,7 @@ For models without pricing support: **Q: Custom pricing not reflecting in dashboard** -- **A**: Upgrade to Model Catalog for UI-based pricing management. Contact enterprise support for custom pricing configuration. +- **A**: Use the [Model Catalog](/aigw/product/model-catalog) for UI-based pricing management. **Q: Token count and cost show as 0 for streaming requests** @@ -156,21 +145,6 @@ This applies to all OpenAI-compatible providers (including Azure OpenAI). Non-st - **A**: Verify that AI Gateway has pricing support for those models. Unsupported models show \$0.00 in cost tracking. -**Q: Air-gapped deployment has stale pricing** - -- **A**: Configure dynamic pricing fetching via the [Air-Gapped Model Pricing guide](/self-hosting/airgapped/model-pricing). You can pull from the AI Gateway's hosted config service, maintain a local copy from the [AI Gateway Models repo](https://github.com/Portkey-AI/models), or mount pricing files as volumes. - -### Enterprise Support - -For enterprise customers requiring: - -- Custom pricing configurations -- Advanced budget management -- Pricing data exports -- Integration assistance - -Contact: [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) - ## Best Practices ### Cost Optimization @@ -186,9 +160,3 @@ Contact: [Portkey support](https://support.portkey.ai/forms/customer-portal-tick 2. **Regular Reviews**: Monitor pricing changes across providers 3. **Custom Pricing**: Configure preferential rates when available 4. **Budget Planning**: Set realistic budgets based on usage patterns - -## Support and Resources - ---- - -_This documentation covers the AI Gateway's pricing management capabilities. For the most current information about specific model pricing or enterprise features, please [Contact support](https://support.portkey.ai/forms/customer-portal-ticket-form)._ diff --git a/aigw/product/observability/feedback.mdx b/aigw/product/observability/feedback.mdx index 65c4283e..dc685b9a 100644 --- a/aigw/product/observability/feedback.mdx +++ b/aigw/product/observability/feedback.mdx @@ -3,10 +3,6 @@ title: "Feedback" description: "The Prisma AIRS AI Gateway's Feedback APIs provide a simple way to get weighted feedback from customers on any request you served, at any stage in your app." --- - - This feature is available on all AI Gateway plans. - - You can capture this feedback on a request or conversation level and analyze it by adding meta data to the relevant request. ## Adding Feedback to Requests diff --git a/aigw/product/observability/filters.mdx b/aigw/product/observability/filters.mdx index 902b022b..7fee286b 100644 --- a/aigw/product/observability/filters.mdx +++ b/aigw/product/observability/filters.mdx @@ -1,10 +1,6 @@ --- title: "Filters" --- - -This feature is available on all Prisma AIRS AI Gateway plans. - - You can filter analytics & logs by the following parameters: 1. **Model Used**: The AI provider and the model used. diff --git a/aigw/product/observability/logs-export.mdx b/aigw/product/observability/logs-export.mdx index 9f002b6a..fa4d4d33 100644 --- a/aigw/product/observability/logs-export.mdx +++ b/aigw/product/observability/logs-export.mdx @@ -4,7 +4,7 @@ description: "Easily access your Prisma AIRS AI Gateway logs data for further an --- - [Data Service](/aigw/changelog/data-service) must be enabled to use the Logs Export feature for Self Hosted Enterprise customers. + [Data Service](/changelog/data-service) must be enabled to use the Logs Export feature in hybrid deployments. At the AI Gateway, we understand the importance of data analysis and reporting for businesses and teams. That's why we provide a comprehensive logs export feature that allows you to download your AI Gateway logs data in a **structured format**, enabling you to gain valuable insights into your LLM usage, performance, costs, and more. @@ -62,7 +62,7 @@ Admins can block specific fields from being exported. See [Restricting exportabl 7. For completed exports, click the **Download** button to get your logs data file. You can - Currently we only support exporting 50k logs per job. For more help reach out to the Palo Alto Networks team via [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) + Currently we only support exporting 50k logs per job. ## Restricting Exportable Fields diff --git a/aigw/product/observability/logs.mdx b/aigw/product/observability/logs.mdx index 285952ee..41a07af2 100644 --- a/aigw/product/observability/logs.mdx +++ b/aigw/product/observability/logs.mdx @@ -36,22 +36,9 @@ The AI Gateway features—[Cache](/aigw/product/ai-gateway/cache-simple-and-sema As you're viewing logs, you can also add manual feedback on the logs to be analysed and filtered later. This data can be viewed on the [feedback analytics dashboards](/aigw/product/observability/analytics#feedback). -## Configs & Prompt IDs in Logs +## Config IDs in Logs -If your request has an attached [Config](/aigw/product/ai-gateway/configs) or if it's originating from a prompt template, you can see the relevant Config or Prompt IDs separately in the log's details on the AI Gateway. And to dig deeper, you can just click on the IDs and the AI Gateway will take you to the respective Config or Prompt playground where you can view the full details. - -## Debug Requests with Log Replay - -You can rerun any buggy request with just one click, straight from the log details page. The `Replay` button opens your request in a fresh prompt playground where you can rerun the request and edit it right there until it works. - - - - `Replay` **button will be inactive for a log in the following cases:** - -1. If the request is sent to any endpoint other than `/chat/completions,` `/completions`, `/embeddings` -2. If the provider used in the log is archived on the AI Gateway -3. If the request originates from a prompt template which is called from inside a Config target - +If your request has an attached [Config](/aigw/product/ai-gateway/configs), you can see the relevant Config ID separately in the log's details on the AI Gateway. And to dig deeper, you can just click on the ID and the AI Gateway will take you to the Config where you can view the full details. ## DO NOT TRACK diff --git a/aigw/product/observability/metadata.mdx b/aigw/product/observability/metadata.mdx index c1b69e78..beef18aa 100644 --- a/aigw/product/observability/metadata.mdx +++ b/aigw/product/observability/metadata.mdx @@ -3,10 +3,6 @@ title: Metadata description: Add custom context to your AI requests for better observability and analytics --- - -This feature is available on all Prisma AIRS AI Gateway plans. - - ## What is Metadata? Metadata in the AI Gateway allows you to attach custom contextual information to your AI requests. Think of it as tagging your requests with important business context that helps you: @@ -129,9 +125,9 @@ HEADERS_TO_METADATA=x-request-id,x-caller-service,x-environment When configured, the gateway extracts the specified header values from each incoming request and merges them into the request metadata. Header names are matched case-insensitively. -## Enterprise Features +## Metadata Governance -For enterprise users, the AI Gateway offers advanced metadata governance and lets you define metadata at multiple levels: +The AI Gateway offers metadata governance and lets you define metadata at multiple levels: 1. **Request level** - Applied to a single request 2. **API key level** - Applied to all requests using that key diff --git a/aigw/product/observability/opentelemetry.mdx b/aigw/product/observability/opentelemetry.mdx index 14075469..775ecaf5 100644 --- a/aigw/product/observability/opentelemetry.mdx +++ b/aigw/product/observability/opentelemetry.mdx @@ -164,7 +164,7 @@ Navigate to the [Logs page](https://stratacloudmanager.paloaltonetworks.com/) to ## Experimental Features -### Push Logs to an OpenTelemetry compatible endpoint `[Enterprise/Self-Hosted]` +### Push Logs to an OpenTelemetry compatible endpoint `[Hybrid]` [OpenTelemetry conventions on GenAI Traces](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/) are still under development and have not been widely adapted, hence the feature is still experimental. diff --git a/aigw/scripts/check_anchors.py b/aigw/scripts/check_anchors.py index e912461b..e3097c94 100644 --- a/aigw/scripts/check_anchors.py +++ b/aigw/scripts/check_anchors.py @@ -34,10 +34,10 @@ # shrinking baseline: an entry that stops matching anything fails the check, same as the # endpoint check, so these cannot quietly rot once someone resolves them. KNOWN = { - # Inherited. "3. Enterprise Governance" does not exist in the Latest version of these - # pages either, so the links arrived broken; we did not break them. 24 links, 14 pages. - "3-enterprise-governance": "section absent upstream too -- needs a destination", - "enterprise-governance": "section absent upstream too -- needs a destination", + # Not broken in the rendered page. "# 3. Set Up Governance" is the first heading of + # snippets/aigw/portkey-advanced-features.mdx, which every linking page imports, and + # this check reads only the page's own headings, not imported snippets. 23 links, 14 pages. + "3-set-up-governance": "heading comes from an imported snippet this check does not read", # Sections removed by our own phases, with the links into them left behind. "auto-instrumentation": "SDK feature removed in Phase 3", "access-control-management": "section gone from list-of-guardrail-checks", diff --git a/aigw/scripts/nav_retirement_ledger.json b/aigw/scripts/nav_retirement_ledger.json index 57dcfc60..0da521da 100644 --- a/aigw/scripts/nav_retirement_ledger.json +++ b/aigw/scripts/nav_retirement_ledger.json @@ -8,5 +8,9 @@ "aigw/product/enterprise-offering/logs-export": "Alias stub. Retired in Phase 6.4; nav points at product/observability/logs-export, which was already in the tree.", "aigw/product/model-catalog/budget-limits": "Alias stub. Retired in Phase 6.3; content now at product/policies/budget-limits.", "aigw/product/model-catalog/rate-limits": "Alias stub. Retired in Phase 6.3; content now at product/policies/rate-limits.", - "aigw/product/observability/budget-limits": "Alias stub. Retired in Phase 6.3; content now at product/policies/budget-limits." + "aigw/product/observability/budget-limits": "Alias stub. Retired in Phase 6.3; content now at product/policies/budget-limits.", + "aigw/self-hosting/hybrid-deployments/aws/marketplace": "Retired 2026-09-28. Prisma AIRS AI Gateway is not sold through AWS Marketplace; the page documented Portkey's listing.", + "aigw/product/administration/configure-prompt-access-permissions": "Retired 2026-09-28. Prompt Management is not part of Prisma AIRS AI Gateway.", + "aigw/changelog/enterprise": "Retired; Prisma AIRS nav points at the standard changelog/enterprise page instead of a separate aigw/ copy.", + "aigw/changelog/data-service": "Retired; Prisma AIRS nav points at the standard changelog/data-service page instead of a separate aigw/ copy." } diff --git a/aigw/self-hosting/cache-behavior.mdx b/aigw/self-hosting/cache-behavior.mdx index 8240f231..a97660ca 100644 --- a/aigw/self-hosting/cache-behavior.mdx +++ b/aigw/self-hosting/cache-behavior.mdx @@ -10,7 +10,7 @@ The Gateway uses a local cache store (Redis or compatible) for two distinct purp 2. **LLM response cache:** stores LLM request/response pairs for reuse across identical requests -**Hybrid vs air-gapped:** In a hybrid deployment, the Management Plane is hosted by Prisma AIRS AI Gateway. In an air-gapped deployment, the Management Plane runs entirely within your own infrastructure. Air-gapped is no longer offered for new deployments; references to it apply to existing air-gapped customers. +**Hybrid deployments:** In a hybrid deployment, the Management Plane is hosted by Prisma AIRS AI Gateway. --- @@ -106,10 +106,10 @@ In each case, an evicted entry behaves the same as an expired one: the next requ ## Related - + Overview of the Data Plane / Management Plane split and data flow between them - + Supported cache backends: Redis, AWS ElastiCache, and more diff --git a/aigw/self-hosting/hybrid-deployments/architecture.mdx b/aigw/self-hosting/hybrid-deployments/architecture.mdx index d1210a9c..4fda2a3c 100644 --- a/aigw/self-hosting/hybrid-deployments/architecture.mdx +++ b/aigw/self-hosting/hybrid-deployments/architecture.mdx @@ -364,5 +364,5 @@ authentication model and the API surface, not the underlying data architecture. - + diff --git a/aigw/self-hosting/hybrid-deployments/aws/ecs.mdx b/aigw/self-hosting/hybrid-deployments/aws/ecs.mdx index 5d5a18af..92dac64b 100644 --- a/aigw/self-hosting/hybrid-deployments/aws/ecs.mdx +++ b/aigw/self-hosting/hybrid-deployments/aws/ecs.mdx @@ -1,6 +1,6 @@ --- title: "ECS" -description: This enterprise-focused document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software on Amazon Elastic Container Service (ECS), tailored to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. +description: This document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software on Amazon Elastic Container Service (ECS), tailored to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. --- ## Components and Sizing Recommendations diff --git a/aigw/self-hosting/hybrid-deployments/aws/eks.mdx b/aigw/self-hosting/hybrid-deployments/aws/eks.mdx index 6bacf5cc..645d284c 100644 --- a/aigw/self-hosting/hybrid-deployments/aws/eks.mdx +++ b/aigw/self-hosting/hybrid-deployments/aws/eks.mdx @@ -1,6 +1,6 @@ --- title: "EKS" -description: This enterprise-focused document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software in a hybrid mode on Amazon EKS clusters, designed to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. +description: This document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software in a hybrid mode on Amazon EKS clusters, designed to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. --- ## Components and Sizing Recommendations diff --git a/aigw/self-hosting/hybrid-deployments/aws/marketplace.mdx b/aigw/self-hosting/hybrid-deployments/aws/marketplace.mdx deleted file mode 100644 index baa9b3ca..00000000 --- a/aigw/self-hosting/hybrid-deployments/aws/marketplace.mdx +++ /dev/null @@ -1,1632 +0,0 @@ ---- -title: "AWS Marketplace" -description: This enterprise-focused document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software using AWS Marketplace. ---- - - It includes specific steps to subscribe AI Gateway Hybrid Enterprise Edition using AWS Marketplace and deploy the AI Gateway on AWS EKS with Quick Launch. - -## Architecture - -## Components and Sizing Recommendations - -| Component | Options | Sizing Recommendations | -| --------------------------------------- | ------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | -| AI Gateway | Deploy as a Docker container in your Kubernetes cluster using Helm Charts | AWS NodeGroup t4g.medium instance, with at least 4GiB of memory and two vCPUs For high reliability, deploy across multiple Availability Zones. | -| Logs store | AWS S3 | Each log document is \~10kb in size (uncompressed) | -| Cache (Prompts, Configs & Providers) | Elasticache or self-hosted Redis | Deploy in the same VPC as the AI Gateway. | - -## Helm Chart - -This deployment uses the AI Gateway hybrid Helm chart to deploy the AI Gateway. You can find more information about the Helm chart in the [AI Gateway Helm Chart GitHub Repository](https://github.com/Portkey-AI/helm/blob/main/charts/portkey-gateway/README.md). - -## Prerequisites - -1. Create a Strata Cloud Manager account on [the AI Gateway](https://stratacloudmanager.paloaltonetworks.com/) -2. Palo Alto Networks team will share the credentials for the private Docker registry. - -## Markeplace Listing - -### Visit AI Gateway AWS Marketplace Listing - -You can find the AI Gateway AWS Marketplace listing [here](https://aws.amazon.com/marketplace/pp/prodview-o2leb4xcrkdqa). - -### Subscribe to AI Gateway Enterprise Edition - -Subscribe to the AI Gateway Enterprise Edition to gain access to the AI Gateway. - -### Quick Launch - -Upon subscribing to the AI Gateway Enterprise Edition, you will be able to select Quick Launch from within your AWS Console Subscriptions. - -### Launch the Cloud Formation Template - -Select the AI Gateway Enterprise Edition and click on Quick Launch. - -### Run the Cloud Formation Template - -Fill the required parameters and click on Next and run the Cloud Formation Template. - -## Cloud Formation Steps - -- Creates a new EKS cluster and NodeGroup in your selected VPC and Subnets -- Sets up IAM Roles needed for S3 bucket access using STS and Lambda execution -- Uses AWS Lambda to: - - Install the AI Gateway Helm chart to your EKS cluster - - Upload the values.yaml file to the S3 bucket -- Allows for changes to the values file or helm chart deployment by updating and re-running the same Lambda function in your AWS account - -### Cloudformation Template - -> Our cloudformation template has passed the AWS Marketplace validation and security review. - -```yaml portkey-hybrid-eks-cloudformation.template.yaml [expandable] -AWSTemplateFormatVersion: "2010-09-09" -Description: Portkey deployment template for AWS Marketplace - -Metadata: - AWS::CloudFormation::Interface: - ParameterGroups: - - Label: - default: "Required Parameters" - Parameters: - - VPCID - - Subnet1ID - - Subnet2ID - - ClusterName - - NodeGroupName - - NodeGroupInstanceType - - SecurityGroupID - - CreateNewCluster - - HelmChartVersion - - PortkeyDockerUsername: - NoEcho: true - - PortkeyDockerPassword: - NoEcho: true - - PortkeyClientAuth: - NoEcho: true - - Label: - default: "Optional Parameters" - Parameters: - - PortkeyOrgId - - PortkeyGatewayIngressEnabled - - PortkeyGatewayIngressSubdomain - - PortkeyFineTuningEnabled - -Parameters: - # Required Parameters - VPCID: - Type: AWS::EC2::VPC::Id - Description: VPC where the EKS cluster will be created - Default: Select a VPC - - Subnet1ID: - Type: AWS::EC2::Subnet::Id - Description: First subnet ID for EKS cluster - Default: Select your subnet - - Subnet2ID: - Type: AWS::EC2::Subnet::Id - Description: Second subnet ID for EKS cluster - Default: Select your subnet - - # Optional Parameters with defaults - ClusterName: - Type: String - Description: Name of the EKS cluster (if not provided, a new EKS cluster will be created) - Default: portkey-eks-cluster - - NodeGroupName: - Type: String - Description: Name of the EKS node group (if not provided, a new EKS node group will be created) - Default: portkey-eks-cluster-node-group - - NodeGroupInstanceType: - Type: String - Description: EC2 instance type for the node group (if not provided, t3.medium will be used) - Default: t3.medium - AllowedValues: - - t3.medium - - t3.large - - t3.xlarge - - PortkeyDockerUsername: - Type: String - Description: Docker username for Portkey (provided by the Portkey team) - Default: portkeyenterprise - - PortkeyDockerPassword: - Type: String - Description: Docker password for Portkey (provided by the Portkey team) - Default: "" - NoEcho: true - - PortkeyClientAuth: - Type: String - Description: Portkey Client ID (provided by the Portkey team) - Default: "" - NoEcho: true - - PortkeyOrgId: - Type: String - Description: Portkey Organisation ID (provided by the Portkey team) - Default: "" - - HelmChartVersion: - Type: String - Description: Version of the Helm chart to deploy - Default: "latest" - AllowedValues: - - latest - - SecurityGroupID: - Type: String - Description: Optional security group ID for the EKS cluster (if not provided, a new security group will be created) - Default: "" - - CreateNewCluster: - Type: String - AllowedValues: [true, false] - Default: true - Description: Whether to create a new EKS cluster or use an existing one - - PortkeyGatewayIngressEnabled: - Type: String - AllowedValues: [true, false] - Default: false - Description: Whether to enable the Portkey Gateway ingress - - PortkeyGatewayIngressSubdomain: - Type: String - Description: Subdomain for the Portkey Gateway ingress - Default: "" - - PortkeyFineTuningEnabled: - Type: String - AllowedValues: [true, false] - Default: false - Description: Whether to enable the Portkey Fine Tuning - -Conditions: - CreateSecurityGroup: !Equals [!Ref SecurityGroupID, ""] - ShouldCreateCluster: !Equals [!Ref CreateNewCluster, true] - -Resources: - PortkeyAM: - Type: AWS::IAM::Role - DeletionPolicy: Delete - Properties: - RoleName: PortkeyAM - AssumeRolePolicyDocument: - Version: "2012-10-17" - Statement: - - Effect: Allow - Principal: - AWS: !Sub "arn:aws:iam::${AWS::AccountId}:root" - Action: sts:AssumeRole - Policies: - - PolicyName: PortkeyEKSAccess - PolicyDocument: - Version: "2012-10-17" - Statement: - - Effect: Allow - Action: - - "eks:DescribeCluster" - - "eks:ListClusters" - - "eks:ListNodegroups" - - "eks:ListFargateProfiles" - - "eks:ListNodegroups" - - "eks:CreateCluster" - - "eks:CreateNodegroup" - - "eks:DeleteCluster" - - "eks:DeleteNodegroup" - - "eks:UpdateClusterConfig" - - "eks:UpdateKubeconfig" - Resource: !Sub "arn:aws:eks:${AWS::Region}:${AWS::AccountId}:cluster/${ClusterName}" - - Effect: Allow - Action: - - "sts:AssumeRole" - Resource: !Sub "arn:aws:iam::${AWS::AccountId}:role/PortkeyAM" - - Effect: Allow - Action: - - "sts:GetCallerIdentity" - Resource: "*" - - Effect: Allow - Action: - - "iam:ListRoles" - - "iam:GetRole" - Resource: "*" - - Effect: Allow - Action: - - "bedrock:InvokeModel" - - "bedrock:InvokeModelWithResponseStream" - Resource: "*" - - Effect: Allow - Action: - - "s3:GetObject" - - "s3:PutObject" - Resource: - - !Sub "arn:aws:s3:::${AWS::AccountId}-${AWS::Region}-portkey-logs/*" - - PortkeyLogsBucket: - Type: AWS::S3::Bucket - DeletionPolicy: Delete - Properties: - BucketName: !Sub "${AWS::AccountId}-${AWS::Region}-portkey-logs" - VersioningConfiguration: - Status: Enabled - PublicAccessBlockConfiguration: - BlockPublicAcls: true - BlockPublicPolicy: true - IgnorePublicAcls: true - RestrictPublicBuckets: true - BucketEncryption: - ServerSideEncryptionConfiguration: - - ServerSideEncryptionByDefault: - SSEAlgorithm: AES256 - - # EKS Cluster Role - EksClusterRole: - Type: AWS::IAM::Role - DeletionPolicy: Delete - Properties: - RoleName: EksClusterRole-Portkey - AssumeRolePolicyDocument: - Version: "2012-10-17" - Statement: - - Effect: Allow - Principal: - Service: eks.amazonaws.com - Action: sts:AssumeRole - ManagedPolicyArns: - - arn:aws:iam::aws:policy/AmazonEKSClusterPolicy - - # EKS Cluster Security Group (if not provided) - EksSecurityGroup: - Type: AWS::EC2::SecurityGroup - Condition: CreateSecurityGroup - DeletionPolicy: Delete - Properties: - GroupDescription: Security group for Portkey EKS cluster - VpcId: !Ref VPCID - SecurityGroupIngress: - - IpProtocol: tcp - FromPort: 8787 - ToPort: 8787 - CidrIp: PORTKEY_IP - SecurityGroupEgress: - - IpProtocol: tcp - FromPort: 443 - ToPort: 443 - CidrIp: 0.0.0.0/0 - - # EKS Cluster - EksCluster: - Type: AWS::EKS::Cluster - Condition: ShouldCreateCluster - DeletionPolicy: Delete - DependsOn: EksClusterRole - Properties: - Name: !Ref ClusterName - Version: "1.32" - RoleArn: !GetAtt EksClusterRole.Arn - ResourcesVpcConfig: - SecurityGroupIds: - - !If - - CreateSecurityGroup - - !Ref EksSecurityGroup - - !Ref SecurityGroupID - SubnetIds: - - !Ref Subnet1ID - - !Ref Subnet2ID - AccessConfig: - AuthenticationMode: API_AND_CONFIG_MAP - - LambdaExecutionRole: - Type: AWS::IAM::Role - DeletionPolicy: Delete - DependsOn: EksCluster - Properties: - RoleName: PortkeyLambdaRole - AssumeRolePolicyDocument: - Version: "2012-10-17" - Statement: - - Effect: Allow - Principal: - Service: lambda.amazonaws.com - Action: sts:AssumeRole - ManagedPolicyArns: - - arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole - Policies: - - PolicyName: EKSAccess - PolicyDocument: - Version: "2012-10-17" - Statement: - - Effect: Allow - Action: - - ec2:DescribeInstances - - ec2:DescribeRegions - Resource: "*" - - Effect: Allow - Action: - - "sts:AssumeRole" - Resource: !GetAtt PortkeyAM.Arn - - Effect: Allow - Action: - - "s3:GetObject" - - "s3:PutObject" - Resource: - - !Sub "arn:aws:s3:::${AWS::AccountId}-${AWS::Region}-portkey-logs/*" - - Effect: Allow - Action: - - "eks:DescribeCluster" - - "eks:ListClusters" - - "eks:ListNodegroups" - - "eks:ListFargateProfiles" - - "eks:ListNodegroups" - - "eks:CreateCluster" - - "eks:CreateNodegroup" - - "eks:DeleteCluster" - - "eks:DeleteNodegroup" - - "eks:CreateFargateProfile" - - "eks:DeleteFargateProfile" - - "eks:DescribeFargateProfile" - - "eks:UpdateClusterConfig" - - "eks:UpdateKubeconfig" - Resource: !Sub "arn:aws:eks:${AWS::Region}:${AWS::AccountId}:cluster/${ClusterName}" - - LambdaClusterAdmin: - Type: AWS::EKS::AccessEntry - DependsOn: EksCluster - Properties: - ClusterName: !Ref ClusterName - PrincipalArn: !GetAtt LambdaExecutionRole.Arn - Type: STANDARD - KubernetesGroups: - - system:masters - AccessPolicies: - - PolicyArn: "arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy" - AccessScope: - Type: "cluster" - - # Node Group Role - NodeGroupRole: - Type: AWS::IAM::Role - DeletionPolicy: Delete - Properties: - RoleName: NodeGroupRole-Portkey - AssumeRolePolicyDocument: - Version: "2012-10-17" - Statement: - - Effect: Allow - Principal: - Service: ec2.amazonaws.com - Action: sts:AssumeRole - ManagedPolicyArns: - - arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy - - arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy - - arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly - - # EKS Node Group - EksNodeGroup: - Type: AWS::EKS::Nodegroup - DependsOn: EksCluster - DeletionPolicy: Delete - Properties: - CapacityType: ON_DEMAND - ClusterName: !Ref ClusterName - NodegroupName: !Ref NodeGroupName - NodeRole: !GetAtt NodeGroupRole.Arn - InstanceTypes: - - !Ref NodeGroupInstanceType - ScalingConfig: - MinSize: 1 - DesiredSize: 1 - MaxSize: 1 - Subnets: - - !Ref Subnet1ID - - !Ref Subnet2ID - - PortkeyInstallerFunction: - Type: AWS::Lambda::Function - DependsOn: EksNodeGroup - DeletionPolicy: Delete - Properties: - FunctionName: portkey-eks-installer - Runtime: nodejs18.x - Handler: index.handler - MemorySize: 1024 - EphemeralStorage: - Size: 1024 - Code: - ZipFile: | - const fs = require('fs'); - const zlib = require('zlib'); - const { pipeline } = require('stream'); - const path = require('path'); - const https = require('https'); - const { promisify } = require('util'); - const { execSync } = require('child_process'); - const { EKSClient, DescribeClusterCommand } = require('@aws-sdk/client-eks'); - - async function unzipAwsCli(zipPath, destPath) { - // ZIP file format: https://en.wikipedia.org/wiki/ZIP_(file_format) - const data = fs.readFileSync(zipPath); - let offset = 0; - - // Find end of central directory record - const EOCD_SIGNATURE = 0x06054b50; - for (let i = data.length - 22; i >= 0; i--) { - if (data.readUInt32LE(i) === EOCD_SIGNATURE) { - offset = i; - break; - } - } - - // Read central directory info - const numEntries = data.readUInt16LE(offset + 10); - let centralDirOffset = data.readUInt32LE(offset + 16); - - // Process each file - for (let i = 0; i < numEntries; i++) { - // Read central directory header - const signature = data.readUInt32LE(centralDirOffset); - if (signature !== 0x02014b50) { - throw new Error('Invalid central directory header'); - } - - const fileNameLength = data.readUInt16LE(centralDirOffset + 28); - const extraFieldLength = data.readUInt16LE(centralDirOffset + 30); - const fileCommentLength = data.readUInt16LE(centralDirOffset + 32); - const localHeaderOffset = data.readUInt32LE(centralDirOffset + 42); - - // Get filename - const fileName = data.slice( - centralDirOffset + 46, - centralDirOffset + 46 + fileNameLength - ).toString(); - - // Read local file header - const localSignature = data.readUInt32LE(localHeaderOffset); - if (localSignature !== 0x04034b50) { - throw new Error('Invalid local file header'); - } - - const localFileNameLength = data.readUInt16LE(localHeaderOffset + 26); - const localExtraFieldLength = data.readUInt16LE(localHeaderOffset + 28); - - // Get file data - const fileDataOffset = localHeaderOffset + 30 + localFileNameLength + localExtraFieldLength; - const compressedSize = data.readUInt32LE(centralDirOffset + 20); - const uncompressedSize = data.readUInt32LE(centralDirOffset + 24); - const compressionMethod = data.readUInt16LE(centralDirOffset + 10); - - // Create directory if needed - const fullPath = path.join(destPath, fileName); - const directory = path.dirname(fullPath); - if (!fs.existsSync(directory)) { - fs.mkdirSync(directory, { recursive: true }); - } - - // Extract file - if (!fileName.endsWith('/')) { // Skip directories - const fileData = data.slice(fileDataOffset, fileDataOffset + compressedSize); - - if (compressionMethod === 0) { // Stored (no compression) - fs.writeFileSync(fullPath, fileData); - } else if (compressionMethod === 8) { // Deflate - const inflated = require('zlib').inflateRawSync(fileData); - fs.writeFileSync(fullPath, inflated); - } else { - throw new Error(`Unsupported compression method: ${compressionMethod}`); - } - } - - // Move to next entry - centralDirOffset += 46 + fileNameLength + extraFieldLength + fileCommentLength; - } - } - - async function extractTarGz(source, destination) { - // First, let's decompress the .gz file - const gunzip = promisify(zlib.gunzip); - console.log('Reading source file...'); - const compressedData = fs.readFileSync(source); - - console.log('Decompressing...'); - const tarData = await gunzip(compressedData); - - // Now we have the raw tar data - // Tar files are made up of 512-byte blocks - let position = 0; - - while (position < tarData.length) { - // Read header block - const header = tarData.slice(position, position + 512); - position += 512; - - // Get filename from header (first 100 bytes) - const filename = header.slice(0, 100) - .toString('utf8') - .replace(/\0/g, '') - .trim(); - - if (!filename) break; // End of tar - - // Get file size from header (bytes 124-136) - const sizeStr = header.slice(124, 136) - .toString('utf8') - .replace(/\0/g, '') - .trim(); - const size = parseInt(sizeStr, 8); // Size is in octal - - console.log(`Found file: ${filename} (${size} bytes)`); - - if (filename === 'linux-amd64/helm') { - console.log('Found helm binary, extracting...'); - // Extract the file content - const content = tarData.slice(position, position + size); - - // Write to destination - const outputPath = path.join(destination, 'helm'); - fs.writeFileSync(outputPath, content); - console.log(`Helm binary extracted to: ${outputPath}`); - return; // We found what we needed - } - - // Move to next file - position += size; - // Move to next 512-byte boundary - position += (512 - (size % 512)) % 512; - } - - throw new Error('Helm binary not found in archive'); - } - - async function downloadFile(url, dest) { - return new Promise((resolve, reject) => { - const file = fs.createWriteStream(dest); - https.get(url, (response) => { - response.pipe(file); - file.on('finish', () => { - file.close(); - resolve(); - }); - }).on('error', reject); - }); - } - - async function setupBinaries() { - const { STSClient, GetCallerIdentityCommand, AssumeRoleCommand } = require("@aws-sdk/client-sts"); - const { SignatureV4 } = require("@aws-sdk/signature-v4"); - const { defaultProvider } = require("@aws-sdk/credential-provider-node"); - const crypto = require('crypto'); - - const tmpDir = '/tmp/bin'; - if (!fs.existsSync(tmpDir)) { - fs.mkdirSync(tmpDir, { recursive: true }); - } - console.log('Setting up AWS CLI...'); - const awsCliUrl = 'https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip'; - const awsZipPath = `${tmpDir}/awscliv2.zip`; - await unzipAwsCli(awsZipPath, tmpDir); - execSync(`chmod +x ${tmpDir}/aws/install ${tmpDir}/aws/dist/aws`); - execSync(`${tmpDir}/aws/install --update --install-dir /tmp/aws-cli --bin-dir /tmp/aws-bin`, { stdio: 'inherit' }); - - try { - await new Promise((resolve, reject) => { - const https = require('https'); - const fs = require('fs'); - - const file = fs.createWriteStream('/tmp/kubectl'); - - const request = https.get('https://dl.k8s.io/release/v1.32.1/bin/linux/amd64/kubectl', response => { - if (response.statusCode === 302 || response.statusCode === 301) { - https.get(response.headers.location, redirectResponse => { - redirectResponse.pipe(file); - file.on('finish', () => { - file.close(); - resolve(); - }); - }).on('error', err => { - fs.unlink('/tmp/kubectl', () => {}); - reject(err); - }); - return; - } - - response.pipe(file); - file.on('finish', () => { - file.close(); - resolve(); - }); - }); - - request.on('error', err => { - fs.unlink('/tmp/kubectl', () => {}); - reject(err); - }); - }); - - execSync('chmod +x /tmp/kubectl', { - stdio: 'inherit' - }); - } catch (error) { - console.error('Error installing kubectl:', error); - throw error; - } - - console.log('Setting up helm...'); - const helmUrl = 'https://get.helm.sh/helm-v3.12.0-linux-amd64.tar.gz'; - const helmTarPath = `${tmpDir}/helm.tar.gz`; - await downloadFile(helmUrl, helmTarPath); - - await extractTarGz(helmTarPath, tmpDir); - execSync(`chmod +x ${tmpDir}/helm`); - fs.unlinkSync(helmTarPath); - - process.env.PATH = `${tmpDir}:${process.env.PATH}`; - - execSync(`/tmp/aws-bin/aws --version`); - } - - exports.handler = async (event, context) => { - try { - - const { CLUSTER_NAME, NODE_GROUP_NAME, CLUSTER_ARN, CHART_VERSION, - PORTKEY_AWS_REGION, PORTKEY_AWS_ACCOUNT_ID, PORTKEYAM_ROLE_ARN, - PORTKEY_DOCKER_USERNAME, PORTKEY_DOCKER_PASSWORD, - PORTKEY_CLIENT_AUTH, ORGANISATIONS_TO_SYNC } = process.env; - - console.log(process.env) - - if (!CLUSTER_NAME || !PORTKEY_AWS_REGION || !CHART_VERSION || - !PORTKEY_AWS_ACCOUNT_ID || !PORTKEYAM_ROLE_ARN) { - throw new Error('Missing one or more required environment variables.'); - } - - await setupBinaries(); - - const awsCredentialsDir = '/tmp/.aws'; - if (!fs.existsSync(awsCredentialsDir)) { - fs.mkdirSync(awsCredentialsDir, { recursive: true }); - } - - // Write AWS credentials file - const credentialsContent = `[default] - aws_access_key_id = ${process.env.AWS_ACCESS_KEY_ID} - aws_secret_access_key = ${process.env.AWS_SECRET_ACCESS_KEY} - aws_session_token = ${process.env.AWS_SESSION_TOKEN} - region = ${process.env.PORTKEY_AWS_REGION} - `; - - fs.writeFileSync(`${awsCredentialsDir}/credentials`, credentialsContent); - - // Write AWS config file - const configContent = `[default] - region = ${process.env.PORTKEY_AWS_REGION} - output = json - `; - fs.writeFileSync(`${awsCredentialsDir}/config`, configContent); - - // Set AWS config environment variables - process.env.AWS_CONFIG_FILE = `${awsCredentialsDir}/config`; - process.env.AWS_SHARED_CREDENTIALS_FILE = `${awsCredentialsDir}/credentials`; - - // Define kubeconfig path - const kubeconfigDir = `/tmp/${CLUSTER_NAME.trim()}`; - const kubeconfigPath = path.join(kubeconfigDir, 'config'); - - // Create the directory if it doesn't exist - if (!fs.existsSync(kubeconfigDir)) { - fs.mkdirSync(kubeconfigDir, { recursive: true }); - } - - console.log(`Updating kubeconfig for cluster: ${CLUSTER_NAME}`); - execSync(`/tmp/aws-bin/aws eks update-kubeconfig --name ${process.env.CLUSTER_NAME} --region ${process.env.PORTKEY_AWS_REGION} --kubeconfig ${kubeconfigPath}`, { - stdio: 'inherit', - env: { - ...process.env, - HOME: '/tmp', - AWS_CONFIG_FILE: `${awsCredentialsDir}/config`, - AWS_SHARED_CREDENTIALS_FILE: `${awsCredentialsDir}/credentials` - } - }); - - // Set KUBECONFIG environment variable - process.env.KUBECONFIG = kubeconfigPath; - - let kubeconfig = fs.readFileSync(kubeconfigPath, 'utf8'); - - // Replace the command line to use full path - kubeconfig = kubeconfig.replace( - 'command: aws', - 'command: /tmp/aws-bin/aws' - ); - - fs.writeFileSync(kubeconfigPath, kubeconfig); - - // Setup Helm repository - console.log('Setting up Helm repository...'); - await new Promise((resolve, reject) => { - try { - execSync(`helm repo add portkey-ai https://portkey-ai.github.io/helm`, { - stdio: 'inherit', - env: { ...process.env, HOME: '/tmp' } - }); - resolve(); - } catch (error) { - reject(error); - } - }); - - await new Promise((resolve, reject) => { - try { - execSync(`helm repo update`, { - stdio: 'inherit', - env: { ...process.env, HOME: '/tmp' } - }); - resolve(); - } catch (error) { - reject(error); - } - }); - - // Create values.yaml - const valuesYAML = ` - replicaCount: 1 - - images: - gatewayImage: - repository: "docker.io/portkeyai/gateway_enterprise" - pullPolicy: IfNotPresent - tag: "1.9.0" - dataserviceImage: - repository: "docker.io/portkeyai/data-service" - pullPolicy: IfNotPresent - tag: "1.0.2" - - imagePullSecrets: [portkeyenterpriseregistrycredentials] - nameOverride: "" - fullnameOverride: "" - - imageCredentials: - - name: portkeyenterpriseregistrycredentials - create: true - registry: https://index.docker.io/v1/ - username: ${PORTKEY_DOCKER_USERNAME} - password: ${PORTKEY_DOCKER_PASSWORD} - - useVaultInjection: false - - environment: - create: true - secret: true - data: - SERVICE_NAME: portkeyenterprise - PORT: "8787" - LOG_STORE: s3_assume - LOG_STORE_REGION: ${PORTKEY_AWS_REGION} - AWS_ROLE_ARN: ${PORTKEYAM_ROLE_ARN} - LOG_STORE_GENERATIONS_BUCKET: portkey-gateway - ANALYTICS_STORE: control_plane - CACHE_STORE: redis - REDIS_URL: redis://redis:6379 - REDIS_TLS_ENABLED: "false" - PORTKEY_CLIENT_AUTH: ${PORTKEY_CLIENT_AUTH} - ORGANISATIONS_TO_SYNC: ${ORGANISATIONS_TO_SYNC} - - serviceAccount: - create: true - automount: true - annotations: {} - name: "" - - podAnnotations: {} - podLabels: {} - - podSecurityContext: {} - securityContext: {} - - service: - type: LoadBalancer - port: 8787 - targetPort: 8787 - protocol: TCP - additionalLabels: {} - annotations: {} - - ingress: - enabled: ${PORTKEY_GATEWAY_INGRESS_ENABLED} - className: "" - annotations: {} - hosts: - - host: ${PORTKEY_GATEWAY_INGRESS_SUBDOMAIN} - paths: - - path: / - pathType: ImplementationSpecific - tls: [] - - resources: {} - - livenessProbe: - httpGet: - path: /v1/health - port: 8787 - initialDelaySeconds: 30 - periodSeconds: 60 - timeoutSeconds: 5 - failureThreshold: 5 - readinessProbe: - httpGet: - path: /v1/health - port: 8787 - initialDelaySeconds: 30 - periodSeconds: 60 - timeoutSeconds: 5 - successThreshold: 1 - failureThreshold: 5 - - autoscaling: - enabled: true - minReplicas: 1 - maxReplicas: 10 - targetCPUUtilizationPercentage: 80 - - volumes: [] - volumeMounts: [] - nodeSelector: {} - tolerations: [] - affinity: {} - autoRestart: false - - dataservice: - name: "dataservice" - enabled: ${PORTKEY_FINE_TUNING_ENABLED} - containerPort: 8081 - finetuneBucket: ${PORTKEY_AWS_ACCOUNT_ID}-${PORTKEY_AWS_REGION}-portkey-logs - logexportsBucket: ${PORTKEY_AWS_ACCOUNT_ID}-${PORTKEY_AWS_REGION}-portkey-logs - deployment: - autoRestart: true - replicas: 1 - labels: {} - annotations: {} - podSecurityContext: {} - securityContext: {} - resources: {} - startupProbe: - httpGet: - path: /health - port: 8081 - initialDelaySeconds: 60 - failureThreshold: 3 - periodSeconds: 10 - timeoutSeconds: 1 - livenessProbe: - httpGet: - path: /health - port: 8081 - failureThreshold: 3 - periodSeconds: 10 - timeoutSeconds: 1 - readinessProbe: - httpGet: - path: /health - port: 8081 - failureThreshold: 3 - periodSeconds: 10 - timeoutSeconds: 1 - extraContainerConfig: {} - nodeSelector: {} - tolerations: [] - affinity: {} - volumes: [] - volumeMounts: [] - service: - type: ClusterIP - port: 8081 - labels: {} - annotations: {} - loadBalancerSourceRanges: [] - loadBalancerIP: "" - serviceAccount: - create: true - name: "" - labels: {} - annotations: {} - autoscaling: - enabled: false - createHpa: false - minReplicas: 1 - maxReplicas: 5 - targetCPUUtilizationPercentage: 80` - - // Write values.yaml - const valuesYamlPath = '/tmp/values.yaml'; - fs.writeFileSync(valuesYamlPath, valuesYAML); - - const { S3Client, PutObjectCommand, GetObjectCommand } = require("@aws-sdk/client-s3"); - const s3Client = new S3Client({ region: process.env.PORTKEY_AWS_REGION }); - try { - const response = await s3Client.send(new GetObjectCommand({ - Bucket: `${process.env.PORTKEY_AWS_ACCOUNT_ID}-${process.env.PORTKEY_AWS_REGION}-portkey-logs`, - Key: 'values.yaml' - })); - const existingValuesYAML = await response.Body.transformToString(); - console.log('Found existing values.yaml in S3, using it instead of default'); - fs.writeFileSync(valuesYamlPath, existingValuesYAML); - } catch (error) { - if (error.name === 'NoSuchKey') { - // Upload the default values.yaml to S3 - await s3Client.send(new PutObjectCommand({ - Bucket: `${process.env.PORTKEY_AWS_ACCOUNT_ID}-${process.env.PORTKEY_AWS_REGION}-portkey-logs`, - Key: 'values.yaml', - Body: valuesYAML, - ContentType: 'text/yaml' - })); - console.log('Default values.yaml written to S3 bucket'); - } else { - throw error; - } - } - - // Install/upgrade Helm chart - console.log('Installing helm chart...'); - await new Promise((resolve, reject) => { - try { - execSync(`helm upgrade --install portkey-ai portkey-ai/gateway -f ${valuesYamlPath} -n portkeyai --create-namespace --kube-context ${process.env.CLUSTER_ARN} --kubeconfig ${kubeconfigPath}`, { - stdio: 'inherit', - env: { - ...process.env, - HOME: '/tmp', - PATH: `/tmp/aws-bin:${process.env.PATH}` - } - }); - resolve(); - } catch (error) { - reject(error); - } - }); - - return { - statusCode: 200, - body: JSON.stringify({ - message: 'EKS installation and helm chart deployment completed successfully', - event: event - }) - }; - } catch (error) { - console.error('Error:', error); - return { - statusCode: 500, - body: JSON.stringify({ - message: 'Error during EKS installation and helm chart deployment', - error: error.message - }) - }; - } - }; - Role: !GetAtt LambdaExecutionRole.Arn - Timeout: 900 - Environment: - Variables: - CLUSTER_NAME: !Ref ClusterName - NODE_GROUP_NAME: !Ref NodeGroupName - CLUSTER_ARN: !GetAtt EksCluster.Arn - CHART_VERSION: !Ref HelmChartVersion - PORTKEY_AWS_REGION: !Ref "AWS::Region" - PORTKEY_AWS_ACCOUNT_ID: !Ref "AWS::AccountId" - PORTKEYAM_ROLE_ARN: !GetAtt PortkeyAM.Arn - PORTKEY_DOCKER_USERNAME: !Ref PortkeyDockerUsername - PORTKEY_DOCKER_PASSWORD: !Ref PortkeyDockerPassword - PORTKEY_CLIENT_AUTH: !Ref PortkeyClientAuth - ORGANISATIONS_TO_SYNC: !Ref PortkeyOrgId - PORTKEY_GATEWAY_INGRESS_ENABLED: !Ref PortkeyGatewayIngressEnabled - PORTKEY_GATEWAY_INGRESS_SUBDOMAIN: !Ref PortkeyGatewayIngressSubdomain - PORTKEY_FINE_TUNING_ENABLED: !Ref PortkeyFineTuningEnabled -``` - -### Lambda Function - -#### Steps - -1. **Sets up required binaries** - Downloads and configures AWS CLI, kubectl, and Helm binaries in the Lambda environment to enable interaction with AWS services and Kubernetes. -2. **Configures AWS credentials** - Creates temporary AWS credential files in the Lambda environment to authenticate with AWS services. -3. **Connects to EKS cluster** - Updates the kubeconfig file to establish a connection with the specified Amazon EKS cluster. -4. **Manages Helm chart deployment** - Adds the AI Gateway Helm repository and deploys/upgrades the AI Gateway using Helm charts. -5. **Handles configuration values** - Creates a values.yaml file with environment-specific configurations and stores it in an S3 bucket for future reference or updates. -6. **Provides idempotent deployment** - Checks for existing configurations in S3 and uses them if available, allowing the function to be run multiple times for updates without losing custom configurations. - -```javascript portkey-hybrid-eks-cloudformation.lambda.js [expandable] - const fs = require('fs'); - const zlib = require('zlib'); - const { pipeline } = require('stream'); - const path = require('path'); - const https = require('https'); - const { promisify } = require('util'); - const { execSync } = require('child_process'); - const { EKSClient, DescribeClusterCommand } = require('@aws-sdk/client-eks'); - - async function unzipAwsCli(zipPath, destPath) { - // ZIP file format: https://en.wikipedia.org/wiki/ZIP_(file_format) - const data = fs.readFileSync(zipPath); - let offset = 0; - - // Find end of central directory record - const EOCD_SIGNATURE = 0x06054b50; - for (let i = data.length - 22; i >= 0; i--) { - if (data.readUInt32LE(i) === EOCD_SIGNATURE) { - offset = i; - break; - } - } - - // Read central directory info - const numEntries = data.readUInt16LE(offset + 10); - let centralDirOffset = data.readUInt32LE(offset + 16); - - // Process each file - for (let i = 0; i < numEntries; i++) { - // Read central directory header - const signature = data.readUInt32LE(centralDirOffset); - if (signature !== 0x02014b50) { - throw new Error('Invalid central directory header'); - } - - const fileNameLength = data.readUInt16LE(centralDirOffset + 28); - const extraFieldLength = data.readUInt16LE(centralDirOffset + 30); - const fileCommentLength = data.readUInt16LE(centralDirOffset + 32); - const localHeaderOffset = data.readUInt32LE(centralDirOffset + 42); - - // Get filename - const fileName = data.slice( - centralDirOffset + 46, - centralDirOffset + 46 + fileNameLength - ).toString(); - - // Read local file header - const localSignature = data.readUInt32LE(localHeaderOffset); - if (localSignature !== 0x04034b50) { - throw new Error('Invalid local file header'); - } - - const localFileNameLength = data.readUInt16LE(localHeaderOffset + 26); - const localExtraFieldLength = data.readUInt16LE(localHeaderOffset + 28); - - // Get file data - const fileDataOffset = localHeaderOffset + 30 + localFileNameLength + localExtraFieldLength; - const compressedSize = data.readUInt32LE(centralDirOffset + 20); - const uncompressedSize = data.readUInt32LE(centralDirOffset + 24); - const compressionMethod = data.readUInt16LE(centralDirOffset + 10); - - // Create directory if needed - const fullPath = path.join(destPath, fileName); - const directory = path.dirname(fullPath); - if (!fs.existsSync(directory)) { - fs.mkdirSync(directory, { recursive: true }); - } - - // Extract file - if (!fileName.endsWith('/')) { // Skip directories - const fileData = data.slice(fileDataOffset, fileDataOffset + compressedSize); - - if (compressionMethod === 0) { // Stored (no compression) - fs.writeFileSync(fullPath, fileData); - } else if (compressionMethod === 8) { // Deflate - const inflated = require('zlib').inflateRawSync(fileData); - fs.writeFileSync(fullPath, inflated); - } else { - throw new Error(`Unsupported compression method: ${compressionMethod}`); - } - } - - // Move to next entry - centralDirOffset += 46 + fileNameLength + extraFieldLength + fileCommentLength; - } - } - - async function extractTarGz(source, destination) { - // First, let's decompress the .gz file - const gunzip = promisify(zlib.gunzip); - console.log('Reading source file...'); - const compressedData = fs.readFileSync(source); - - console.log('Decompressing...'); - const tarData = await gunzip(compressedData); - - // Now we have the raw tar data - // Tar files are made up of 512-byte blocks - let position = 0; - - while (position < tarData.length) { - // Read header block - const header = tarData.slice(position, position + 512); - position += 512; - - // Get filename from header (first 100 bytes) - const filename = header.slice(0, 100) - .toString('utf8') - .replace(/\0/g, '') - .trim(); - - if (!filename) break; // End of tar - - // Get file size from header (bytes 124-136) - const sizeStr = header.slice(124, 136) - .toString('utf8') - .replace(/\0/g, '') - .trim(); - const size = parseInt(sizeStr, 8); // Size is in octal - - console.log(`Found file: ${filename} (${size} bytes)`); - - if (filename === 'linux-amd64/helm') { - console.log('Found helm binary, extracting...'); - // Extract the file content - const content = tarData.slice(position, position + size); - - // Write to destination - const outputPath = path.join(destination, 'helm'); - fs.writeFileSync(outputPath, content); - console.log(`Helm binary extracted to: ${outputPath}`); - return; // We found what we needed - } - - // Move to next file - position += size; - // Move to next 512-byte boundary - position += (512 - (size % 512)) % 512; - } - - throw new Error('Helm binary not found in archive'); - } - - async function downloadFile(url, dest) { - return new Promise((resolve, reject) => { - const file = fs.createWriteStream(dest); - https.get(url, (response) => { - response.pipe(file); - file.on('finish', () => { - file.close(); - resolve(); - }); - }).on('error', reject); - }); - } - - async function setupBinaries() { - const { STSClient, GetCallerIdentityCommand, AssumeRoleCommand } = require("@aws-sdk/client-sts"); - const { SignatureV4 } = require("@aws-sdk/signature-v4"); - const { defaultProvider } = require("@aws-sdk/credential-provider-node"); - const crypto = require('crypto'); - - const tmpDir = '/tmp/bin'; - if (!fs.existsSync(tmpDir)) { - fs.mkdirSync(tmpDir, { recursive: true }); - } - - // Download and setup AWS CLI - console.log('Setting up AWS CLI...'); - const awsCliUrl = 'https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip'; - const awsZipPath = `${tmpDir}/awscliv2.zip`; - await downloadFile(awsCliUrl, awsZipPath); - // Extract using our custom unzip function - await unzipAwsCli(awsZipPath, tmpDir); - execSync(`chmod +x ${tmpDir}/aws/install ${tmpDir}/aws/dist/aws`); - // Install AWS CLI - execSync(`${tmpDir}/aws/install --update --install-dir /tmp/aws-cli --bin-dir /tmp/aws-bin`, { stdio: 'inherit' }); - - // Download and setup kubectl - try { - // Download kubectl binary using Node.js https - await new Promise((resolve, reject) => { - const https = require('https'); - const fs = require('fs'); - - const file = fs.createWriteStream('/tmp/kubectl'); - - const request = https.get('https://dl.k8s.io/release/v1.32.1/bin/linux/amd64/kubectl', response => { - if (response.statusCode === 302 || response.statusCode === 301) { - https.get(response.headers.location, redirectResponse => { - redirectResponse.pipe(file); - file.on('finish', () => { - file.close(); - resolve(); - }); - }).on('error', err => { - fs.unlink('/tmp/kubectl', () => {}); - reject(err); - }); - return; - } - - response.pipe(file); - file.on('finish', () => { - file.close(); - resolve(); - }); - }); - - request.on('error', err => { - fs.unlink('/tmp/kubectl', () => {}); - reject(err); - }); - }); - - execSync('chmod +x /tmp/kubectl', { - stdio: 'inherit' - }); - } catch (error) { - console.error('Error installing kubectl:', error); - throw error; - } - - console.log('Setting up helm...'); - const helmUrl = 'https://get.helm.sh/helm-v3.12.0-linux-amd64.tar.gz'; - const helmTarPath = `${tmpDir}/helm.tar.gz`; - await downloadFile(helmUrl, helmTarPath); - - await extractTarGz(helmTarPath, tmpDir); - execSync(`chmod +x ${tmpDir}/helm`); - fs.unlinkSync(helmTarPath); - - process.env.PATH = `${tmpDir}:${process.env.PATH}`; - - execSync(`/tmp/aws-bin/aws --version`); - } - - exports.handler = async (event, context) => { - try { - - const { CLUSTER_NAME, NODE_GROUP_NAME, CLUSTER_ARN, CHART_VERSION, - PORTKEY_AWS_REGION, PORTKEY_AWS_ACCOUNT_ID, PORTKEYAM_ROLE_ARN, - PORTKEY_DOCKER_USERNAME, PORTKEY_DOCKER_PASSWORD, - PORTKEY_CLIENT_AUTH, ORGANISATIONS_TO_SYNC } = process.env; - - console.log(process.env) - - if (!CLUSTER_NAME || !PORTKEY_AWS_REGION || !CHART_VERSION || - !PORTKEY_AWS_ACCOUNT_ID || !PORTKEYAM_ROLE_ARN) { - throw new Error('Missing one or more required environment variables.'); - } - - await setupBinaries(); - - const awsCredentialsDir = '/tmp/.aws'; - if (!fs.existsSync(awsCredentialsDir)) { - fs.mkdirSync(awsCredentialsDir, { recursive: true }); - } - - // Write AWS credentials file - const credentialsContent = `[default] - aws_access_key_id = ${process.env.AWS_ACCESS_KEY_ID} - aws_secret_access_key = ${process.env.AWS_SECRET_ACCESS_KEY} - aws_session_token = ${process.env.AWS_SESSION_TOKEN} - region = ${process.env.PORTKEY_AWS_REGION} - `; - - fs.writeFileSync(`${awsCredentialsDir}/credentials`, credentialsContent); - - // Write AWS config file - const configContent = `[default] - region = ${process.env.PORTKEY_AWS_REGION} - output = json - `; - fs.writeFileSync(`${awsCredentialsDir}/config`, configContent); - - // Set AWS config environment variables - process.env.AWS_CONFIG_FILE = `${awsCredentialsDir}/config`; - process.env.AWS_SHARED_CREDENTIALS_FILE = `${awsCredentialsDir}/credentials`; - - // Define kubeconfig path - const kubeconfigDir = `/tmp/${CLUSTER_NAME.trim()}`; - const kubeconfigPath = path.join(kubeconfigDir, 'config'); - - // Create the directory if it doesn't exist - if (!fs.existsSync(kubeconfigDir)) { - fs.mkdirSync(kubeconfigDir, { recursive: true }); - } - - console.log(`Updating kubeconfig for cluster: ${CLUSTER_NAME}`); - execSync(`/tmp/aws-bin/aws eks update-kubeconfig --name ${process.env.CLUSTER_NAME} --region ${process.env.PORTKEY_AWS_REGION} --kubeconfig ${kubeconfigPath}`, { - stdio: 'inherit', - env: { - ...process.env, - HOME: '/tmp', - AWS_CONFIG_FILE: `${awsCredentialsDir}/config`, - AWS_SHARED_CREDENTIALS_FILE: `${awsCredentialsDir}/credentials` - } - }); - - // Set KUBECONFIG environment variable - process.env.KUBECONFIG = kubeconfigPath; - - let kubeconfig = fs.readFileSync(kubeconfigPath, 'utf8'); - - // Replace the command line to use full path - kubeconfig = kubeconfig.replace( - 'command: aws', - 'command: /tmp/aws-bin/aws' - ); - - fs.writeFileSync(kubeconfigPath, kubeconfig); - - // Setup Helm repository - console.log('Setting up Helm repository...'); - await new Promise((resolve, reject) => { - try { - execSync(`helm repo add portkey-ai https://portkey-ai.github.io/helm`, { - stdio: 'inherit', - env: { ...process.env, HOME: '/tmp' } - }); - resolve(); - } catch (error) { - reject(error); - } - }); - - await new Promise((resolve, reject) => { - try { - execSync(`helm repo update`, { - stdio: 'inherit', - env: { ...process.env, HOME: '/tmp' } - }); - resolve(); - } catch (error) { - reject(error); - } - }); - - // Create values.yaml - const valuesYAML = ` - replicaCount: 1 - - images: - gatewayImage: - repository: "docker.io/portkeyai/gateway_enterprise" - pullPolicy: IfNotPresent - tag: "1.9.0" - dataserviceImage: - repository: "docker.io/portkeyai/data-service" - pullPolicy: IfNotPresent - tag: "1.0.2" - - imagePullSecrets: [portkeyenterpriseregistrycredentials] - nameOverride: "" - fullnameOverride: "" - - imageCredentials: - - name: portkeyenterpriseregistrycredentials - create: true - registry: https://index.docker.io/v1/ - username: ${PORTKEY_DOCKER_USERNAME} - password: ${PORTKEY_DOCKER_PASSWORD} - - useVaultInjection: false - - environment: - create: true - secret: true - data: - SERVICE_NAME: portkeyenterprise - PORT: "8787" - LOG_STORE: s3_assume - LOG_STORE_REGION: ${PORTKEY_AWS_REGION} - AWS_ROLE_ARN: ${PORTKEYAM_ROLE_ARN} - LOG_STORE_GENERATIONS_BUCKET: portkey-gateway - ANALYTICS_STORE: control_plane - CACHE_STORE: redis - REDIS_URL: redis://redis:6379 - REDIS_TLS_ENABLED: "false" - PORTKEY_CLIENT_AUTH: ${PORTKEY_CLIENT_AUTH} - ORGANISATIONS_TO_SYNC: ${ORGANISATIONS_TO_SYNC} - - serviceAccount: - create: true - automount: true - annotations: {} - name: "" - - podAnnotations: {} - podLabels: {} - - podSecurityContext: {} - securityContext: {} - - service: - type: LoadBalancer - port: 8787 - targetPort: 8787 - protocol: TCP - additionalLabels: {} - annotations: {} - - ingress: - enabled: ${PORTKEY_GATEWAY_INGRESS_ENABLED} - className: "" - annotations: {} - hosts: - - host: ${PORTKEY_GATEWAY_INGRESS_SUBDOMAIN} - paths: - - path: / - pathType: ImplementationSpecific - tls: [] - - resources: {} - - livenessProbe: - httpGet: - path: /v1/health - port: 8787 - initialDelaySeconds: 30 - periodSeconds: 60 - timeoutSeconds: 5 - failureThreshold: 5 - readinessProbe: - httpGet: - path: /v1/health - port: 8787 - initialDelaySeconds: 30 - periodSeconds: 60 - timeoutSeconds: 5 - successThreshold: 1 - failureThreshold: 5 - - autoscaling: - enabled: true - minReplicas: 1 - maxReplicas: 10 - targetCPUUtilizationPercentage: 80 - - volumes: [] - volumeMounts: [] - nodeSelector: {} - tolerations: [] - affinity: {} - autoRestart: false - - dataservice: - name: "dataservice" - enabled: ${PORTKEY_FINE_TUNING_ENABLED} - containerPort: 8081 - finetuneBucket: ${PORTKEY_AWS_ACCOUNT_ID}-${PORTKEY_AWS_REGION}-portkey-logs - logexportsBucket: ${PORTKEY_AWS_ACCOUNT_ID}-${PORTKEY_AWS_REGION}-portkey-logs - deployment: - autoRestart: true - replicas: 1 - labels: {} - annotations: {} - podSecurityContext: {} - securityContext: {} - resources: {} - startupProbe: - httpGet: - path: /health - port: 8081 - initialDelaySeconds: 60 - failureThreshold: 3 - periodSeconds: 10 - timeoutSeconds: 1 - livenessProbe: - httpGet: - path: /health - port: 8081 - failureThreshold: 3 - periodSeconds: 10 - timeoutSeconds: 1 - readinessProbe: - httpGet: - path: /health - port: 8081 - failureThreshold: 3 - periodSeconds: 10 - timeoutSeconds: 1 - extraContainerConfig: {} - nodeSelector: {} - tolerations: [] - affinity: {} - volumes: [] - volumeMounts: [] - service: - type: ClusterIP - port: 8081 - labels: {} - annotations: {} - loadBalancerSourceRanges: [] - loadBalancerIP: "" - serviceAccount: - create: true - name: "" - labels: {} - annotations: {} - autoscaling: - enabled: false - createHpa: false - minReplicas: 1 - maxReplicas: 5 - targetCPUUtilizationPercentage: 80` - - // Write values.yaml - const valuesYamlPath = '/tmp/values.yaml'; - fs.writeFileSync(valuesYamlPath, valuesYAML); - - const { S3Client, PutObjectCommand, GetObjectCommand } = require("@aws-sdk/client-s3"); - const s3Client = new S3Client({ region: process.env.PORTKEY_AWS_REGION }); - try { - const response = await s3Client.send(new GetObjectCommand({ - Bucket: `${process.env.PORTKEY_AWS_ACCOUNT_ID}-${process.env.PORTKEY_AWS_REGION}-portkey-logs`, - Key: 'values.yaml' - })); - const existingValuesYAML = await response.Body.transformToString(); - console.log('Found existing values.yaml in S3, using it instead of default'); - fs.writeFileSync(valuesYamlPath, existingValuesYAML); - } catch (error) { - if (error.name === 'NoSuchKey') { - // Upload the default values.yaml to S3 - await s3Client.send(new PutObjectCommand({ - Bucket: `${process.env.PORTKEY_AWS_ACCOUNT_ID}-${process.env.PORTKEY_AWS_REGION}-portkey-logs`, - Key: 'values.yaml', - Body: valuesYAML, - ContentType: 'text/yaml' - })); - console.log('Default values.yaml written to S3 bucket'); - } else { - throw error; - } - } - - // Install/upgrade Helm chart - console.log('Installing helm chart...'); - await new Promise((resolve, reject) => { - try { - execSync(`helm upgrade --install portkey-ai portkey-ai/gateway -f ${valuesYamlPath} -n portkeyai --create-namespace --kube-context ${process.env.CLUSTER_ARN} --kubeconfig ${kubeconfigPath}`, { - stdio: 'inherit', - env: { - ...process.env, - HOME: '/tmp', - PATH: `/tmp/aws-bin:${process.env.PATH}` - } - }); - resolve(); - } catch (error) { - reject(error); - } - }); - - return { - statusCode: 200, - body: JSON.stringify({ - message: 'EKS installation and helm chart deployment completed successfully', - event: event - }) - }; - } catch (error) { - console.error('Error:', error); - return { - statusCode: 500, - body: JSON.stringify({ - message: 'Error during EKS installation and helm chart deployment', - error: error.message - }) - }; - } - }; -``` - -### Post Deployment Verification - -#### Verify AI Gateway Deployment - -```bash -kubectl get all -n portkeyai -``` - -#### Verify AI Gateway Endpoint - -```bash -export POD_NAME=$(kubectl get pods -n portkeyai -l app.kubernetes.io/name=gateway -o jsonpath="{.items[0].metadata.name}") -kubectl port-forward $POD_NAME 8787:8787 -n portkeyai -``` - -Visiting localhost:8787/v1/health will return `Server is healthy` - -Your AI Gateway is now ready to use! diff --git a/aigw/self-hosting/hybrid-deployments/azure/aca.mdx b/aigw/self-hosting/hybrid-deployments/azure/aca.mdx index 6cb1788e..869d2fc6 100644 --- a/aigw/self-hosting/hybrid-deployments/azure/aca.mdx +++ b/aigw/self-hosting/hybrid-deployments/azure/aca.mdx @@ -1,6 +1,6 @@ --- title: "ACA" -description: This enterprise-focused document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software on Azure Container Apps (ACA), tailored to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. +description: This document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software on Azure Container Apps (ACA), tailored to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. --- ## Components and Sizing Recommendations diff --git a/aigw/self-hosting/hybrid-deployments/azure/aks.mdx b/aigw/self-hosting/hybrid-deployments/azure/aks.mdx index 577ff11b..5ceaf4f0 100644 --- a/aigw/self-hosting/hybrid-deployments/azure/aks.mdx +++ b/aigw/self-hosting/hybrid-deployments/azure/aks.mdx @@ -1,14 +1,8 @@ --- title: "AKS" -description: This enterprise-focused document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software on Azure Kubernetes Service (AKS), tailored to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. +description: This document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software on Azure Kubernetes Service (AKS), tailored to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. --- - - The AI Gateway is also available on the Azure Marketplace. You can deploy the AI Gateway directly through your Azure console, which streamlines procurement and deployment processes. - - [Deploy via Azure Marketplace →](https://azuremarketplace.microsoft.com/en-in/marketplace/apps/portkey.enterprise-saas?tab=Overview) - - ## Components and Sizing Recommendations | Component | Options | Sizing Recommendations | diff --git a/aigw/self-hosting/hybrid-deployments/gateway-registration.mdx b/aigw/self-hosting/hybrid-deployments/gateway-registration.mdx index 8b814d45..85ae3530 100644 --- a/aigw/self-hosting/hybrid-deployments/gateway-registration.mdx +++ b/aigw/self-hosting/hybrid-deployments/gateway-registration.mdx @@ -4,10 +4,6 @@ sidebarTitle: "Gateway Registration" description: "Register your self-hosted AI Gateway with Prisma AIRS AI Gateway to enable configuration sync, analytics, and workspace access control." --- - - This feature is available for select Enterprise customers only. - - ## Overview Gateway Registration connects your self-hosted AI Gateway to the Management Plane. Once registered, the Data Plane can pull promt templates, routing configs,integrations, API keys etc. diff --git a/aigw/self-hosting/hybrid-deployments/gcp.mdx b/aigw/self-hosting/hybrid-deployments/gcp.mdx index 4196b59d..4eb1167b 100644 --- a/aigw/self-hosting/hybrid-deployments/gcp.mdx +++ b/aigw/self-hosting/hybrid-deployments/gcp.mdx @@ -1,6 +1,6 @@ --- title: "GCP" -description: This enterprise-focused document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software in a hybrid mode on Google Kubernetes Engine clusters, designed to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. +description: This document provides comprehensive instructions for deploying Prisma AIRS AI Gateway software in a hybrid mode on Google Kubernetes Engine clusters, designed to meet the needs of large-scale, mission-critical applications. It includes specific recommendations for component sizing, high availability, and integration with monitoring systems. --- ## Components and Sizing Recommendations diff --git a/aigw/self-hosting/private-network-access.mdx b/aigw/self-hosting/private-network-access.mdx index 2a565884..a4fa6ebb 100644 --- a/aigw/self-hosting/private-network-access.mdx +++ b/aigw/self-hosting/private-network-access.mdx @@ -13,7 +13,7 @@ You need this when the gateway calls any of these on a private address: - A [custom guardrail webhook](/aigw/integrations/guardrails/bring-your-own-guardrails) -`TRUSTED_CUSTOM_HOSTS` is available only on hybrid and air-gapped deployments. On AI Gateway SaaS, upstream URLs must be publicly reachable. +`TRUSTED_CUSTOM_HOSTS` is available only on hybrid deployments. On AI Gateway SaaS, upstream URLs must be publicly reachable. ## Recognize the error diff --git a/aigw/self-hosting/prometheus-metrics.mdx b/aigw/self-hosting/prometheus-metrics.mdx index 1743bb19..95f81cc5 100644 --- a/aigw/self-hosting/prometheus-metrics.mdx +++ b/aigw/self-hosting/prometheus-metrics.mdx @@ -1,11 +1,11 @@ --- title: "Prometheus Metrics" -description: "Comprehensive monitoring and observability for Prisma AIRS AI Gateway Enterprise Gateway through Prometheus metrics" +description: "Comprehensive monitoring and observability for Prisma AIRS AI Gateway through Prometheus metrics" --- ## Overview -The AI Gateway Enterprise Gateway exposes detailed telemetry data through Prometheus metrics, enabling comprehensive observability for LLM gateway operations. These metrics cover the entire request lifecycle from authentication through response delivery, including cost tracking, performance monitoring, and cache analytics. +The AI Gateway exposes detailed telemetry data through Prometheus metrics, enabling comprehensive observability for LLM gateway operations. These metrics cover the entire request lifecycle from authentication through response delivery, including cost tracking, performance monitoring, and cache analytics. This monitoring capability is essential for: - **Performance Optimization**: Identify bottlenecks and optimize gateway performance @@ -57,7 +57,7 @@ These buckets are specifically tuned for applications handling variable-length L ## Custom Application Metrics -The AI Gateway Enterprise Gateway exposes 15 custom metrics designed to provide deep visibility into LLM gateway operations, performance characteristics, and business metrics. +The AI Gateway exposes 15 custom metrics designed to provide deep visibility into LLM gateway operations, performance characteristics, and business metrics. ### Universal Label Schema @@ -497,7 +497,7 @@ Measures the performance of converting incoming gRPC requests to HTTP format bef **Default**: `true` (enabled) **Values**: `true` | `false` -The AI Gateway Enterprise Gateway allows you to completely disable Prometheus metrics collection if not needed for your deployment. When disabled, both the metrics middleware and the `/metrics` endpoint are deactivated, reducing overhead in environments where Prometheus monitoring is not required. +The AI Gateway allows you to completely disable Prometheus metrics collection if not needed for your deployment. When disabled, both the metrics middleware and the `/metrics` endpoint are deactivated, reducing overhead in environments where Prometheus monitoring is not required. **Configuration**: ```bash @@ -566,7 +566,7 @@ This exports: ### Dynamic Metadata Label System -The AI Gateway Enterprise Gateway supports dynamic metadata labelling through request-specific metadata injection. This powerful feature enables fine-grained observability across custom dimensions specific to your organisation's structure and use cases. +The AI Gateway supports dynamic metadata labelling through request-specific metadata injection. This powerful feature enables fine-grained observability across custom dimensions specific to your organisation's structure and use cases. **Metadata labels are disabled by default** to prevent cardinality issues. Enable them with `PROMETHEUS_INCLUDE_METADATA_LABELS`, then restrict which keys are promoted to labels using `PROMETHEUS_LABELS_METADATA_ALLOWED_KEYS`. @@ -876,5 +876,5 @@ When deploying metrics collection in production: ## Related Documentation - [Analytics Dashboard](/aigw/product/observability/analytics) - SaaS monitoring and analytics -- [Private Cloud Architecture](/product/enterprise-offering/private-cloud-deployments/architecture) - Deployment architecture overview +- [Architecture](/aigw/self-hosting/hybrid-deployments/architecture) - Deployment architecture overview - [Observability](/aigw/product/observability) - General observability features diff --git a/api-reference/inference-api/supported-providers.mdx b/api-reference/inference-api/supported-providers.mdx index 9bebd3a2..8b9140b9 100644 --- a/api-reference/inference-api/supported-providers.mdx +++ b/api-reference/inference-api/supported-providers.mdx @@ -1,6 +1,5 @@ --- title: "Supported Providers" -mode: "wide" --- diff --git a/docs.json b/docs.json index dc7df0ed..13fc5e51 100644 --- a/docs.json +++ b/docs.json @@ -1715,7 +1715,6 @@ "aigw/product/administration/configure-data-visibility-settings", "aigw/product/administration/configure-virtual-key-access-permissions", "aigw/product/administration/configure-api-key-access-permissions", - "aigw/product/administration/configure-prompt-access-permissions", "aigw/product/administration/configure-guardrail-access-permissions", "aigw/product/administration/configure-integration-access-permissions" ] @@ -2147,8 +2146,7 @@ "group": "AWS", "pages": [ "aigw/self-hosting/hybrid-deployments/aws/eks", - "aigw/self-hosting/hybrid-deployments/aws/ecs", - "aigw/self-hosting/hybrid-deployments/aws/marketplace" + "aigw/self-hosting/hybrid-deployments/aws/ecs" ] }, { @@ -2804,8 +2802,8 @@ { "group": "Changelog", "pages": [ - "aigw/changelog/enterprise", - "aigw/changelog/data-service" + "changelog/enterprise", + "changelog/data-service" ] } ] @@ -3987,6 +3985,14 @@ } }, "redirects": [ + { + "source": "/aigw/changelog/enterprise", + "destination": "/changelog/enterprise" + }, + { + "source": "/aigw/changelog/data-service", + "destination": "/changelog/data-service" + }, { "source": "/integrations/observability-integrations", "destination": "/product/observability/opentelemetry/list-of-supported-otel-instrumenters" diff --git a/enterprise/pricing.mdx b/enterprise/pricing.mdx index adaee8d1..2df42b23 100644 --- a/enterprise/pricing.mdx +++ b/enterprise/pricing.mdx @@ -1,6 +1,5 @@ --- title: Portkey Enterprise Pricing Guide -mode: "center" noindex: "true" --- diff --git a/enterprise/security.mdx b/enterprise/security.mdx index ef88bfa6..47b79dde 100644 --- a/enterprise/security.mdx +++ b/enterprise/security.mdx @@ -1,7 +1,6 @@ --- title: "Security" description: "Compare SaaS and Hybrid security postures, data flows, and controls" -mode: "center" noindex: "true" --- diff --git a/enterprise/support.mdx b/enterprise/support.mdx index 63e94699..94993ad8 100644 --- a/enterprise/support.mdx +++ b/enterprise/support.mdx @@ -2,7 +2,6 @@ title: Portkey Enterprise Support Plan description: Support plans and service-level commitments for Portkey Enterprise customers noindex: "true" -mode: "center" --- ## Support Tiers diff --git a/help-center/mcp-gateway-troubleshooting.mdx b/help-center/mcp-gateway-troubleshooting.mdx index edce98af..62118c70 100644 --- a/help-center/mcp-gateway-troubleshooting.mdx +++ b/help-center/mcp-gateway-troubleshooting.mdx @@ -212,6 +212,51 @@ If you run the gateway yourself (instead of `mcp.portkey.ai`): 3. **Set `MCP_GATEWAY_BASE_URL`** to your public gateway URL (for example `https://`). The gateway uses this to construct callback and discovery URLs—if it's wrong, OAuth and discovery fail. 4. Make sure your ingress forwards the `Host` header so the server advertises its public URL (not an internal address). +With `SERVER_MODE=unified` (gateway 2.20.0 or later), both gateways share one port. Use `https:///m/{slug}/mcp` as the client URL in step 1. `/.well-known/*` and `/oauth/*` stay at the root, not under `/m`. + +### Client registration fails with "Invalid API Key. Error Code: 03" + +An API key works, but connecting without one fails before any login page opens. Claude Code reports it as: + +``` +SDK auth failed: Dynamic Client Registration rejected (HTTP 401): +{"status":"failure","message":"Portkey Error: Invalid API Key. Error Code: 03", ...} +``` + +**Why:** that response body comes from the AI Gateway, not the MCP Gateway. The MCP Gateway's `/.well-known/*`, `/oauth/register` and CORS preflight (`OPTIONS`) endpoints do not require an API key. When the MCP Gateway itself rejects a request, it answers `{"error":"unauthorized", ...}` with a `WWW-Authenticate` header. So the OAuth requests are reaching the AI Gateway service. API-key access still works because only the `/{slug}/mcp` path is routed correctly. + +**Confirm:** these checks are for `SERVER_MODE=all` or `mcp`, where the MCP Gateway has its own port (`8788`). In unified mode one port serves both paths, so this routing fault does not occur. + +1. Fetch the authorization server metadata and check that every URL in it uses your public MCP Gateway host: + + ```bash + curl https:///.well-known/oauth-authorization-server/{slug}/mcp + ``` + +2. Register a test client at the `registration_endpoint` it returns. Expect `201` with a `client_id`: + + ```bash + curl -i -X POST -H 'Content-Type: application/json' \ + -d '{"client_name":"probe","redirect_uris":["http://localhost:33418/callback"],"token_endpoint_auth_method":"none","grant_types":["authorization_code","refresh_token"],"response_types":["code"]}' + ``` + +3. If that returns the `401` above, repeat it from inside the gateway pod, which bypasses the load balancer and ingress: + + ```bash + kubectl exec -- wget -qO- --header 'Content-Type: application/json' \ + --post-data '{"client_name":"probe","redirect_uris":["http://localhost:33418/callback"],"token_endpoint_auth_method":"none","grant_types":["authorization_code","refresh_token"],"response_types":["code"]}' \ + http://localhost:8788/oauth/register + ``` + + A JSON body with a `client_id` here means the gateway is fine and the Service, load balancer or ingress is sending MCP traffic to the AI Gateway port (`8787`). + +**Fix:** + +- Set `MCP_GATEWAY_BASE_URL` to the public MCP Gateway URL, including the scheme and no path, then restart the pods. It is read at startup. +- With `SERVER_MODE=all`, route **every** path on the MCP Gateway host to the MCP port (`8788`), not only `/{slug}/mcp`. `/.well-known/*` and `/oauth/*` must reach it too. +- Or switch to `SERVER_MODE=unified`, which serves everything on one port and needs no host-based routing. +- In the client, remove the server and add it again with no API key header, so it starts a fresh OAuth flow. + --- ## Enterprise: `external_auth_config` returns 403 diff --git a/introduction/what-is-portkey.mdx b/introduction/what-is-portkey.mdx index d6f8da2a..372581d4 100644 --- a/introduction/what-is-portkey.mdx +++ b/introduction/what-is-portkey.mdx @@ -1,7 +1,6 @@ --- title: "What is Portkey?" description: Portkey AI is a comprehensive platform designed to streamline and enhance AI integration for developers and organizations. It serves as a unified interface for interacting with over 250 AI models, offering advanced tools for control, visibility, and security in your Generative AI apps. -mode: "wide" --- It takes 2 mins to integrate and with that, it starts monitoring all of your LLM requests and makes your app resilient, secure, performant, and more accurate at the same time. diff --git a/product/enterprise-offering.mdx b/product/enterprise-offering.mdx index 41c7c26d..c4499e65 100644 --- a/product/enterprise-offering.mdx +++ b/product/enterprise-offering.mdx @@ -1,6 +1,5 @@ --- title: "Enterprise Offering" -mode: "wide" --- diff --git a/product/product-feature-comparison.mdx b/product/product-feature-comparison.mdx index cdedb02f..884a5992 100644 --- a/product/product-feature-comparison.mdx +++ b/product/product-feature-comparison.mdx @@ -1,7 +1,6 @@ --- title: "Feature Comparison" description: Comparing Portkey's Open-source version and Dev, Pro, Enterprise plans. -mode: "wide" --- import AirgappedLegacy from "/snippets/airgapped-legacy.mdx"; diff --git a/snippets/aigw/coming-soon.mdx b/snippets/aigw/coming-soon.mdx index 9a8fdbb4..ab30905c 100644 --- a/snippets/aigw/coming-soon.mdx +++ b/snippets/aigw/coming-soon.mdx @@ -1,3 +1,3 @@ -This feature is coming soon. Contact [Portkey support](https://support.portkey.ai/forms/customer-portal-ticket-form) for early access. - +This feature is coming soon. + \ No newline at end of file diff --git a/snippets/aigw/portkey-advanced-features.mdx b/snippets/aigw/portkey-advanced-features.mdx index 70a4ba2e..33bd878c 100644 --- a/snippets/aigw/portkey-advanced-features.mdx +++ b/snippets/aigw/portkey-advanced-features.mdx @@ -1,16 +1,16 @@ -# 3. Set Up Enterprise Governance +# 3. Set Up Governance -**Why Enterprise Governance?** +**Why Governance?** - **Cost Management**: Controlling and tracking AI spending across teams - **Access Control**: Managing team access and workspaces -- **Usage Analytics**: Understanding how AI is being used across the organization -- **Security & Compliance**: Maintaining enterprise security standards +- **Usage Analytics**: Understanding how AI is being used across the organisation +- **Security & Compliance**: Maintaining your organisation's security standards - **Reliability**: Ensuring consistent service across all users - **Model Management**: Managing what models are being used in your setup -Portkey adds a comprehensive governance layer to address these enterprise needs. +The AI Gateway adds a governance layer to address these needs. -**Enterprise Implementation Guide** +**Implementation Guide** @@ -22,7 +22,7 @@ Model Catalog enables you to have granular control over LLM access at the team/d - Track departmental spending #### Setting Up Department-Specific Controls: -1. Navigate to [Model Catalog](https://app.portkey.ai/model-catalog) in Portkey dashboard +1. Navigate to [Model Catalog](https://stratacloudmanager.paloaltonetworks.com/) in Strata Cloud Manager 2. Create new Provider for each engineering team with budget limits and rate limits 3. Configure department-specific limits @@ -31,16 +31,18 @@ Model Catalog enables you to have granular control over LLM access at the team/d ### Step 2: Define Model Access Rules -As your AI usage scales, controlling which teams can access specific models becomes crucial. You can simply manage AI models in your org by provisioning model at the top integration level. +As your AI usage scales, controlling which teams can access specific models becomes crucial. You can manage AI models in your organisation by provisioning models at the top integration level. - - Portkey allows you to control your routing logic very simply with it's Configs feature. Portkey Configs provide this control layer with things like: + +### Step 3: Set Routing Configuration + +The AI Gateway lets you control your routing logic with its Configs feature. AI Gateway Configs provide this control layer with things like: - **Data Protection**: Implement guardrails for sensitive code and data - **Reliability Controls**: Add fallbacks, load-balance, retry and smart conditional routing logic -- **Caching**: Implement Simple and Semantic Caching. and more.... +- **Caching**: Implement Simple and Semantic Caching, and more #### Example Configuration: Here's a basic configuration to load-balance requests to OpenAI and Anthropic: @@ -65,7 +67,7 @@ Here's a basic configuration to load-balance requests to OpenAI and Anthropic: } ``` -Create your config on the [Configs page](https://app.portkey.ai/configs) in your Portkey dashboard. You'll need the config ID for connecting. +Create your config on the [Configs](https://stratacloudmanager.paloaltonetworks.com/) page in Strata Cloud Manager. You'll need the config ID for connecting. Configs can be updated anytime to adjust controls without affecting running applications. @@ -74,7 +76,7 @@ Configs can be updated anytime to adjust controls without affecting running appl -### Step 3: Implement Access Controls +### Step 4: Implement Access Controls Create User-specific API keys that automatically: - Track usage per developer/team with the help of metadata @@ -83,7 +85,7 @@ Create User-specific API keys that automatically: - Enforce access permissions Create API keys through: -- [Portkey App](https://app.portkey.ai/) +- [Strata Cloud Manager](https://stratacloudmanager.paloaltonetworks.com/) - [API Key Management API](/api-reference/admin-api/control-plane/api-keys/create-api-key) Example using Python SDK: @@ -112,10 +114,11 @@ For detailed key management instructions, see our [API Keys documentation](/api- -### Step 4: Deploy & Monitor -After distributing API keys to your engineering teams, your enterprise-ready setup is ready to go. Each developer can now use their designated API keys with appropriate access levels and budget controls. -Apply your governance setup using the integration steps from earlier sections -Monitor usage in Portkey dashboard: +### Step 5: Deploy & Monitor +After distributing API keys to your engineering teams, your setup is ready to go. Each developer can now use their designated API keys with appropriate access levels and budget controls. +Apply your governance setup using the integration steps from earlier sections. + +Monitor usage in the AI Gateway dashboard: - Cost tracking by engineering team - Model usage patterns for AI agent tasks - Request volumes @@ -125,7 +128,7 @@ Monitor usage in Portkey dashboard: -### Enterprise Features Now Available +### What's Now Available **You now have:** - Departmental budget controls @@ -136,14 +139,14 @@ Monitor usage in Portkey dashboard: -# Portkey Features -Now that you have an enterprise-grade setup, let's explore the comprehensive features Portkey provides to ensure secure, efficient, and cost-effective AI operations. +# AI Gateway Features +Now that your setup is in place, let's explore the features the AI Gateway provides to ensure secure, efficient, and cost-effective AI operations. ### 1. Comprehensive Metrics -Using Portkey you can track 40+ key metrics including cost, token usage, response time, and performance across all your LLM providers in real time. You can also filter these metrics based on custom metadata that you can set in your configs. Learn more about metadata here. +Using the AI Gateway you can track 40+ key metrics including cost, token usage, response time, and performance across all your LLM providers in real time. You can also filter these metrics based on custom metadata that you can set in your configs. Learn more about metadata here. ### 2. Advanced Logs -Portkey's logging dashboard provides detailed logs for every request made to your LLMs. These logs include: +The AI Gateway's logging dashboard provides detailed logs for every request made to your LLMs. These logs include: - Complete request and response tracking - Metadata tags for filtering - Cost attribution and much more... @@ -153,12 +156,12 @@ Portkey's logging dashboard provides detailed logs for every request made to you You can easily switch between 1600+ LLMs. Call various LLMs such as Anthropic, Gemini, Mistral, Azure OpenAI, Google Vertex AI, AWS Bedrock, and many more by simply changing the `provider` slug in your default `config` object. ### 4. Advanced Metadata Tracking -Using Portkey, you can add custom metadata to your LLM requests for detailed tracking and analytics. Use metadata tags to filter logs, track usage, and attribute costs across departments and teams. +Using the AI Gateway, you can add custom metadata to your LLM requests for detailed tracking and analytics. Use metadata tags to filter logs, track usage, and attribute costs across departments and teams. - + -### 5. Enterprise Access Management +### 5. Access Management @@ -166,11 +169,11 @@ Set and manage spending limits across teams and departments. Control costs with -Enterprise-grade SSO integration with support for SAML 2.0, Okta, Azure AD, and custom providers for secure authentication. +SSO integration with support for SAML 2.0, Okta, Azure AD, and custom providers for secure authentication. - -Hierarchical organization structure with workspaces, teams, and role-based access control for enterprise-scale deployments. + +Hierarchical organisation structure with workspaces, teams, and role-based access control for large-scale deployments. @@ -204,26 +207,26 @@ Automatic retry handling with exponential backoff for failed requests Protect your Project's data and enhance reliability with real-time checks on LLM inputs and outputs. Leverage guardrails to: - Prevent sensitive data leaks -- Enforce compliance with organizational policies +- Enforce compliance with organisational policies - PII detection and masking - Content filtering - Custom security rules - Data compliance checks -Implement real-time protection for your LLM interactions with automatic detection and filtering of sensitive content, PII, and custom security rules. Enable comprehensive data protection while maintaining compliance with organizational policies. +Implement real-time protection for your LLM interactions with automatic detection and filtering of sensitive content, PII, and custom security rules. Enable comprehensive data protection while maintaining compliance with organisational policies. # FAQs - Update AI Provider limits at any time from [Model Catalog](https://app.portkey.ai/model-catalog): 1. Open the provider you want to modify. 2. Update the budget or rate limits. 3. Save your changes. + Update AI Provider limits at any time from [Model Catalog](https://stratacloudmanager.paloaltonetworks.com/): 1. Open the provider you want to modify. 2. Update the budget or rate limits. 3. Save your changes. Yes! Add multiple AI Providers to Model Catalog (one for each provider) and attach them to a single config. This config can then be connected to your API key, allowing you to use multiple providers through a single API key. -Portkey provides several ways to track team costs: +The AI Gateway provides several ways to track team costs: - Create separate AI Providers for each team - Use metadata tags in your configs - Set up team-specific API keys @@ -242,7 +245,3 @@ When a team reaches their budget limit: **Join our Community** - [GitHub Repository](https://github.com/Portkey-AI) - - -For enterprise support and custom features, contact our [enterprise team](https://www.paloaltonetworks.com/ai-security/ai-gateway#:~:text=See%20Prisma%20AIRS%20AI%20Gateway%20in%20Action - \ No newline at end of file diff --git a/virtual_key_old/introduction/what-is-portkey.mdx b/virtual_key_old/introduction/what-is-portkey.mdx index 678765c9..6bf7a479 100644 --- a/virtual_key_old/introduction/what-is-portkey.mdx +++ b/virtual_key_old/introduction/what-is-portkey.mdx @@ -1,7 +1,6 @@ --- title: "What is Portkey?" description: Portkey AI is a comprehensive platform designed to streamline and enhance AI integration for developers and organizations. It serves as a unified interface for interacting with over 250 AI models, offering advanced tools for control, visibility, and security in your Generative AI apps. -mode: "wide" --- It takes 2 mins to integrate and with that, it starts monitoring all of your LLM requests and makes your app resilient, secure, performant, and more accurate at the same time. diff --git a/virtual_key_old/product/enterprise-offering.mdx b/virtual_key_old/product/enterprise-offering.mdx index f935e7c2..72bf993c 100644 --- a/virtual_key_old/product/enterprise-offering.mdx +++ b/virtual_key_old/product/enterprise-offering.mdx @@ -1,6 +1,5 @@ --- title: "Enterprise Offering" -mode: "wide" --- diff --git a/virtual_key_old/product/product-feature-comparison.mdx b/virtual_key_old/product/product-feature-comparison.mdx index 816523b0..6b3ca788 100644 --- a/virtual_key_old/product/product-feature-comparison.mdx +++ b/virtual_key_old/product/product-feature-comparison.mdx @@ -1,7 +1,6 @@ --- title: "Feature Comparison" description: Comparing Portkey's Open-source version and Dev, Pro, Enterprise plans. -mode: "wide" --- Portkey has a generous free tier (10k requests/month) on our **Dev** plan — but, as you move to production-scale, you may benefit from Portkey's **Pro** or **Enterprise** plans.