Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ This component computes the embeddings of a list of documents and stores the obt

The vectors computed by this component are necessary to perform embedding retrieval on a collection of documents. At retrieval time, the vector representing the query is compared with those of the documents to find the most similar or relevant documents.

To see the list of compatible embedding models, head over to Azure [documentation](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models?source=recommendations). The default model for `AzureOpenAITextEmbedder` is `text-embedding-ada-002`.
To see the list of compatible embedding models, head over to Azure [documentation](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models?source=recommendations). The default model for `AzureOpenAIDocumentEmbedder` is `text-embedding-3-small`.

This component should be used to embed a list of documents. To embed a string, you should use the [`AzureOpenAITextEmbedder`](azureopenaitextembedder.mdx).

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ When you perform embedding retrieval, you use this component to transform your q

`AzureOpenAITextEmbedder` transforms a string into a vector that captures its semantics using an OpenAI embedding model. It uses Azure cognitive services for text and document embedding with models deployed on Azure.

To see the list of compatible embedding models, head over to Azure [documentation](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models?source=recommendations). The default model for `AzureOpenAITextEmbedder` is `text-embedding-ada-002`.
To see the list of compatible embedding models, head over to Azure [documentation](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models?source=recommendations). The default model for `AzureOpenAITextEmbedder` is `text-embedding-3-small`.

Use `AzureOpenAITextEmbedder` to embed a simple string (such as a query) into a vector. For embedding lists of documents, use the [`AzureOpenAIDocumentEmbedder`](azureopenaidocumentembedder.mdx), which enriches the documents with the computed embedding, also known as vector.

Expand Down Expand Up @@ -63,7 +63,7 @@ text_embedder = AzureOpenAITextEmbedder()
print(text_embedder.run(text_to_embed))

# {'embedding': [0.017020374536514282, -0.023255806416273117, ...],
# 'meta': {'model': 'text-embedding-ada-002-v2',
# 'meta': {'model': 'text-embedding-3-small',
# 'usage': {'prompt_tokens': 4, 'total_tokens': 4}}}
```

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ The vectors computed by this component are necessary to perform embedding retrie

## Overview

To see the list of compatible OpenAI embedding models, head over to OpenAI [documentation](https://platform.openai.com/docs/guides/embeddings). The default model for `OpenAIDocumentEmbedder` is `text-embedding-ada-002`. You can specify another model with the `model` parameter when initializing this component.
To see the list of compatible OpenAI embedding models, head over to OpenAI [documentation](https://platform.openai.com/docs/guides/embeddings). The default model for `OpenAIDocumentEmbedder` is `text-embedding-3-small`. You can specify another model with the `model` parameter when initializing this component.

This component should be used to embed a list of documents. To embed a string, use the [OpenAITextEmbedder](openaitextembedder.mdx).

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ When you perform embedding retrieval, you use this component to transform your q

## Overview

To see the list of compatible OpenAI embedding models, head over to OpenAI [documentation](https://platform.openai.com/docs/guides/embeddings). The default model for `OpenAITextEmbedder` is `text-embedding-ada-002`. You can specify another model with the `model` parameter when initializing this component.
To see the list of compatible OpenAI embedding models, head over to OpenAI [documentation](https://platform.openai.com/docs/guides/embeddings). The default model for `OpenAITextEmbedder` is `text-embedding-3-small`. You can specify another model with the `model` parameter when initializing this component.

Use `OpenAITextEmbedder` to embed a simple string (such as a query) into a vector. For embedding lists of documents, use the [OpenAIDocumentEmbedder](openaidocumentembedder.mdx), which enriches the document with the computed embedding, also known as vector.

Expand All @@ -54,7 +54,7 @@ text_embedder = OpenAITextEmbedder(api_key=Secret.from_token("<your-api-key>"))
print(text_embedder.run(text_to_embed))

# {'embedding': [0.017020374536514282, -0.023255806416273117, ...],
# 'meta': {'model': 'text-embedding-ada-002-v2',
# 'meta': {'model': 'text-embedding-3-small',
# 'usage': {'prompt_tokens': 4, 'total_tokens': 4}}}
```

Expand Down
4 changes: 2 additions & 2 deletions haystack/components/embedders/azure_document_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ def __init__( # noqa: PLR0913, PLR0917 (too-many-arguments, too-many-positional
self,
azure_endpoint: str | None = None,
api_version: str | None = "2023-05-15",
azure_deployment: str = "text-embedding-ada-002",
azure_deployment: str = "text-embedding-3-small",
dimensions: int | None = None,
api_key: Secret | None = Secret.from_env_var("AZURE_OPENAI_API_KEY", strict=False),
azure_ad_token: Secret | None = Secret.from_env_var("AZURE_OPENAI_AD_TOKEN", strict=False),
Expand All @@ -67,7 +67,7 @@ def __init__( # noqa: PLR0913, PLR0917 (too-many-arguments, too-many-positional
:param api_version:
The version of the API to use.
:param azure_deployment:
The name of the model deployed on Azure. The default model is text-embedding-ada-002.
The name of the model deployed on Azure. The default is `text-embedding-3-small`.
:param dimensions:
The number of dimensions of the resulting embeddings. Only supported in text-embedding-3
and later models.
Expand Down
6 changes: 3 additions & 3 deletions haystack/components/embedders/azure_text_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ class AzureOpenAITextEmbedder(OpenAITextEmbedder):
print(text_embedder.run(text_to_embed))

# {'embedding': [0.017020374536514282, -0.023255806416273117, ...],
# 'meta': {'model': 'text-embedding-ada-002-v2',
# 'meta': {'model': 'text-embedding-3-small',
# 'usage': {'prompt_tokens': 4, 'total_tokens': 4}}}
```
"""
Expand All @@ -38,7 +38,7 @@ def __init__( # noqa: PLR0913
self,
azure_endpoint: str | None = None,
api_version: str | None = "2023-05-15",
azure_deployment: str = "text-embedding-ada-002",
azure_deployment: str = "text-embedding-3-small",
dimensions: int | None = None,
api_key: Secret | None = Secret.from_env_var("AZURE_OPENAI_API_KEY", strict=False),
azure_ad_token: Secret | None = Secret.from_env_var("AZURE_OPENAI_AD_TOKEN", strict=False),
Expand All @@ -60,7 +60,7 @@ def __init__( # noqa: PLR0913
:param api_version:
The version of the API to use.
:param azure_deployment:
The name of the model deployed on Azure. The default model is text-embedding-ada-002.
The name of the model deployed on Azure. The default is `text-embedding-3-small`.
:param dimensions:
The number of dimensions the resulting output embeddings should have. Only supported in text-embedding-3
and later models.
Expand Down
4 changes: 2 additions & 2 deletions haystack/components/embedders/openai_document_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ class OpenAIDocumentEmbedder:
def __init__( # noqa: PLR0913, PLR0917 (too-many-arguments, too-many-positional-arguments)
self,
api_key: Secret = Secret.from_env_var("OPENAI_API_KEY"),
model: str = "text-embedding-ada-002",
model: str = "text-embedding-3-small",
dimensions: int | None = None,
api_base_url: str | None = None,
organization: str | None = None,
Expand Down Expand Up @@ -71,7 +71,7 @@ def __init__( # noqa: PLR0913, PLR0917 (too-many-arguments, too-many-positional
during initialization.
:param model:
The name of the model to use for calculating embeddings.
The default model is `text-embedding-ada-002`.
The default model is `text-embedding-3-small`.
:param dimensions:
The number of dimensions of the resulting embeddings. Only `text-embedding-3` and
later models support this parameter.
Expand Down
6 changes: 3 additions & 3 deletions haystack/components/embedders/openai_text_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,15 +31,15 @@ class OpenAITextEmbedder:
print(text_embedder.run(text_to_embed))

# {'embedding': [0.017020374536514282, -0.023255806416273117, ...],
# 'meta': {'model': 'text-embedding-ada-002-v2',
# 'meta': {'model': 'text-embedding-3-small',
# 'usage': {'prompt_tokens': 4, 'total_tokens': 4}}}
```
"""

def __init__(
self,
api_key: Secret = Secret.from_env_var("OPENAI_API_KEY"),
model: str = "text-embedding-ada-002",
model: str = "text-embedding-3-small",
dimensions: int | None = None,
api_base_url: str | None = None,
organization: str | None = None,
Expand All @@ -62,7 +62,7 @@ def __init__(
during initialization.
:param model:
The name of the model to use for calculating embeddings.
The default model is `text-embedding-ada-002`.
The default model is `text-embedding-3-small`.
:param dimensions:
The number of dimensions of the resulting embeddings. Only `text-embedding-3` and
later models support this parameter.
Expand Down

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need double backtick for in-line code in release notes.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed

All inline code in the release note now uses double backticks

and the deprecation-page reference was replaced with a link to OpenAI's new-embedding-models announcement

The original "legacy/deprecated" framing was inaccurate ada-002 is not on the formal deprecation schedule

Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
---
upgrade:
- |
The default model for ``OpenAITextEmbedder``, ``OpenAIDocumentEmbedder``,
``AzureOpenAITextEmbedder``, and ``AzureOpenAIDocumentEmbedder`` has been changed
from ``text-embedding-ada-002`` to ``text-embedding-3-small``.

``text-embedding-3-small`` is roughly 5x cheaper per token than the previous
generation ``text-embedding-ada-002`` and scores higher on the MTEB benchmark
(see OpenAI's [announcement](https://openai.com/index/new-embedding-models-and-api-updates)).

To preserve previous behavior, pass ``model="text-embedding-ada-002"``
explicitly (or, for the Azure variants,
``azure_deployment="text-embedding-ada-002"``).

Note for Azure users: ``azure_deployment`` refers to the deployment name in
your Azure OpenAI resource, not a model identifier. If you do not already
have a ``text-embedding-3-small`` deployment in Azure, you will need to
create one or pass the name of an existing deployment explicitly.

Embeddings generated with the new default model are not compatible with
embeddings produced by ``text-embedding-ada-002``. If you have a document
store populated with previously-generated embeddings, either keep using
the old model explicitly or re-embed your corpus.
12 changes: 6 additions & 6 deletions test/components/embedders/test_azure_document_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,8 @@ class TestAzureOpenAIDocumentEmbedder:
def test_init_default(self, monkeypatch):
monkeypatch.setenv("AZURE_OPENAI_API_KEY", "fake-api-key")
embedder = AzureOpenAIDocumentEmbedder(azure_endpoint="https://example-resource.azure.openai.com/")
assert embedder.azure_deployment == "text-embedding-ada-002"
assert embedder.model == "text-embedding-ada-002"
assert embedder.azure_deployment == "text-embedding-3-small"
assert embedder.model == "text-embedding-3-small"
assert embedder.dimensions is None
assert embedder.organization is None
assert embedder.prefix == ""
Expand All @@ -43,8 +43,8 @@ def test_init_with_0_max_retries(self, monkeypatch):
embedder = AzureOpenAIDocumentEmbedder(
azure_endpoint="https://example-resource.azure.openai.com/", max_retries=0
)
assert embedder.azure_deployment == "text-embedding-ada-002"
assert embedder.model == "text-embedding-ada-002"
assert embedder.azure_deployment == "text-embedding-3-small"
assert embedder.model == "text-embedding-3-small"
assert embedder.dimensions is None
assert embedder.organization is None
assert embedder.prefix == ""
Expand All @@ -69,7 +69,7 @@ def test_to_dict(self, monkeypatch):
"api_key": {"env_vars": ["AZURE_OPENAI_API_KEY"], "strict": False, "type": "env_var"},
"azure_ad_token": {"env_vars": ["AZURE_OPENAI_AD_TOKEN"], "strict": False, "type": "env_var"},
"api_version": "2023-05-15",
"azure_deployment": "text-embedding-ada-002",
"azure_deployment": "text-embedding-3-small",
"dimensions": None,
"azure_endpoint": "https://example-resource.azure.openai.com/",
"organization": None,
Expand Down Expand Up @@ -262,7 +262,7 @@ def test_run(self):
Document(content="I love cheese", meta={"topic": "Cuisine"}),
Document(content="A transformer is a deep learning architecture", meta={"topic": "ML"}),
]
# the default model is text-embedding-ada-002 even if we don't specify it, but let's be explicit
# set the deployment explicitly instead of relying on the default
embedder = AzureOpenAIDocumentEmbedder(
azure_deployment="text-embedding-ada-002",
meta_fields_to_embed=["topic"],
Expand Down
12 changes: 6 additions & 6 deletions test/components/embedders/test_azure_text_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,8 @@ def test_init_default(self, monkeypatch):
embedder = AzureOpenAITextEmbedder(azure_endpoint="https://example-resource.azure.openai.com/")

assert embedder.api_key.resolve_value() == "fake-api-key"
assert embedder.azure_deployment == "text-embedding-ada-002"
assert embedder.model == "text-embedding-ada-002"
assert embedder.azure_deployment == "text-embedding-3-small"
assert embedder.model == "text-embedding-3-small"
assert embedder.dimensions is None
assert embedder.organization is None
assert embedder.prefix == ""
Expand All @@ -39,8 +39,8 @@ def test_init_with_zero_max_retries(self, monkeypatch):
embedder = AzureOpenAITextEmbedder(azure_endpoint="https://example-resource.azure.openai.com/", max_retries=0)

assert embedder.api_key.resolve_value() == "fake-api-key"
assert embedder.azure_deployment == "text-embedding-ada-002"
assert embedder.model == "text-embedding-ada-002"
assert embedder.azure_deployment == "text-embedding-3-small"
assert embedder.model == "text-embedding-3-small"
assert embedder.dimensions is None
assert embedder.organization is None
assert embedder.prefix == ""
Expand All @@ -60,7 +60,7 @@ def test_to_dict_default(self, monkeypatch):
"init_parameters": {
"api_key": {"env_vars": ["AZURE_OPENAI_API_KEY"], "strict": False, "type": "env_var"},
"azure_ad_token": {"env_vars": ["AZURE_OPENAI_AD_TOKEN"], "strict": False, "type": "env_var"},
"azure_deployment": "text-embedding-ada-002",
"azure_deployment": "text-embedding-3-small",
"dimensions": None,
"organization": None,
"azure_endpoint": "https://example-resource.azure.openai.com/",
Expand Down Expand Up @@ -188,7 +188,7 @@ def test_from_dict_with_parameters(self, monkeypatch):
),
)
def test_run(self):
# the default model is text-embedding-ada-002 even if we don't specify it, but let's be explicit
# set the deployment explicitly instead of relying on the default
embedder = AzureOpenAITextEmbedder(
azure_deployment="text-embedding-ada-002", prefix="prefix ", suffix=" suffix", organization="HaystackCI"
)
Expand Down
4 changes: 2 additions & 2 deletions test/components/embedders/test_openai_document_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ def test_init_default(self, monkeypatch):
monkeypatch.setenv("OPENAI_API_KEY", "fake-api-key")
embedder = OpenAIDocumentEmbedder()
assert embedder.api_key.resolve_value() == "fake-api-key"
assert embedder.model == "text-embedding-ada-002"
assert embedder.model == "text-embedding-3-small"
assert embedder.organization is None
assert embedder.prefix == ""
assert embedder.suffix == ""
Expand Down Expand Up @@ -100,7 +100,7 @@ def test_to_dict(self, monkeypatch):
"init_parameters": {
"api_key": {"env_vars": ["OPENAI_API_KEY"], "strict": True, "type": "env_var"},
"api_base_url": None,
"model": "text-embedding-ada-002",
"model": "text-embedding-3-small",
"dimensions": None,
"organization": None,
"http_client_kwargs": None,
Expand Down
6 changes: 3 additions & 3 deletions test/components/embedders/test_openai_text_embedder.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ def test_init_default(self, monkeypatch):
embedder = OpenAITextEmbedder()

assert embedder.api_key.resolve_value() == "fake-api-key"
assert embedder.model == "text-embedding-ada-002"
assert embedder.model == "text-embedding-3-small"
assert embedder.api_base_url is None
assert embedder.organization is None
assert embedder.prefix == ""
Expand Down Expand Up @@ -87,7 +87,7 @@ def test_to_dict(self, monkeypatch):
"api_key": {"env_vars": ["OPENAI_API_KEY"], "strict": True, "type": "env_var"},
"api_base_url": None,
"dimensions": None,
"model": "text-embedding-ada-002",
"model": "text-embedding-3-small",
"organization": None,
"http_client_kwargs": None,
"prefix": "",
Expand Down Expand Up @@ -157,7 +157,7 @@ def test_prepare_input(self, monkeypatch):
inp = "The food was delicious"
prepared_input = embedder._prepare_input(inp)
assert prepared_input == {
"model": "text-embedding-ada-002",
"model": "text-embedding-3-small",
"input": "The food was delicious",
"encoding_format": "float",
"dimensions": 1536,
Expand Down
Loading