diff --git a/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx b/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx index e9c5de9fc9..e961cf24ca 100644 --- a/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx +++ b/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx @@ -16,7 +16,7 @@ Use this component to retrieve neighboring sentences around relevant sentences t | **Most common position in a pipeline** | Used after the main Retriever component, like the `InMemoryEmbeddingRetriever` or any other Retriever. | | **Mandatory init variables** | `document_store`: An instance of a Document Store | | **Mandatory run variables** | `retrieved_documents`: A list of already retrieved documents for which you want to get a context window | -| **Output variables** | `context_windows`: A list of strings

`context_documents`: A list of documents ordered by `split_idx_start` | +| **Output variables** | `context_windows`: A list of strings, one per retrieved document

`context_documents`: A list of documents, grouped by retrieved document and ordered by `split_id` within each window | | **API reference** | [Retrievers](/reference/retrievers-api) | | **GitHub link** | https://github.com/deepset-ai/haystack/blob/main/haystack/components/retrievers/sentence_window_retriever.py | | **Package name** | `haystack-ai` | @@ -25,13 +25,24 @@ Use this component to retrieve neighboring sentences around relevant sentences t ## Overview -The "sentence window" is a retrieval technique that allows for the retrieval of the context around relevant sentences. +Sentence-window retrieval is a technique for retrieving the context around relevant text. During indexing, documents are split into small chunks, such as sentences, and written to a Document Store. During retrieval, a Retriever such as `InMemoryEmbeddingRetriever` or `InMemoryBM25Retriever` finds the chunks most relevant to the query. `SentenceWindowRetriever` then fetches up to `window_size` chunks before and after each of them (3 by default) from the same source document. -During indexing, documents are broken into smaller chunks or sentences and indexed. During retrieval, the sentences most relevant to a given query, based on a certain similarity metric, are retrieved. +This combines the strengths of both chunk sizes: small chunks match a query precisely, and the surrounding window gives the LLM enough context to answer. Despite its name, the component works with chunks of any size, not only sentences. You can override `window_size` for a single call by passing it to `run()`. -Once we have the relevant sentences, we can retrieve neighboring sentences to provide full context. The number of neighboring sentences to retrieve is defined by a fixed number of sentences before and after the relevant sentence. +### Required metadata -This component is meant to be used with other Retrievers, such as the `InMemoryEmbeddingRetriever`. These Retrievers find relevant sentences by comparing a query against indexed sentences using a similarity metric. Then, the `SentenceWindowRetriever` component retrieves neighboring sentences around the relevant ones by leveraging metadata stored in the `Document` object. +The component uses these `meta` fields to find neighboring chunks: + +- `source_id`: identifies the original document a chunk comes from. To match on several fields, pass a list to `source_id_meta_field`. +- `split_id`: the chunk's position within its source document. To use a different field, set `split_id_meta_field`. +- `split_idx_start` (optional): the chunk's start position in the source text. If every chunk in a window has it, text that overlaps between chunks is removed when they're merged. Otherwise, the chunks are joined in `split_id` order without removing any overlap. + +[`DocumentSplitter`](../preprocessors/documentsplitter.mdx) and [`RecursiveDocumentSplitter`](../preprocessors/recursivesplitter.mdx) add all three fields. If a retrieved document is missing `source_id` or `split_id`, the component raises a `ValueError`. To pass such documents through unchanged instead, set `raise_on_missing_meta_fields=False`. + +### Outputs + +- `context_windows`: one string per retrieved document, in the same order, containing the merged text of its window. +- `context_documents`: the documents in each window, grouped by retrieved document and ordered by `split_id` within each window. If two retrieved documents are close together, their windows overlap and the shared chunks appear in both. ## Usage diff --git a/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx b/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx index e9c5de9fc9..d0d5ccbe8b 100644 --- a/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx +++ b/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx @@ -16,7 +16,7 @@ Use this component to retrieve neighboring sentences around relevant sentences t | **Most common position in a pipeline** | Used after the main Retriever component, like the `InMemoryEmbeddingRetriever` or any other Retriever. | | **Mandatory init variables** | `document_store`: An instance of a Document Store | | **Mandatory run variables** | `retrieved_documents`: A list of already retrieved documents for which you want to get a context window | -| **Output variables** | `context_windows`: A list of strings

`context_documents`: A list of documents ordered by `split_idx_start` | +| **Output variables** | `context_windows`: A list of strings, one per retrieved document

`context_documents`: A list of documents, grouped by retrieved document and ordered by `split_id` within each window | | **API reference** | [Retrievers](/reference/retrievers-api) | | **GitHub link** | https://github.com/deepset-ai/haystack/blob/main/haystack/components/retrievers/sentence_window_retriever.py | | **Package name** | `haystack-ai` | diff --git a/haystack/components/retrievers/sentence_window_retriever.py b/haystack/components/retrievers/sentence_window_retriever.py index 3e60a1c39e..b5a89fa3df 100644 --- a/haystack/components/retrievers/sentence_window_retriever.py +++ b/haystack/components/retrievers/sentence_window_retriever.py @@ -192,9 +192,11 @@ def run(self, retrieved_documents: list[Document], window_size: int | None = Non A dictionary with the following keys: - `context_windows`: A list of strings, where each string represents the concatenated text from the context window of the corresponding document in `retrieved_documents`. - - `context_documents`: A list `Document` objects, containing the retrieved documents plus the context - document surrounding them. The documents are sorted by the `split_idx_start` - meta field. + - `context_documents`: A list of `Document` objects, containing the retrieved documents plus the + context documents surrounding them, grouped by window in the order of + `retrieved_documents`. Within each window, the documents are sorted by the + meta field set in `split_id_meta_field`. Documents shared by overlapping + windows appear once per window. """ window_size = self.window_size if window_size is None else window_size @@ -223,9 +225,11 @@ async def run_async(self, retrieved_documents: list[Document], window_size: int A dictionary with the following keys: - `context_windows`: A list of strings, where each string represents the concatenated text from the context window of the corresponding document in `retrieved_documents`. - - `context_documents`: A list `Document` objects, containing the retrieved documents plus the context - document surrounding them. The documents are sorted by the `split_idx_start` - meta field. + - `context_documents`: A list of `Document` objects, containing the retrieved documents plus the + context documents surrounding them, grouped by window in the order of + `retrieved_documents`. Within each window, the documents are sorted by the + meta field set in `split_id_meta_field`. Documents shared by overlapping + windows appear once per window. """ window_size = self.window_size if window_size is None else window_size @@ -322,8 +326,9 @@ def _assemble_context( for fetched in fetched_documents if self._is_in_window(fetched, source_ids, split_id - window_size, split_id + window_size) ] - context_text.append(self.merge_documents_text(context_docs)) - context_documents.extend(sorted(context_docs, key=lambda d: d.meta[self.split_id_meta_field])) + context_docs_sorted = sorted(context_docs, key=lambda d: d.meta[self.split_id_meta_field]) + context_text.append(self.merge_documents_text(context_docs_sorted)) + context_documents.extend(context_docs_sorted) return {"context_windows": context_text, "context_documents": context_documents} diff --git a/releasenotes/notes/fix-sentence-window-retriever-context-order-a1267b171dea93bb.yaml b/releasenotes/notes/fix-sentence-window-retriever-context-order-a1267b171dea93bb.yaml new file mode 100644 index 0000000000..6d04fbcdd2 --- /dev/null +++ b/releasenotes/notes/fix-sentence-window-retriever-context-order-a1267b171dea93bb.yaml @@ -0,0 +1,8 @@ +--- +fixes: + - | + Fixed ``SentenceWindowRetriever`` returning ``context_windows`` with the chunks in the wrong order when the + documents have no ``split_idx_start`` meta field, for example chunks produced by ``MarkdownHeaderSplitter``. + The chunks were concatenated in the order returned by the Document Store instead of by ``split_id``, so the + context passed to an LLM could be scrambled while ``context_documents`` was correctly sorted. The chunks are now + sorted by ``split_id`` before being merged, in both ``run`` and ``run_async``. diff --git a/test/components/retrievers/test_sentence_window_retriever.py b/test/components/retrievers/test_sentence_window_retriever.py index 9cc05f3046..c2560a5014 100644 --- a/test/components/retrievers/test_sentence_window_retriever.py +++ b/test/components/retrievers/test_sentence_window_retriever.py @@ -284,6 +284,10 @@ def test_run_custom_fields(self, in_memory_doc_store): # run the retriever with a document whose content = "Sentence 4." result = retriever.run(retrieved_documents=[doc for doc in docs if doc.content == "Sentence 4."]) assert len(result["context_documents"]) == 7 + assert [doc.meta["split_id_test"] for doc in result["context_documents"]] == [1, 2, 3, 4, 5, 6, 7] + assert result["context_windows"] == [ + "Sentence 1.Sentence 2.Sentence 3.Sentence 4.Sentence 5.Sentence 6.Sentence 7." + ] def test_run_with_multiple_source_ids(self, in_memory_doc_store): docs = [ diff --git a/test/components/retrievers/test_sentence_window_retriever_async.py b/test/components/retrievers/test_sentence_window_retriever_async.py index 4d3c728eeb..81a85c9b31 100644 --- a/test/components/retrievers/test_sentence_window_retriever_async.py +++ b/test/components/retrievers/test_sentence_window_retriever_async.py @@ -176,6 +176,10 @@ async def test_run_async_custom_fields(self, in_memory_doc_store): # run the retriever with a document whose content = "Sentence 4." result = await retriever.run_async(retrieved_documents=[doc for doc in docs if doc.content == "Sentence 4."]) assert len(result["context_documents"]) == 7 + assert [doc.meta["split_id_test"] for doc in result["context_documents"]] == [1, 2, 3, 4, 5, 6, 7] + assert result["context_windows"] == [ + "Sentence 1.Sentence 2.Sentence 3.Sentence 4.Sentence 5.Sentence 6.Sentence 7." + ] @pytest.mark.asyncio async def test_run_async_with_multiple_source_ids(self, in_memory_doc_store):