diff --git a/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx b/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx
index e9c5de9fc9..e961cf24ca 100644
--- a/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx
+++ b/docs-website/docs/pipeline-components/retrievers/sentencewindowretriever.mdx
@@ -16,7 +16,7 @@ Use this component to retrieve neighboring sentences around relevant sentences t
| **Most common position in a pipeline** | Used after the main Retriever component, like the `InMemoryEmbeddingRetriever` or any other Retriever. |
| **Mandatory init variables** | `document_store`: An instance of a Document Store |
| **Mandatory run variables** | `retrieved_documents`: A list of already retrieved documents for which you want to get a context window |
-| **Output variables** | `context_windows`: A list of strings
`context_documents`: A list of documents ordered by `split_idx_start` |
+| **Output variables** | `context_windows`: A list of strings, one per retrieved document
`context_documents`: A list of documents, grouped by retrieved document and ordered by `split_id` within each window |
| **API reference** | [Retrievers](/reference/retrievers-api) |
| **GitHub link** | https://github.com/deepset-ai/haystack/blob/main/haystack/components/retrievers/sentence_window_retriever.py |
| **Package name** | `haystack-ai` |
@@ -25,13 +25,24 @@ Use this component to retrieve neighboring sentences around relevant sentences t
## Overview
-The "sentence window" is a retrieval technique that allows for the retrieval of the context around relevant sentences.
+Sentence-window retrieval is a technique for retrieving the context around relevant text. During indexing, documents are split into small chunks, such as sentences, and written to a Document Store. During retrieval, a Retriever such as `InMemoryEmbeddingRetriever` or `InMemoryBM25Retriever` finds the chunks most relevant to the query. `SentenceWindowRetriever` then fetches up to `window_size` chunks before and after each of them (3 by default) from the same source document.
-During indexing, documents are broken into smaller chunks or sentences and indexed. During retrieval, the sentences most relevant to a given query, based on a certain similarity metric, are retrieved.
+This combines the strengths of both chunk sizes: small chunks match a query precisely, and the surrounding window gives the LLM enough context to answer. Despite its name, the component works with chunks of any size, not only sentences. You can override `window_size` for a single call by passing it to `run()`.
-Once we have the relevant sentences, we can retrieve neighboring sentences to provide full context. The number of neighboring sentences to retrieve is defined by a fixed number of sentences before and after the relevant sentence.
+### Required metadata
-This component is meant to be used with other Retrievers, such as the `InMemoryEmbeddingRetriever`. These Retrievers find relevant sentences by comparing a query against indexed sentences using a similarity metric. Then, the `SentenceWindowRetriever` component retrieves neighboring sentences around the relevant ones by leveraging metadata stored in the `Document` object.
+The component uses these `meta` fields to find neighboring chunks:
+
+- `source_id`: identifies the original document a chunk comes from. To match on several fields, pass a list to `source_id_meta_field`.
+- `split_id`: the chunk's position within its source document. To use a different field, set `split_id_meta_field`.
+- `split_idx_start` (optional): the chunk's start position in the source text. If every chunk in a window has it, text that overlaps between chunks is removed when they're merged. Otherwise, the chunks are joined in `split_id` order without removing any overlap.
+
+[`DocumentSplitter`](../preprocessors/documentsplitter.mdx) and [`RecursiveDocumentSplitter`](../preprocessors/recursivesplitter.mdx) add all three fields. If a retrieved document is missing `source_id` or `split_id`, the component raises a `ValueError`. To pass such documents through unchanged instead, set `raise_on_missing_meta_fields=False`.
+
+### Outputs
+
+- `context_windows`: one string per retrieved document, in the same order, containing the merged text of its window.
+- `context_documents`: the documents in each window, grouped by retrieved document and ordered by `split_id` within each window. If two retrieved documents are close together, their windows overlap and the shared chunks appear in both.
## Usage
diff --git a/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx b/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx
index e9c5de9fc9..d0d5ccbe8b 100644
--- a/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx
+++ b/docs-website/versioned_docs/version-3.3/pipeline-components/retrievers/sentencewindowretriever.mdx
@@ -16,7 +16,7 @@ Use this component to retrieve neighboring sentences around relevant sentences t
| **Most common position in a pipeline** | Used after the main Retriever component, like the `InMemoryEmbeddingRetriever` or any other Retriever. |
| **Mandatory init variables** | `document_store`: An instance of a Document Store |
| **Mandatory run variables** | `retrieved_documents`: A list of already retrieved documents for which you want to get a context window |
-| **Output variables** | `context_windows`: A list of strings
`context_documents`: A list of documents ordered by `split_idx_start` |
+| **Output variables** | `context_windows`: A list of strings, one per retrieved document
`context_documents`: A list of documents, grouped by retrieved document and ordered by `split_id` within each window |
| **API reference** | [Retrievers](/reference/retrievers-api) |
| **GitHub link** | https://github.com/deepset-ai/haystack/blob/main/haystack/components/retrievers/sentence_window_retriever.py |
| **Package name** | `haystack-ai` |
diff --git a/haystack/components/retrievers/sentence_window_retriever.py b/haystack/components/retrievers/sentence_window_retriever.py
index 3e60a1c39e..b5a89fa3df 100644
--- a/haystack/components/retrievers/sentence_window_retriever.py
+++ b/haystack/components/retrievers/sentence_window_retriever.py
@@ -192,9 +192,11 @@ def run(self, retrieved_documents: list[Document], window_size: int | None = Non
A dictionary with the following keys:
- `context_windows`: A list of strings, where each string represents the concatenated text from the
context window of the corresponding document in `retrieved_documents`.
- - `context_documents`: A list `Document` objects, containing the retrieved documents plus the context
- document surrounding them. The documents are sorted by the `split_idx_start`
- meta field.
+ - `context_documents`: A list of `Document` objects, containing the retrieved documents plus the
+ context documents surrounding them, grouped by window in the order of
+ `retrieved_documents`. Within each window, the documents are sorted by the
+ meta field set in `split_id_meta_field`. Documents shared by overlapping
+ windows appear once per window.
"""
window_size = self.window_size if window_size is None else window_size
@@ -223,9 +225,11 @@ async def run_async(self, retrieved_documents: list[Document], window_size: int
A dictionary with the following keys:
- `context_windows`: A list of strings, where each string represents the concatenated text from the
context window of the corresponding document in `retrieved_documents`.
- - `context_documents`: A list `Document` objects, containing the retrieved documents plus the context
- document surrounding them. The documents are sorted by the `split_idx_start`
- meta field.
+ - `context_documents`: A list of `Document` objects, containing the retrieved documents plus the
+ context documents surrounding them, grouped by window in the order of
+ `retrieved_documents`. Within each window, the documents are sorted by the
+ meta field set in `split_id_meta_field`. Documents shared by overlapping
+ windows appear once per window.
"""
window_size = self.window_size if window_size is None else window_size
@@ -322,8 +326,9 @@ def _assemble_context(
for fetched in fetched_documents
if self._is_in_window(fetched, source_ids, split_id - window_size, split_id + window_size)
]
- context_text.append(self.merge_documents_text(context_docs))
- context_documents.extend(sorted(context_docs, key=lambda d: d.meta[self.split_id_meta_field]))
+ context_docs_sorted = sorted(context_docs, key=lambda d: d.meta[self.split_id_meta_field])
+ context_text.append(self.merge_documents_text(context_docs_sorted))
+ context_documents.extend(context_docs_sorted)
return {"context_windows": context_text, "context_documents": context_documents}
diff --git a/releasenotes/notes/fix-sentence-window-retriever-context-order-a1267b171dea93bb.yaml b/releasenotes/notes/fix-sentence-window-retriever-context-order-a1267b171dea93bb.yaml
new file mode 100644
index 0000000000..6d04fbcdd2
--- /dev/null
+++ b/releasenotes/notes/fix-sentence-window-retriever-context-order-a1267b171dea93bb.yaml
@@ -0,0 +1,8 @@
+---
+fixes:
+ - |
+ Fixed ``SentenceWindowRetriever`` returning ``context_windows`` with the chunks in the wrong order when the
+ documents have no ``split_idx_start`` meta field, for example chunks produced by ``MarkdownHeaderSplitter``.
+ The chunks were concatenated in the order returned by the Document Store instead of by ``split_id``, so the
+ context passed to an LLM could be scrambled while ``context_documents`` was correctly sorted. The chunks are now
+ sorted by ``split_id`` before being merged, in both ``run`` and ``run_async``.
diff --git a/test/components/retrievers/test_sentence_window_retriever.py b/test/components/retrievers/test_sentence_window_retriever.py
index 9cc05f3046..c2560a5014 100644
--- a/test/components/retrievers/test_sentence_window_retriever.py
+++ b/test/components/retrievers/test_sentence_window_retriever.py
@@ -284,6 +284,10 @@ def test_run_custom_fields(self, in_memory_doc_store):
# run the retriever with a document whose content = "Sentence 4."
result = retriever.run(retrieved_documents=[doc for doc in docs if doc.content == "Sentence 4."])
assert len(result["context_documents"]) == 7
+ assert [doc.meta["split_id_test"] for doc in result["context_documents"]] == [1, 2, 3, 4, 5, 6, 7]
+ assert result["context_windows"] == [
+ "Sentence 1.Sentence 2.Sentence 3.Sentence 4.Sentence 5.Sentence 6.Sentence 7."
+ ]
def test_run_with_multiple_source_ids(self, in_memory_doc_store):
docs = [
diff --git a/test/components/retrievers/test_sentence_window_retriever_async.py b/test/components/retrievers/test_sentence_window_retriever_async.py
index 4d3c728eeb..81a85c9b31 100644
--- a/test/components/retrievers/test_sentence_window_retriever_async.py
+++ b/test/components/retrievers/test_sentence_window_retriever_async.py
@@ -176,6 +176,10 @@ async def test_run_async_custom_fields(self, in_memory_doc_store):
# run the retriever with a document whose content = "Sentence 4."
result = await retriever.run_async(retrieved_documents=[doc for doc in docs if doc.content == "Sentence 4."])
assert len(result["context_documents"]) == 7
+ assert [doc.meta["split_id_test"] for doc in result["context_documents"]] == [1, 2, 3, 4, 5, 6, 7]
+ assert result["context_windows"] == [
+ "Sentence 1.Sentence 2.Sentence 3.Sentence 4.Sentence 5.Sentence 6.Sentence 7."
+ ]
@pytest.mark.asyncio
async def test_run_async_with_multiple_source_ids(self, in_memory_doc_store):