Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ Use this component to retrieve neighboring sentences around relevant sentences t
| **Most common position in a pipeline** | Used after the main Retriever component, like the `InMemoryEmbeddingRetriever` or any other Retriever. |
| **Mandatory init variables** | `document_store`: An instance of a Document Store |
| **Mandatory run variables** | `retrieved_documents`: A list of already retrieved documents for which you want to get a context window |
| **Output variables** | `context_windows`: A list of strings <br /> <br />`context_documents`: A list of documents ordered by `split_idx_start` |
| **Output variables** | `context_windows`: A list of strings, one per retrieved document <br /> <br />`context_documents`: A list of documents, grouped by retrieved document and ordered by `split_id` within each window |
| **API reference** | [Retrievers](/reference/retrievers-api) |
| **GitHub link** | https://github.com/deepset-ai/haystack/blob/main/haystack/components/retrievers/sentence_window_retriever.py |
| **Package name** | `haystack-ai` |
Expand All @@ -25,13 +25,24 @@ Use this component to retrieve neighboring sentences around relevant sentences t

## Overview

The "sentence window" is a retrieval technique that allows for the retrieval of the context around relevant sentences.
Sentence-window retrieval is a technique for retrieving the context around relevant text. During indexing, documents are split into small chunks, such as sentences, and written to a Document Store. During retrieval, a Retriever such as `InMemoryEmbeddingRetriever` or `InMemoryBM25Retriever` finds the chunks most relevant to the query. `SentenceWindowRetriever` then fetches up to `window_size` chunks before and after each of them (3 by default) from the same source document.

During indexing, documents are broken into smaller chunks or sentences and indexed. During retrieval, the sentences most relevant to a given query, based on a certain similarity metric, are retrieved.
This combines the strengths of both chunk sizes: small chunks match a query precisely, and the surrounding window gives the LLM enough context to answer. Despite its name, the component works with chunks of any size, not only sentences. You can override `window_size` for a single call by passing it to `run()`.

Once we have the relevant sentences, we can retrieve neighboring sentences to provide full context. The number of neighboring sentences to retrieve is defined by a fixed number of sentences before and after the relevant sentence.
### Required metadata

This component is meant to be used with other Retrievers, such as the `InMemoryEmbeddingRetriever`. These Retrievers find relevant sentences by comparing a query against indexed sentences using a similarity metric. Then, the `SentenceWindowRetriever` component retrieves neighboring sentences around the relevant ones by leveraging metadata stored in the `Document` object.
The component uses these `meta` fields to find neighboring chunks:

- `source_id`: identifies the original document a chunk comes from. To match on several fields, pass a list to `source_id_meta_field`.
- `split_id`: the chunk's position within its source document. To use a different field, set `split_id_meta_field`.
- `split_idx_start` (optional): the chunk's start position in the source text. If every chunk in a window has it, text that overlaps between chunks is removed when they're merged. Otherwise, the chunks are joined in `split_id` order without removing any overlap.

[`DocumentSplitter`](../preprocessors/documentsplitter.mdx) and [`RecursiveDocumentSplitter`](../preprocessors/recursivesplitter.mdx) add all three fields. If a retrieved document is missing `source_id` or `split_id`, the component raises a `ValueError`. To pass such documents through unchanged instead, set `raise_on_missing_meta_fields=False`.

### Outputs

- `context_windows`: one string per retrieved document, in the same order, containing the merged text of its window.
- `context_documents`: the documents in each window, grouped by retrieved document and ordered by `split_id` within each window. If two retrieved documents are close together, their windows overlap and the shared chunks appear in both.

## Usage

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ Use this component to retrieve neighboring sentences around relevant sentences t
| **Most common position in a pipeline** | Used after the main Retriever component, like the `InMemoryEmbeddingRetriever` or any other Retriever. |
| **Mandatory init variables** | `document_store`: An instance of a Document Store |
| **Mandatory run variables** | `retrieved_documents`: A list of already retrieved documents for which you want to get a context window |
| **Output variables** | `context_windows`: A list of strings <br /> <br />`context_documents`: A list of documents ordered by `split_idx_start` |
| **Output variables** | `context_windows`: A list of strings, one per retrieved document <br /> <br />`context_documents`: A list of documents, grouped by retrieved document and ordered by `split_id` within each window |
| **API reference** | [Retrievers](/reference/retrievers-api) |
| **GitHub link** | https://github.com/deepset-ai/haystack/blob/main/haystack/components/retrievers/sentence_window_retriever.py |
| **Package name** | `haystack-ai` |
Expand Down
21 changes: 13 additions & 8 deletions haystack/components/retrievers/sentence_window_retriever.py
Original file line number Diff line number Diff line change
Expand Up @@ -192,9 +192,11 @@ def run(self, retrieved_documents: list[Document], window_size: int | None = Non
A dictionary with the following keys:
- `context_windows`: A list of strings, where each string represents the concatenated text from the
context window of the corresponding document in `retrieved_documents`.
- `context_documents`: A list `Document` objects, containing the retrieved documents plus the context
document surrounding them. The documents are sorted by the `split_idx_start`
meta field.
- `context_documents`: A list of `Document` objects, containing the retrieved documents plus the
context documents surrounding them, grouped by window in the order of
`retrieved_documents`. Within each window, the documents are sorted by the
meta field set in `split_id_meta_field`. Documents shared by overlapping
windows appear once per window.

"""
window_size = self.window_size if window_size is None else window_size
Expand Down Expand Up @@ -223,9 +225,11 @@ async def run_async(self, retrieved_documents: list[Document], window_size: int
A dictionary with the following keys:
- `context_windows`: A list of strings, where each string represents the concatenated text from the
context window of the corresponding document in `retrieved_documents`.
- `context_documents`: A list `Document` objects, containing the retrieved documents plus the context
document surrounding them. The documents are sorted by the `split_idx_start`
meta field.
- `context_documents`: A list of `Document` objects, containing the retrieved documents plus the
context documents surrounding them, grouped by window in the order of
`retrieved_documents`. Within each window, the documents are sorted by the
meta field set in `split_id_meta_field`. Documents shared by overlapping
windows appear once per window.

"""
window_size = self.window_size if window_size is None else window_size
Expand Down Expand Up @@ -322,8 +326,9 @@ def _assemble_context(
for fetched in fetched_documents
if self._is_in_window(fetched, source_ids, split_id - window_size, split_id + window_size)
]
context_text.append(self.merge_documents_text(context_docs))
context_documents.extend(sorted(context_docs, key=lambda d: d.meta[self.split_id_meta_field]))
context_docs_sorted = sorted(context_docs, key=lambda d: d.meta[self.split_id_meta_field])
context_text.append(self.merge_documents_text(context_docs_sorted))
context_documents.extend(context_docs_sorted)

return {"context_windows": context_text, "context_documents": context_documents}

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
---
fixes:
- |
Fixed ``SentenceWindowRetriever`` returning ``context_windows`` with the chunks in the wrong order when the
documents have no ``split_idx_start`` meta field, for example chunks produced by ``MarkdownHeaderSplitter``.
The chunks were concatenated in the order returned by the Document Store instead of by ``split_id``, so the
context passed to an LLM could be scrambled while ``context_documents`` was correctly sorted. The chunks are now
sorted by ``split_id`` before being merged, in both ``run`` and ``run_async``.
Original file line number Diff line number Diff line change
Expand Up @@ -284,6 +284,10 @@ def test_run_custom_fields(self, in_memory_doc_store):
# run the retriever with a document whose content = "Sentence 4."
result = retriever.run(retrieved_documents=[doc for doc in docs if doc.content == "Sentence 4."])
assert len(result["context_documents"]) == 7
assert [doc.meta["split_id_test"] for doc in result["context_documents"]] == [1, 2, 3, 4, 5, 6, 7]
assert result["context_windows"] == [
"Sentence 1.Sentence 2.Sentence 3.Sentence 4.Sentence 5.Sentence 6.Sentence 7."
]

def test_run_with_multiple_source_ids(self, in_memory_doc_store):
docs = [
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,10 @@ async def test_run_async_custom_fields(self, in_memory_doc_store):
# run the retriever with a document whose content = "Sentence 4."
result = await retriever.run_async(retrieved_documents=[doc for doc in docs if doc.content == "Sentence 4."])
assert len(result["context_documents"]) == 7
assert [doc.meta["split_id_test"] for doc in result["context_documents"]] == [1, 2, 3, 4, 5, 6, 7]
assert result["context_windows"] == [
"Sentence 1.Sentence 2.Sentence 3.Sentence 4.Sentence 5.Sentence 6.Sentence 7."
]

@pytest.mark.asyncio
async def test_run_async_with_multiple_source_ids(self, in_memory_doc_store):
Expand Down
Loading