Skip to content

Milestone 4A: structure-aware chunking - #1

Merged
princerm06 merged 3 commits into
mainfrom
milestone-4a-structure-aware-chunking
Oct 5, 2026
Merged

princerm06 merged 3 commits into
mainfrom
milestone-4a-structure-aware-chunking

Conversation

@princerm06

Copy link
Copy Markdown
Owner

Summary

  • replace fixed character slicing with paragraph-aware chunk packing
  • split oversized paragraphs at sentence boundaries, with word-boundary fallback
  • preserve structural boundaries and use complete trailing blocks for overlap
  • validate invalid chunking configurations
  • add focused unit tests for boundaries, oversized paragraphs, overlap, and configuration

Why

The previous chunker cut text every 1,000 characters, which produced messy search excerpts and could split concepts mid-sentence. This change makes retrieval chunks substantially cleaner without hiding retrieval quality behind an LLM.

Testing note

The tests are committed under backend/tests/test_chunking.py. The repository currently does not include pytest in its pinned requirements, and an attempted dependency-file update was blocked by the tool safety layer, so the dependency change is intentionally not included in this PR yet.

@princerm06
princerm06 marked this pull request as ready for review October 5, 2026 09:00
@princerm06
princerm06 merged commit 2dbe885 into main Oct 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant