fix: keep a CSV column named row_number in CSVToDocument row mode - #12882
burakeyler wants to merge 1 commit into
Conversation
|
@burakeyler is attempting to deploy a commit to the deepset Team on Vercel. A member of the Team first needs to authorize it. |
Signed-off-by: Burak Eyler <burakeyler@gmail.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
|
Hi @burakeyler, thanks for your interest in contributing to Haystack! 🙏 ⛔ First-time contributors can have at most 1 open pull request in this repository until it has been approved, so this PR was closed automatically. Your open pull request #12879 is unaffected. Once it has been approved by a maintainer, you are welcome to open more PRs. Feel free to reopen this one at that point. See the contributing guidelines for details. This is an automated message to help us keep the review queue healthy. |
|
Hi @burakeyler, thanks for your interest in contributing to Haystack! 🙏 This is an automated message to help us keep the review queue healthy. |
Related Issues
Fixes #12787
Proposed Changes
_build_document_from_rowmerged the CSV columns into the metadata first and wrote the generatedrow_numberafterwards, so a CSV column of that name was silently dropped:The generated row number is now reserved before the merge, so the existing collision handling gives the CSV column the
csv_prefix like any other clashing column. Choosingrow_numberas the content column is unaffected, since the content column is skipped during the merge.How did you test it?
Three tests added: the ordinary case, a case where
csv_row_numberandcsv_row_number_1are already taken (so the next suffix is used), and a control withcontent_column="row_number". The first two fail onmain.pytest test/components/converters/test_csv_to_document.py→ 20 passed.Checklist
🤖 Generated with Claude Code