fix(XLSXToDocument): keep cells whose text is a default NA string - #12951
mukktinaadh wants to merge 1 commit into
Conversation
`XLSXToDocument` read sheets with `pandas.read_excel` and pandas' default NA-string parsing, so text values such as "NA" or "N/A" were turned into missing cells. Those can be real content, like a region code or an explicit status, and once the document is produced the original text is gone. Only genuinely empty cells count as missing now, via `na_values=[""]`, so the existing missing-value handling (`missingval` for markdown, empty fields for csv) still applies to blank cells. `read_excel_kwargs` is applied after these defaults, so passing `keep_default_na=True` restores the previous behaviour. Fixes deepset-ai#12945
|
@mukktinaadh is attempting to deploy a commit to the deepset Team on Vercel. A member of the Team first needs to authorize it. |
|
|
|
Hi @mukktinaadh, thanks a lot for your contribution! 🙏 We noticed that the Contributor License Agreement (CLA) check ( To get your PR reviewed, please sign the CLA via the link in the |
|
Hi @mukktinaadh, just a friendly reminder: this PR is still in draft because the Contributor License Agreement (CLA) hasn't been signed yet. We'd love to review your contribution! Please sign the CLA via the link in the |
|
Thank you for your efforts! We're closing this PR because it makes the same |
Fixes #12945.
Problem
XLSXToDocumentreads sheets throughpandas.read_excelwith pandas' default NA-stringparsing, so cells holding text like
NAorN/Abecome missing values:Those are real values (a region code, an explicit status), and the original text cannot be
recovered from the resulting Document.
Change
Two keys in the
read_exceldefaults:keep_default_na=False— stop parsing NA strings as missing.na_values=[""]— keep only genuinely empty cells as missing, so the existingmissing-value handling still applies to them.
They are placed before
**self.read_excel_kwargs, so a caller passingread_excel_kwargs={"keep_default_na": True}still gets the old behaviour.After:
Why both keys
keep_default_na=Falseon its own regresses an existing test. In that configuration blankcells become
"", sotable_format_kwargs={"missingval": "N/A"}no longer reaches them andtest_run_markdown_missing_valuefails. Addingna_values=[""]keeps blank cells asNaN,which is what
missingvalrelies on:Tests
Four tests added to
test/components/converters/test_xlsx_to_document.py, covering csv andmarkdown output, that genuinely empty cells are still missing, and that
read_excel_kwargscan restore the previous parsing.
Against the parent commit, three of them fail:
After the change the file is
29 passed, with the pre-existingtest_run_markdown_missing_valuestill green.Also checked:
ruff checkandruff format --checkclean at ruff 0.16.0 (the pinnedruff-pre-commitrev), and a release note added underreleasenotes/notes/.