-
Notifications
You must be signed in to change notification settings - Fork 116
Keep NULL rows of single-column CSV results in ArrowCursor #1031
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -308,7 +308,8 @@ def _read_csv(self) -> Table: | |
| parse_opts = csv.ParseOptions( | ||
| delimiter=",", | ||
| quote_char='"', | ||
| ignore_empty_lines=not binary_columns, | ||
| # Athena writes a single-column row with a NULL value as an empty line. | ||
| ignore_empty_lines=False, | ||
|
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Self-review round one (implementation behavior) Base Result: CLEAN.
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Independent review (relayed result; static review) Reviewer: Codex CLI 0.160.0, model Covered: sync, threaded async, and aio Arrow callers; table/fetch conversion; single/multi-column and binary results; explicit/inferred dtypes; all-NULL and header-only CSV; CRLF, quoted newlines, Result: FINDINGS (2, both pre-existing; no regression in the changed line's own behavior).
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Correction to finding 2:
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Disposition update for finding 2: filed as #1040, reproduced on Athena with |
||
| double_quote=True, | ||
| escape_char=False, | ||
| ) | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Self-review round two (claims, callers, operations)
Base
de8cc52ac2a43cdba72bb4571c185883287741fc, head6b924fe7161e6d0e4b61bc601beeadb78f075793.Result: FINDINGS (PR description only; no code change).
Claims checked:
git show v2.7.0:pyathena/arrow/result_set.pybuilds the CSVParseOptionswithoutignore_empty_lines(pyarrow defaultTrue). Holds.varbinarycolumns — incorrect as written:ignore_empty_linesis absent in v3.34.0–v3.36.0 and present from v3.37.0 (and on3.x). The release note now says v3.37.0 kept the rows only forvarbinarycolumns and that3.xhas the same code.pyathena/arrow/async_cursor.py:166andpyathena/aio/arrow/cursor.py:177constructAthenaArrowResultSet. Holds; their tests ran in the 106-test local run, but the new regression test is sync only.''— matchesdocs/null_handling.md(ArrowCursor CSV row); no documentation becomes obsolete.Callers/operators: no API, default, or exception change; no extra AWS requests (one CSV read as before). Callers that relied on dropped rows now see the NULL rows, which is the fix.