Read CSV results for the pandas PyArrow engine with pyarrow.csv - #1057
Conversation
pandas.read_csv(engine="pyarrow") builds its pyarrow ParseOptions without newlines_in_values and does not let callers set it, so a quoted value containing a newline that crosses one of pyarrow's 1 MiB read blocks made PandasCursor raise "CSV parser got out of sync with chunker" or return wrong values. Read the file with pyarrow.csv and newlines_in_values=True, and convert the table as pandas' PyArrow engine does for the options PandasCursor sets. The PyArrow engine is now used only when keep_default_na is False and the read_csv() options given to execute() are limited to dtype (a mapping), parse_dates, and storage_options; other options use the C engine. Document that PolarsCursor with use_pyarrow=True has the same limitation, which comes from pl.read_csv(). Closes #1048 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| return df | ||
|
|
||
|
|
||
| def _read_csv_with_pyarrow(source: str | IOBase, read_csv_kwargs: dict[str, Any]) -> DataFrame: |
There was a problem hiding this comment.
Self-review round one (implementation behavior): base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, head 5339aff7e2993ebbc72bd0385891a3588bdb1b1d.
Covered: _read_csv_with_pyarrow() compared step by step with pandas 3.0.6 ArrowParserWrapper.read(), arrow_table_to_pandas(), _post_convert_dtypes(), _maybe_convert_string_to_object(), and date_converter(); the _get_csv_engine() conditions and their only caller, _read_csv(); source opening for the PyArrow path (self._fs.open without storage_options, fsspec with them, even None); binary/converter columns (never on this path); .txt results (header-less, names); the auto-chunk and as_pandas() paths (the PyArrow engine is never chunked); sync, async, and aio pandas cursors (shared result set).
Edge cases compared with pandas.read_csv(engine="pyarrow") outside the suite: a header-only file, an all-NULL date column, an all-NULL integer column, and "" versus NULL all gave equal frames; dtype as a single type or None made pandas raise, so the engine condition now requires a mapping and the function assumes one.
Result: FINDINGS
- Duplicate column names (
SELECT 1 AS x, 2 AS x): master's PyArrow engine silently returned{'x': [2]}, and_read_csv_with_pyarrow()raisesAttributeError(OperationalErrorto the caller) becausedf[column]is a DataFrame. Deferred to Keep columns with the same name, and key CSV column types by the reader's column labels #1050, which adds a duplicate-name condition to thisis_compatibleexpression that sends such results to the C engine; whichever PR merges second takes both conditions. The PR body says so.
Simplifications already applied before publishing: the reader is a module function (it uses no result-set state), and the non-mapping dtype branches were removed in favor of the engine condition.
Validation: tests/pyathena/pandas/ + tests/pyathena/aio/pandas 327 passed live; the new live test fails on the pandas route with CSV parser got out of sync with chunker.
| so s3fs is also not required. | ||
|
|
||
| Options passed to `execute()`, other than PolarsCursor's own options such as `chunksize`, are passed to the Polars read function. | ||
| `pl.read_csv()` reads the CSV result with pyarrow when `use_pyarrow=True` is given with `schema_overrides=None`, which replaces PyAthena's column types, and Polars' other conditions for it hold. |
There was a problem hiding this comment.
Self-review round two (claims, compatibility, operations): base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, head e272719123fce03c291f22f015ba00ff04b8eb0d.
Claims checked:
- pandas builds
ParseOptionsfromdelimiter,quote_char,escape_char, andignore_empty_linesonly: confirmed in pandas 3.0.6arrow_parser_wrapper.py, which also mapson_bad_linestoinvalid_row_handler. The PR text now names it. - pandas' default NA strings are private:
STR_NA_VALUESlives inpandas._libs.parsersand is not exported bypandas.io.parsers, hence thekeep_default_na=Falsecondition. - Silent wrong values: reproduced 19 of 5,000 with
pyarrow.csvdirectly (newlines_in_values=False), and all were correct withTrue. - BUG: read_csv with pyarrow engine treats newlines in quotes as newline when they should be ignored pandas-dev/pandas#56396 and #59009: open (GitHub search, 2026-10-04).
- Performance: 0.58 s versus 0.63 s (pandas PyArrow) and 1.49 s (C) on 70 MiB locally; this measures CPU only, not S3 transfer.
docs/polars.mdsaid "withuse_pyarrow=Trueandschema_overrides=None" as if those two were sufficient. Polars 1.44.2 also requiresn_rows,n_threads, andnull_valuesto be unset andlow_memory=False, so the text now namesschema_overrides=Noneand "Polars' other conditions" (e272719). Live,PolarsCursorwith those options raisesOperationalError: CSV parser got out of sync with chunkeron the 2,000-row query, while the default reader returns all rows correctly.docs/pandas.md: the engine conditions match_get_csv_engine(), and the conversion sentence matches the function.
Existing callers: results without multi-line values give frames equal to before (assert_frame_equal(check_exact=True) over every eligible type, infer_string on and off, and 200 random files). Callers that passed other read_csv options or keep_default_na=True with engine="pyarrow" now get the C engine, which is listed as a release-note item. No AWS requests are added (one GET of the result file, as before).
Result: CLEAN after the wording corrections. The duplicate-name interaction from round one stays deferred to #1050.
The independent review found that _read_csv_with_pyarrow() diverged from pandas.read_csv(engine="pyarrow"), and that sending other cases to the C engine broke callers: - pandas applies the dtype mapping again after parsing dates, so a dtype entry for a parse_dates column wins; apply it again. - Duplicate column names made df[column] a DataFrame; look up a column's dtype only when it has no dtype entry, as pandas does, and convert strings to objects by position. - Numeric na_values, keep_default_na, storage_options (with pandas' anonymous retry), and PyArrow-only options such as a callable on_bad_lines are not reproduced, and the C engine rejects some of them. Keep the engine selection unchanged, and read the file with _read_csv_with_pyarrow() only when keep_default_na and na_values are the defaults and execute() passes no read_csv() options other than a dtype mapping and parse_dates; pandas reads it otherwise, as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| return df | ||
|
|
||
|
|
||
| def _read_csv_with_pyarrow(source: str | IOBase, read_csv_kwargs: dict[str, Any]) -> DataFrame: |
There was a problem hiding this comment.
Independent review (relayed): Codex CLI 0.160.0, model gpt-6-astra, codex exec -s read-only --ephemeral, session 01a1046c-5f68-7ce1-8882-f37d65fb08a4. Base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, head e272719123fce03c291f22f015ba00ff04b8eb0d, detached snapshot (unchanged afterwards). The prompt held the literal diff and repository conventions plus read-only access to the installed pandas, pyarrow, and Polars sources; it held no PR text or prior findings. This was a static review: no tests, Python, builds, or network access.
Reviewer coverage: the five-file diff, CSV/TSV conversion, engine selection, dtype/date/NULL handling, sync and async callers, stream ownership, fsspec options, tests, and the pandas/Polars docs, compared with pandas 3.0.6, pyarrow 25.0.1, and Polars 1.44.2.
Verdict: FINDINGS (5 × P2 regressions)
result_set.py:135: pandas appliesdtypeagain after parsing dates (_finalize_pandas_output()), sodtype={"d": "string"}, parse_dates=["d"]returns strings on master and datetimes here.result_set.py:107: with duplicate column names,df[column].dtypehits a DataFrame and raises (OperationalError); pandas skips that lookup when the name has a dtype entry and returns both columns.result_set.py:73: pandas expands numericna_values("-1"also matches-1.0), and a NumPy array of markers raises onor [].result_set.py:508: sending other options to the C engine breaks PyArrow-only options, e.g. a callableon_bad_linesis rejected by the C engine.result_set.py:702: the direct fsspec open withstorage_optionsdrops pandas' retry withanon=True(pandas/io/common.py).
The reviewer also judged the DataFrame-equivalence sentence indocs/pandas.mdtoo broad until these are fixed.
Author verification: all five confirmed. Findings 1–4 were run against pandas 3.0.6: findings 1 and 3 reproduce as described, master returns ['x', 'x'] with both values for duplicate names, and the C engine raises on_bad_line can only be a callable function if engine='python' or 'pyarrow'. Finding 5 was confirmed in the pandas source, not run. My round-one record was wrong about duplicate names: it read master's result through to_dict(), which collapses duplicate keys.
Repair in 84e2fca, recorded in the reply below.
There was a problem hiding this comment.
Repair 84e2fca (old head e272719123fce03c291f22f015ba00ff04b8eb0d → new head 84e2fca36c98d8a6306d1ef3ecfd0a9a591e51c8, same base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1)
- Findings 3–5 led to a routing change.
_get_csv_engine()is back to master._read_csv()calls_read_csv_with_pyarrow()only when the new_reads_csv_with_pyarrow()holds:keep_default_nais false,na_valuesis the default("",)as a list or tuple, andexecute()passes only adtypemapping and/orparse_dates. Every other case goes topandas.read_csv(engine="pyarrow")as on master, so numericna_values,keep_default_na,storage_options(with its anonymous retry), and PyArrow-only options such as a callableon_bad_linesbehave as before; no case moves to the C engine any more. Thestorage_optionsopen branch was removed, and thequoting/converterspops for the pandas PyArrow call are restored. - Finding 1: the
dtypemapping (including the NumPy integer entries) is applied again after date parsing, as_finalize_pandas_output()does. - Finding 2: the integer lookup checks
column not in dtypefirst, as_post_convert_dtypes()does, and the string-to-object loop works by position withisetitem(). - Tests:
test_reads_csv_with_pyarrowreplaces the engine-selection test; new parity cases for adtypeentry on a date column and for duplicate names fail on e272719 and pass now; the live test no longer has astorage_optionscase (pandas reads that one).
Self-review of the repair, round one (behavior): parity harness all equal (types, TSV, infer_string off, 200 random files, user dtype/parse_dates); the second astype() mirrors pandas, including its cost; an ndarray na_values is not a list/tuple, so pandas reads it as before. Validation: just lint passed; offline 22 targeted tests passed; tests/pyathena/pandas/ + tests/pyathena/aio/pandas 331 passed live.
Round two (claims): the PR text no longer lists C-engine fallbacks as release-note items. Its only behavior change is the multi-line fix on the default path. docs/pandas.md keeps master's engine conditions and adds when PyAthena reads the file and that pandas' engine still has the limitation otherwise, which also scopes the DataFrame-equivalence sentence the reviewer flagged. The Polars limitation was reported upstream as pola-rs/polars#29717, reproduced on Polars 1.44.2 and main 9ee0dc5b81.
The independent follow-up on this repair is pending.
There was a problem hiding this comment.
Independent follow-up 1 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a10488-30bd-78b0-98b8-37d67ec7ed57. This was a full review of 84e2fca36c98d8a6306d1ef3ecfd0a9a591e51c8 (base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, detached snapshot unchanged), static only.
The reviewer found the five earlier findings resolved: dates followed by a final cast; duplicate names via short-circuit and positional conversion; non-default NA settings, other options, and storage_options all delegated to pandas; master's engine selection restored.
Verdict: FINDINGS
- P2
result_set.py:106:dtypemapping values reachedastype()raw. pandas normalizes them withpandas_dtype(), sodtype={"x": None}gives float64 with NaN in pandas and nullable Int64 here. - P3
result_set.py:124:parse_dates="d"or("d",)was silently ignored by the helper, while pandas raisesTypeError: Only booleans and lists are accepted for the 'parse_dates' parameter.
The reviewer also noted, as pre-existing (pandas itself fails): duplicate names without a dtype entry, and duplicate date-column names.
Author verification: both reproduced against pandas 3.0.6 (float64 vs Int64; pandas TypeError for a string and a tuple).
Repair fbceb11b: the retained values go through pd.api.types.pandas_dtype(), and _reads_csv_with_pyarrow() requires parse_dates given to execute() to be a list, so other forms reach pandas (and its TypeError) as before. The helper now iterates parse_dates directly. docs/pandas.md says "parse_dates as a list". New cases: parity dtype_none, and the predicate's parse_dates_string and parse_dates_tuple. All 4 fail on 84e2fca and pass now.
Self-review of the repair: behavior: pandas_dtype() is the public form of what _post_convert_dtypes() applies, the parity harness is all equal, and lint passed. Claims: the docs and PR text name the list requirement. No AWS calls change. The narrow independent follow-up on 84e2fca3..fbceb11b is running.
There was a problem hiding this comment.
Independent follow-up 2 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a1049b-7b5b-7670-991a-0c16591db2ae. This was a narrow review of 84e2fca36c98d8a6306d1ef3ecfd0a9a591e51c8..fbceb11b894f3b208554bc53ba86af3e87545119 (detached snapshot unchanged), static only.
Verdict: CLEAN. Both points are resolved, and the reviewer found no new defect.
result_set.py:107normalizes the values as pandas does (None→ float64); the regression compares values and dtypes withinfer_stringon and off.result_set.py:731admits only a listparse_dates; strings and tuples reach pandas unchanged and raise itsTypeError, which_read_csv()wraps inOperationalError. The routing cases and the docs match.
Validation at fbceb11: tests/pyathena/pandas/ + tests/pyathena/aio/pandas 335 passed live.
There was a problem hiding this comment.
Merge of master (15399b5): master gained #1044 (_JSONConverter, managed integer dtypes) and other PRs since base e49d576c. The only conflict was in pyathena/pandas/result_set.py, where both sides added a module-level definition at the same place, so both were kept unchanged.
Upstream contracts checked: _read_csv() now sets self._csv_converters before reading and restores json NULLs afterwards. The PyArrow engine is chosen only without converters, so that set is empty on this path, and the _read_csv_with_pyarrow() dispatch is unchanged.
Validation after the merge: just lint passed; the offline pandas result-set and engine/routing tests passed (39); the parity harness is all equal. AWS CI reruns on this push.
…w reader pandas normalizes dtype mapping values with pandas_dtype(), so None means float64, and rejects a parse_dates value other than a boolean or a list. Normalize the values the same way, and let pandas read the file when parse_dates given to execute() is not a list. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ltiline-values # Conflicts: # pyathena/pandas/result_set.py
pandas' PyArrow engine names the columns of a header-less file beyond the given names by their positions, while _read_csv_with_pyarrow() assigned the names directly and raised a length mismatch. Build the parity tests' options with the real _get_csv_read_options(), and cover the extra fields and a parse_dates column that does not parse. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| df = table.cast(schema).to_pandas(types_mapper=integer_dtypes.get) | ||
| if header is None: | ||
| # pandas names the columns beyond the given names by their positions. | ||
| df.columns = [str(index) for index in range(len(df.columns) - len(names))] + names |
There was a problem hiding this comment.
Additional review requested by the maintainer (relayed): Claude Code 2.1.287, model claude-fable-5-1, effort high, profile max (first-party claude.ai Max, not Enterprise), session 4a0ff4d9-d1d0-4569-9085-d3f6f96c5219. Snapshot 15399b55355736b8c4bf6711494e61e4f4063b78 (detached, unchanged afterwards), base a57325b1fbc911d928bbe2c7e7054f6bdcb42c92. Tools were limited to Read, Grep, Glob, and git show/diff/log; the pandas, pyarrow, and Polars sources were readable. The prompt held the literal diff only, with no PR text or prior findings. This was a static review. Fable is another Claude model, so this complements the Codex reviews above rather than replacing them.
Reviewer coverage: the helper, the predicate and _read_csv(), _get_csv_read_options(), _get_csv_engine(), _finish_csv_frame(), and _trunc_date(), compared step by step with pandas 3.0.6 (ArrowParserWrapper, arrow_table_to_pandas, _post_convert_dtypes, _maybe_convert_string_to_object, date_converter, _clean_na_values, pandas_dtype); the Polars 1.44.2 use_pyarrow branch; the tests and docs.
Verdict: FINDINGS (all low, no blocker)
- Low regression: for a header-less (
.txt) result with more fields thannames, pandas prefixes"0","1", … (_adjust_column_names), whiledf.columns = namesraised a length mismatch. - Low test gap: the
to_datetimefailure fallback was not covered. - Low test design:
_pyarrow_read_csv_kwargs()re-implemented_get_csv_read_options(), so a change to the builder would go unnoticed. - Info, docs:
newlines_in_values=Truereduces pyarrow's multi-threaded parsing, so the reviewer suggested a performance note.
Pre-existing, as noted by the reviewer: withfuture.infer_string=False,dtype=strturns NULL into a string in pandas itself; duplicate names without a dtype entry fail in pandas too.
Author verification and repair 2302049:
- Reproduced offline: pandas gives
['0', '1', 'col_name'], the helper raised. The reviewer'sDESCRIBEscenario does not occur live: on AthenaDESCRIBEdescribes three columns (col_name,data_type,comment), and the helper read it correctly (5 rows). A header-less result whose values contain tabs would still diverge, so the helper now prefixes positional names as pandas does. The new casetab_separated_extra_fieldsfails on 15399b5. - The fallback already matched pandas. Added the
unparsed_datescase. - The parity tests now build their options with the real
_get_csv_read_options()on anAthenaPandasResultSetwith patcheddescription/output_location, and assert_reads_csv_with_pyarrow(). - Not applied. Measured locally on 70 MiB, the helper with
newlines_in_values=Truetakes 0.58 s versus 0.63 s forpandas.read_csv(engine="pyarrow"), so the documented path is not slower than before; the docs state usage facts only.
Pre-existing note verified: withinfer_string=Falseanddtype=str, pandas' PyArrow engine returns'nan'strings for NULL (the C engine gives NaN). The helper reproduces this; it is out of scope here.
Self-review of the repair: behavior: the prefix mirrors_adjust_column_names(PyAthena always passes names for.txt);just lintpassed;test_result_set.py32 passed. Claims: none of the PR text changes except the test list. AWS CI reruns on this push.
There was a problem hiding this comment.
Independent follow-up on the repair (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a104da-610d-7211-b8d2-395579d34123. This was a narrow review of 15399b55355736b8c4bf6711494e61e4f4063b78..2302049431ff2ee9d2b49ba271ae634cc5252a10 (detached snapshot unchanged), static only.
Verdict: CLEAN. All three points are resolved, and the reviewer found no new defect.
result_set.py:99matches pandas' positional prefix; the caller always supplies names for header-less results, and the new case has three fields with one name.test_result_set.py:228reaches theto_datetime()failure path with"plain"under both string-inference settings.test_result_set.py:190uses the real eligibility check and option builder; the expected read still goes through pandas' PyArrow engine (engine="pyarrow"comes from the builder), and the separate dtype copies keep the comparison independent.
AWS CI at 2302049: test / run (3.14) passed (test-sqla* are skipped by the path filter, as pandas-only changes).
WHAT
With
engine="pyarrow",PandasCursornow reads the CSV result withpyarrow.csvitself, withnewlines_in_values=True, instead of callingpandas.read_csv(engine="pyarrow"), when_reads_csv_with_pyarrow()holds:keep_default_naisFalseandna_valuesis PyAthena's default("",);execute()passes no pandas.read_csv() options other thandtypeas a mapping andparse_datesas a list.The new
_read_csv_with_pyarrow()converts the table as pandas' PyArrow engine does for those options: NULL-typed columns become float64; integer columns without adtypeentry get NumPy integer types; thedtypemapping, normalized withpandas_dtype(), is applied before and after theparse_datescolumns go throughpandas.to_datetime()on their string values; and withfuture.infer_stringdisabled, strings and string categories become objects.In all other cases pandas reads the file as before, and the engine selection in
_get_csv_engine()is unchanged.docs/pandas.mddescribes when PyAthena reads the file and the remaining limitation of pandas' PyArrow engine.docs/polars.mdnow says thatexecute()options reach the Polars read function and thatpl.read_csv()has the same limitation whenuse_pyarrow=Truewithschema_overrides=Nonemakes it read with pyarrow; the PolarsCursor code is unchanged (reported upstream as pola-rs/polars#29717).Release note: with
engine="pyarrow"and default read options, multi-line values that cross a pyarrow read block are read correctly; previously this raisedOperationalErroror returned wrong values.WHY
Closes #1048.
pandas builds
pyarrow.csv.ParseOptionsfromdelimiter,quote_char,escape_char, andignore_empty_lines(plusinvalid_row_handlerforon_bad_lines) only, so pyarrow's chunker splits its 1 MiB blocks at any newline, including one inside a quoted value (pandas-dev/pandas#56396, pandas-dev/pandas#59009 are open).That raises
CSV parser got out of sync with chunker, or returns wrong values without an error: reading 5,000 randomized small files with quoted multi-line values throughpyarrow.csv.read_csv()withnewlines_in_values=Falseat various block sizes, 19 returned the right row count with wrong values (for example[' ', '', ',x"']for a multi-line row), whilenewlines_in_values=Trueread them all correctly.The maintainer chose reading with
pyarrow.csvover limiting the PyArrow engine to results without string-like columns, and documentation plus an upstream report for Polars'use_pyarrow, a passthrough option.Cases that
_read_csv_with_pyarrow()does not reproduce stay with pandas rather than moving to the C engine, which rejects PyArrow-only options such as a callableon_bad_lines. These cases are numericna_values(pandas expands"-1"to"-1.0"),keep_default_na=True(pandas' default NA strings are private),storage_options(pandas retries a failed open withanon=True), and other read_csv() options.Measured locally on a 70 MiB CSV of 1,000,000 rows (bigint, double, varchar, date, timestamp; best of 3):
pandas.read_csv(engine="pyarrow")0.63 s,_read_csv_with_pyarrow()0.58 s,engine="c"1.49 s.TEST
Tested commit: fbceb11 (live suite); 2302049 adds a header-less naming fix and test changes, validated offline and by AWS CI.
just lint: passed.just docs lint: 0 errors.uv run --env-file .env pytest -n 2 tests/pyathena/pandas/ tests/pyathena/aio/pandas: 335 passed (live Athena, 2026-10-04).TestPandasCursor.test_pyarrow_engine_multiline_values_across_blocksreads 2,000 values of 1,201 bytes (a newline and quotes each, 2.4 MB) withengine="pyarrow"and checks that_read_csv_with_pyarrow()was used. Routed topandas.read_csv()as on master, it fails withCSV parser got out of sync with chunker.test_result_set.pybuild their options with the real_get_csv_read_options()and compare_read_csv_with_pyarrow()withpandas.read_csv(engine="pyarrow")(assert_frame_equal,check_exact=True), each withfuture.infer_stringon and off. The inputs are an Athena CSV result with every type the PyArrow engine reads plus a NULL row, a partialdtypeoverride (float32, category, a missing column), positionalparse_dates, adtypeentry for a date column,dtype={"x": None}with NULLs, duplicate column names, aparse_datescolumn that does not parse, a header-less result with more fields than names, and a tab-separated result. The date-column and duplicate-name cases fail on the pre-review implementation e272719, and theNonecase on 84e2fca. The tests also read 2.4 MB of multi-line values across the default 1 MiB blocks.test_reads_csv_with_pyarrowcovers which options PyAthena reads itself, includingparse_datesas a string or tuple (pandas reads and rejects those).PolarsCursorwithuse_pyarrow=True, schema_overrides=Noneon the same 2,000-row query raisesOperationalError: CSV parser got out of sync with chunker, and the default reader reads all rows (the fact stated indocs/polars.md).🤖 Generated with Claude Code