Read string columns as text for the pandas PyArrow engine - #1065
Conversation
pandas' PyArrow engine lets pyarrow infer a type for each column and only
then applies dtype=str, so a VARCHAR column whose values look like
numbers came back changed ("007" as "7.0", "1e3" as "1000.0", the text
"nan" as NaN), and without future.infer_string, astype(str) turned NULL
into the string "nan". _read_csv_with_pyarrow() reproduced both.
Read the columns with a string dtype as pyarrow strings, and give the
columns with a NumPy str dtype object values that keep missing values
missing, so these columns have the C engine's values. pandas' own
PyArrow engine, used for options PyAthena does not read itself, keeps
pandas' behavior (pandas-dev/pandas#57666), which the docs now state.
Closes #1062
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A dtype mapping can have keys that are not column names, such as positions, which pandas ignores; pyarrow's column_types raised TypeError for them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| f"f{index}": pa.string() for index, name in enumerate(names) if name in string_columns | ||
| } | ||
| else: | ||
| # pandas ignores dtype keys that are not column names, such as positions. |
There was a problem hiding this comment.
Self-review round one (implementation behavior): base 8bbf8cb5e2d566bfe3dda86273b498168cfe4360, head a3fcbb3e → repaired in 51129926.
Covered: _read_csv_with_pyarrow() (string-column detection, column_types for header and header-less files, the NumPy-str path, both astype() calls, the infer_string=False object conversion, date parsing); its callers _read_csv() and _reads_csv_with_pyarrow(); DefaultPandasTypeConverter dtypes; the C engine as the reference for string columns.
Checked outside the suite (pandas 3.0.6, pyarrow 25.0.1): string columns equal the C engine's with infer_string on and off; a dtype entry for a missing column is ignored by pyarrow; duplicate names both read as text; a header-less .txt keeps "007".
Result: FINDINGS
- A
dtypekey that is not a column name (e.g.{0: str}, which pandas ignores) made pyarrow'scolumn_typesraiseTypeError: expected bytes, int found. Repaired in 5112992 (onlystrkeys go tocolumn_types) with thedtype_position_keyparity case.
Known limit, not changed: for a header-less result with more fields than names, which pandas names from the end, column_types maps names from the start, so the shifted columns keep inferred types. This needs a tab inside a value of a DDL text result.
Validation: lint passed; test_result_set.py 36 passed. Live runs are listed in the PR body (one unnamed failure in the first run, which did not recur in --lf or a full rerun).
| With the PyArrow engine, when `keep_default_na` and `na_values` are the defaults and the only pandas.read_csv() options passed to `execute()` are `dtype` as a mapping and `parse_dates` as a list, PyAthena reads the file with `pyarrow.csv`, allowing newlines in quoted values, and converts it to a DataFrame as `pandas.read_csv(engine="pyarrow")` does. | ||
| With the PyArrow engine, when `keep_default_na` and `na_values` are the defaults and the only pandas.read_csv() options passed to `execute()` are `dtype` as a mapping and `parse_dates` as a list, PyAthena reads the file with `pyarrow.csv`, allowing newlines in quoted values, and converts it to a DataFrame as `pandas.read_csv(engine="pyarrow")` does, except that columns with a string dtype keep their text and their missing values, as with the C engine. | ||
| Otherwise, pandas reads the file, and its PyArrow engine raises an error or returns wrong values when a quoted value containing a newline crosses one of pyarrow's read blocks. | ||
| It also changes the values of a column with a string dtype whose values all look like numbers, such as `"007"` to `"7.0"`, and without `future.infer_string`, turns NULL in a column with a string dtype into the string `"nan"`. |
There was a problem hiding this comment.
Self-review round two (claims, compatibility, operations): base 8bbf8cb5e2d566bfe3dda86273b498168cfe4360, head 487d88ef71399f48639da61af83eca8a9b806822.
Claims checked:
- "pandas does not expose it" (BUG: pyarrow read_csv engine stripping leading zeros with dtype=str pandas-dev/pandas#57666) was too strong, because the issue is closed upstream. The PR now states that it is reported there and still reproduces with pandas 3.0.6 / pyarrow 25.0.1 (measured).
Series.astype(str)turningNaNinto"nan"withinfer_string=Falsewas measured for StringDtype and object columns; "long-standing" was removed as unverified.- PyAthena maps
char/varchar/string/array/map/rowtostr: confirmed inDefaultPandasTypeConverter._dtypes. - The docs said the PyArrow engine "changes strings that look like numbers"; it happens only when the column's values all look like numbers (pyarrow infers per column), and the NULL case only in string-dtype columns. The wording was narrowed (487d88e).
Existing callers: string-dtype columns change only where master returned changed values (numeric-looking text, the text "nan" read as NaN, NULL as "nan"). All other columns convert exactly as before, as the parity tests confirm. Results that pandas reads itself are unchanged. No AWS requests change.
Result: CLEAN after the wording corrections.
…yArrow reader The independent review found that the string-column change: - applied the NumPy str handling to Arrow string dtypes, whose kind is also "U", returning objects instead of the requested dtype; - typed the first fields of a header-less file while pandas gives the names to the last ones, changing other columns' types when a row has extra fields; - left parse_dates columns with a NumPy str dtype as datetimes, while the dtype used to apply again after parsing; - failed on invalid dtype entries for columns not in the result, which pandas ignores. Apply the dtype mapping through _astype_keeping_missing_strings() before and after parsing dates, limit its NumPy str handling to NumPy dtypes, count the fields of a header-less file's first line to type the named ones, and skip dtype values that do not resolve when choosing string columns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| return df | ||
|
|
||
|
|
||
| def _read_csv_with_pyarrow(source: IOBase, read_csv_kwargs: dict[str, Any]) -> DataFrame: |
There was a problem hiding this comment.
Independent review (relayed): Codex CLI 0.160.0, model gpt-6-astra, codex exec -s read-only --ephemeral, session 01a10518-1ab0-7b41-a3dd-4026904daeb3. Base 8bbf8cb5e2d566bfe3dda86273b498168cfe4360, head 487d88ef71399f48639da61af83eca8a9b806822, detached snapshot (unchanged afterwards). The prompt held the literal diff and conventions only, with no PR text. This was a static review.
Reviewer coverage: the four-file diff, sync/Future/asyncio PandasCursor callers, engine selection, CSV/TSV options, date conversion, tests, docs, and the pandas/pyarrow sources (UNLOAD and the Arrow/Polars result sets do not call the helper).
Verdict: FINDINGS (4 × P2 regressions, 1 × P3)
result_set.py:137:pd.ArrowDtype(pa.string())/large_string()havekind == "U"too, so they got the NumPy-str object path instead of the requested Arrow dtype.result_set.py:87: header-less files typedf0…from the start, while names are right-aligned when rows have extra fields (001\t2\t003withnames=["v"]typedf0as text and leftvnumeric), changing other columns' output types.result_set.py:138: a NumPy-strparse_datescolumn that parses successfully stayed datetime; master re-applied the dtype after dates.result_set.py:81:pandas_dtype()on every mapping value made an invalid entry for an absent column ({{"unused": "not-a-dtype"}}) fail, while pandas ignores it.- P3
docs/pandas.md:559: the C-engine equivalence needs theparse_datesexception, and the"nan"warning applies only to NumPy str dtypes.
The reviewer suggested keeping the cast mapping separate from the pyarrow column types.
Author verification: 1, 3, and 4 reproduced (Arrow dtype → object with NaN, pandas gives string[pyarrow] with <NA>; dates → Timestamps, master gives ['2024-01-01', 'None']; TypeError: data type 'not-a-dtype' not understood). 2 was confirmed by the reviewer's example. It was my round-one "known limit", which I had judged rare, but it changes other columns, so it is repaired, not deferred. 5 accepted.
Repair 7bad8fe:
_astype_keeping_missing_strings()applies the normalized mapping, and only NumPy (non-extension) str dtypes become object strings with missing values kept. It runs before and after date parsing, so 1 and 3 follow master's dtype-after-dates order; NaT and NULL become NaN rather than master's'None'/'nan'.- Header-less files: the first line's field count gives the offset of the named fields (
readline()thenseek(0)on the opened stream)._read_csv()passes the helper only anIOBasesource, which the PyArrow path always opens. - String-column detection skips values that
pandas_dtype()rejects; validation stays on the result's columns. - Docs: string-dtype columns "are read as text", and the NULL wording names dtype
str(NumPy str withoutinfer_string). - Tests: parity cases
arrow_string_dtypesandtab_separated_numeric_fields, plustest_read_csv_with_pyarrow_string_dtype_after_datesandtest_read_csv_with_pyarrow_ignores_unused_dtype_entries(the C engine validates unused entries, so that case compares with pandas' PyArrow engine only). 6 of the new cases fail on 487d88e.
Self-review of the repair: behavior: test_result_set.py + routing tests 54 passed offline; live, DESCRIBE of a 6-column table through the header-less path read correctly, and test_cursor.py -k "pyarrow_engine or show_columns or describe" 8 passed. Claims: the docs and PR text now match the narrowed wording. A narrow independent follow-up on 487d88ef..7bad8fe5 is running.
There was a problem hiding this comment.
Independent follow-up 1 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a1052d-3908-7762-a8f3-6bb4fea80ce4. This was a narrow review of 487d88ef71399f48639da61af83eca8a9b806822..7bad8fe55571c5f22ea610c666e1db2b6a351f59, static only. The reviewer marked prior findings 1, 3, and 5 resolved, and 2 and 4 partially resolved.
Verdict: FINDINGS (2 × P2)
result_set.py:123: splitting the first line on the delimiter miscounts a quoted delimiter (007\t"a\tb"counts 3 fields, while pyarrow parses 2), sof1was typed instead off0andvlost"007"; a quoted newline can undercount.result_set.py:115:pandas_dtype("decimal128(10, 2)[pyarrow]")raisesNotImplementedError, which the detection did not catch, so an unused entry still failed.
Author verification: both confirmed (NotImplementedError for that string; the quoted-tab case typed the wrong field).
Repair d09bccd: the first record is read with csv.reader (same " quoting and doubling as the reader) over decoded readline() lines, then seek(0); an empty first record counts as one field. Detection skips values raising TypeError, ValueError, or NotImplementedError. New cases: tab_separated_quoted_tab and tab_separated_quoted_newline, plus the decimal128 entry in the unused-entries test. The quoted-tab case and the unused-entries test fail on 7bad8fe.
Self-review of the repair: lint passed; 58 offline tests passed; live DESCRIBE through the header-less path on S3File read correctly. A full live run at 7bad8fe (with these edits partly in the working tree) gave 360 passed. A clean full run at d09bccd and a narrow Codex follow-up are running.
There was a problem hiding this comment.
Independent follow-up 2 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a1053b-6e0a-7943-b4fa-de184b235565. This was a narrow review of 7bad8fe55571c5f22ea610c666e1db2b6a351f59..d09bccd4e9d655791a041ad9cbcd9e13fb108573, static only. The reviewer marked the absent-column dtype issue resolved.
Verdict: FINDINGS (3 × P2), all in the header-less field-count probe:
result_set.py:125:csv.readerlimits a field to 131,072 characters, so a long first field that pyarrow accepts now raises_csv.Error.result_set.py:124: bare-CR records (b"007\r010\r") arrive in one binaryreadline(), andcsv.readerraises "new-line character seen in unquoted field".result_set.py:124: a UTF-8 BOM before a quoted first field is kept by the decode, so the count disagrees with pyarrow (which strips it), and the wrong field is typed.
The reviewer also noted that the multiline test put the newline in the last field.
Decision (author), 6bd7711: these three come from counting fields before reading, which was needed only to type header-less fields. Each counting method so far (delimiter split, then csv.reader) has disagreed with pyarrow on some input, so further probing would keep widening this PR. Header-less files are the tab-separated results of DDL statements, so their fields now keep the types pyarrow infers, as on master, and the probe is gone. They keep only the NumPy-str missing-value handling (NULL stays missing without infer_string). Numeric-looking text in DDL results therefore still follows pandas' PyArrow engine; the docs say so ("other than in the tab-separated results of DDL statements").
Also in 6bd7711: the NumPy-str conversion happens once after date parsing, inline, instead of in the separate _astype_keeping_missing_strings() helper. Both astype() calls skip those columns, and date parsing goes through astype("string"), so the result is the same. The parity test compares header-less string columns with pandas' PyArrow engine, except that missing values follow the C engine. The two probe-only cases were removed.
Validation: lint passed; 54 offline tests passed; a clean live run at d09bccd gave 364 passed. A live run at 6bd7711 and a Codex follow-up on d09bccd4..6bd7711b are running.
There was a problem hiding this comment.
Independent follow-up 3 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a10548-f947-7573-b0bd-1cfbb7e48175. This covered d09bccd4e9d655791a041ad9cbcd9e13fb108573..6bd7711b3b5b51e3035a7755db53dbc072969291 plus the full diff, static only. The reviewer resolved the long-field, bare-CR, and BOM findings (the probe is gone) and found no defect in the single conversion after dates.
Verdict: FINDINGS (1 × P2)
result_set.py:180: in a header-less result withinfer_string=False, a field pyarrow infers as float keeps the literal textnanas a float NaN, and the newnotna()mask makes it missing like NULL; master returned the string"nan"for it.
Author verification and decision: deferred, no change. Measured on 1.5\nnan\n\n (names=["v"], dtype={"v": str}):
infer_string |
master / pandas PyArrow engine | this PR |
|---|---|---|
| True (default) | ['1.5', nan, nan] |
['1.5', nan, nan] |
| False | ['1.5', 'nan', 'nan'] |
['1.5', nan, nan] |
With the default setting, master already turns the literal into a missing value: pyarrow's float inference cannot tell it from NULL. With infer_string=False, master returned 'nan' for both the literal and NULL, so it did not keep them apart either. This PR makes infer_string=False match the default. Telling them apart would need Arrow's null bitmap carried through conversion, only for inferred header-less (DDL result) fields, which goes past this issue. Query results with a header read string columns as text and are not affected. |
Validation at 6bd7711: tests/pyathena/pandas/ + tests/pyathena/aio/pandas 360 passed live (clean working tree).
There was a problem hiding this comment.
Merge of master (1f6fc7a): master gained #1066 (#1051, duplicate column types) and other PRs. The merge had no conflicts.
Upstream contracts checked: _read_csv_header_as_labels() rewrites names/header/skiprows only for results with duplicate column names. Since #1050, _get_csv_engine() never selects the PyArrow engine for those (_needs_csv_column_name_resolution()), so _read_csv_with_pyarrow() never receives these options; its duplicate-name parity cases call the function directly.
Validation after the merge: just lint passed; 54 offline tests passed; live tests/pyathena/pandas/test_cursor.py -k "pyarrow or duplicate_column" 46 passed. AWS CI reruns on this push.
…type The follow-up review found that splitting the first line of a header-less file on the delimiter miscounted a quoted delimiter, which typed the wrong field as text, and that dtype values pandas_dtype() rejects with NotImplementedError still failed for absent columns. Read the first record with the csv module, which follows the same quoting as the reader, and skip dtype values that raise TypeError, ValueError, or NotImplementedError when choosing string columns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Typing the named fields of a header-less file needs their positions before reading, and each way of counting the first record's fields (delimiter split, the csv module) disagreed with pyarrow on some input: quoted delimiters, long fields, bare CR line ends, a byte order mark. Header-less files are the tab-separated results of DDL statements, so leave their fields with the types pyarrow infers, as before, and keep only the missing-value handling of NumPy str columns for them. Also convert NumPy str columns once, after the dates, instead of through a separate helper. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
WHAT
_read_csv_with_pyarrow(), which reads CSV results forPandasCursorwithengine="pyarrow"and default read options (#1057), no longer changes string values:ConvertOptions(column_types=...)), so values that look like numbers keep their text:"007","1e3", and the text"nan"stay as they are, as with pandas' C engine.strresolves to whenfuture.infer_stringis disabled, become object strings with missing values kept missing, after date parsing, instead of going throughastype(str), which turned NULL into the string"nan".A column with a string dtype is one whose
dtypeentry resolves to a pandasStringDtypeor a NumPy str dtype; Arrow string dtypes such aspd.ArrowDtype(pa.string())are read as text and keep their dtype. PyAthena mapschar,varchar,string,array,map, androwtostr.Two cases keep pandas' PyArrow engine behavior:
parse_datescolumn with a string dtype: the dtype is applied again after the dates are parsed.dtypeentries for columns not in the result, including invalid ones, are ignored as before, and other columns convert aspandas.read_csv(engine="pyarrow")does.When pandas reads the file itself (read options PyAthena does not handle), its PyArrow engine keeps both behaviors;
docs/pandas.mdstates them.Release note: with
engine="pyarrow"and default read options, string columns of query results are no longer changed. Numeric-looking text such as"007"was previously returned as"7.0". Withfuture.infer_stringdisabled, NULL was previously the string"nan".WHY
Closes #1062.
pandas' PyArrow engine lets pyarrow infer each column's type and applies
dtype=strafterwards, so a VARCHAR column whose values all look like numbers is read as float64 and turned back into"7.0"-style text (reported as pandas-dev/pandas#57666, closed upstream, but still reproduced with pandas 3.0.6 and pyarrow 25.0.1).With
future.infer_stringdisabled,strmeans a NumPy str dtype, andSeries.astype(str)turnsNaNinto"nan"(measured with pandas 3.0.6, also for object columns). That is pandas'astypesemantics rather than a CSV reader defect, so this part is not reported upstream.#1057 made
_read_csv_with_pyarrow()reproduce pandas' PyArrow engine, so it inherited both; the maintainer chose to read string columns as text instead of sending results with string columns to the C engine.Measured offline with pandas 3.0.6 and pyarrow 25.0.1, a
vcolumn of"1", NULL,"nan","007","1e3"(dtype=str, PyAthena's read options):infer_string['1', nan, 'nan', '007', '1e3']['1.0', nan, nan, '7.0', '1000.0']['1', nan, 'nan', '007', '1e3']['1', nan, 'nan', '007', '1e3']['1.0', 'nan', 'nan', '7.0', '1000.0']['1', nan, 'nan', '007', '1e3']Performance, measured locally (pandas 3.0.6, pyarrow 25.0.1, Apple silicon) on 1,000,000-row CSVs of 76–78 MiB. Columns: bigint, double, two varchar, date, timestamp, with 10% NULL strings. Times are medians of 15 runs alternating master and this PR's
_read_csv_with_pyarrow(), except the numeric-looking rows, which are best of 3; S3 transfer is not included.infer_stringpyarrow's read itself takes the same time with or without
column_types(0.05–0.06 s, both givingstring). The gains come from skipping the float round trip for numeric-looking columns and from converting NumPy-str columns once instead ofastype(str)before and after the dates. pandas' C engine took 1.0–1.3 s on the same data.TEST
Tested commit: 6bd7711.
just lint: passed.just docs lint: 0 errors.uv run --env-file .env pytest -n 2 tests/pyathena/pandas/ tests/pyathena/aio/pandas(live Athena): 360 passed at 6bd7711. Earlier runs: 364 passed at d09bccd. The first run at a3fcbb3 had one failure whose name was not captured, and it did not recur in any later run.TestPandasCursor.test_pyarrow_engine_string_values(live,infer_stringon and off) reads'1', NULL,'nan','007','1e3'withengine="pyarrow", checks that_read_csv_with_pyarrow()was used, and compares the column with the C engine's.test_read_csv_with_pyarrow_matches_pandasbuilds its options with the real_get_csv_read_options(). It expects the C engine's values for string-dtype columns of files with a header, and for header-less files the PyArrow engine's values with the C engine's missing values. All other columns follow the PyArrow engine. New cases:numeric_looking_strings,arrow_string_dtypes,tab_separated_numeric_fields,dtype_position_key.test_read_csv_with_pyarrow_string_dtype_after_datesandtest_read_csv_with_pyarrow_ignores_unused_dtype_entries(an invalidTypeErrorvalue and aNotImplementedErrorone,decimal128(10, 2)[pyarrow]).DESCRIBEof a 6-column table goes through_read_csv_with_pyarrow()and reads correctly.🤖 Generated with Claude Code