Skip to content

Read string columns as text for the pandas PyArrow engine - #1065

Merged
laughingman7743 merged 7 commits into
masterfrom
fix/1062-pyarrow-string-columns
Oct 4, 2026
Merged

laughingman7743 merged 7 commits into
masterfrom
fix/1062-pyarrow-string-columns

Conversation

@laughingman7743

@laughingman7743 laughingman7743 commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

WHAT

_read_csv_with_pyarrow(), which reads CSV results for PandasCursor with engine="pyarrow" and default read options (#1057), no longer changes string values:

  • Columns with a string dtype are read as pyarrow strings (ConvertOptions(column_types=...)), so values that look like numbers keep their text: "007", "1e3", and the text "nan" stay as they are, as with pandas' C engine.
  • Columns with a NumPy str dtype, which str resolves to when future.infer_string is disabled, become object strings with missing values kept missing, after date parsing, instead of going through astype(str), which turned NULL into the string "nan".

A column with a string dtype is one whose dtype entry resolves to a pandas StringDtype or a NumPy str dtype; Arrow string dtypes such as pd.ArrowDtype(pa.string()) are read as text and keep their dtype. PyAthena maps char, varchar, string, array, map, and row to str.
Two cases keep pandas' PyArrow engine behavior:

  • Header-less files, the tab-separated results of DDL statements: their fields keep the types pyarrow infers, because their names are assigned only after reading. Only the NULL handling applies to them.
  • A parse_dates column with a string dtype: the dtype is applied again after the dates are parsed.

dtype entries for columns not in the result, including invalid ones, are ignored as before, and other columns convert as pandas.read_csv(engine="pyarrow") does.
When pandas reads the file itself (read options PyAthena does not handle), its PyArrow engine keeps both behaviors; docs/pandas.md states them.

Release note: with engine="pyarrow" and default read options, string columns of query results are no longer changed. Numeric-looking text such as "007" was previously returned as "7.0". With future.infer_string disabled, NULL was previously the string "nan".

WHY

Closes #1062.
pandas' PyArrow engine lets pyarrow infer each column's type and applies dtype=str afterwards, so a VARCHAR column whose values all look like numbers is read as float64 and turned back into "7.0"-style text (reported as pandas-dev/pandas#57666, closed upstream, but still reproduced with pandas 3.0.6 and pyarrow 25.0.1).
With future.infer_string disabled, str means a NumPy str dtype, and Series.astype(str) turns NaN into "nan" (measured with pandas 3.0.6, also for object columns). That is pandas' astype semantics rather than a CSV reader defect, so this part is not reported upstream.
#1057 made _read_csv_with_pyarrow() reproduce pandas' PyArrow engine, so it inherited both; the maintainer chose to read string columns as text instead of sending results with string columns to the C engine.

Measured offline with pandas 3.0.6 and pyarrow 25.0.1, a v column of "1", NULL, "nan", "007", "1e3" (dtype=str, PyAthena's read options):

infer_string C engine PyArrow engine (master) this PR
True ['1', nan, 'nan', '007', '1e3'] ['1.0', nan, nan, '7.0', '1000.0'] ['1', nan, 'nan', '007', '1e3']
False ['1', nan, 'nan', '007', '1e3'] ['1.0', 'nan', 'nan', '7.0', '1000.0'] ['1', nan, 'nan', '007', '1e3']

Performance, measured locally (pandas 3.0.6, pyarrow 25.0.1, Apple silicon) on 1,000,000-row CSVs of 76–78 MiB. Columns: bigint, double, two varchar, date, timestamp, with 10% NULL strings. Times are medians of 15 runs alternating master and this PR's _read_csv_with_pyarrow(), except the numeric-looking rows, which are best of 3; S3 transfer is not included.

varchar values infer_string master this PR
ordinary text True 0.57 s 0.54 s
ordinary text False 0.88 s 0.71 s
one numeric-looking column (best of 3) True 0.60 s 0.49 s
one numeric-looking column (best of 3) False 0.99 s 0.80 s

pyarrow's read itself takes the same time with or without column_types (0.05–0.06 s, both giving string). The gains come from skipping the float round trip for numeric-looking columns and from converting NumPy-str columns once instead of astype(str) before and after the dates. pandas' C engine took 1.0–1.3 s on the same data.

TEST

Tested commit: 6bd7711.

  • just lint: passed. just docs lint: 0 errors.
  • uv run --env-file .env pytest -n 2 tests/pyathena/pandas/ tests/pyathena/aio/pandas (live Athena): 360 passed at 6bd7711. Earlier runs: 364 passed at d09bccd. The first run at a3fcbb3 had one failure whose name was not captured, and it did not recur in any later run.
  • New TestPandasCursor.test_pyarrow_engine_string_values (live, infer_string on and off) reads '1', NULL, 'nan', '007', '1e3' with engine="pyarrow", checks that _read_csv_with_pyarrow() was used, and compares the column with the C engine's.
  • test_read_csv_with_pyarrow_matches_pandas builds its options with the real _get_csv_read_options(). It expects the C engine's values for string-dtype columns of files with a header, and for header-less files the PyArrow engine's values with the C engine's missing values. All other columns follow the PyArrow engine. New cases: numeric_looking_strings, arrow_string_dtypes, tab_separated_numeric_fields, dtype_position_key.
  • New test_read_csv_with_pyarrow_string_dtype_after_dates and test_read_csv_with_pyarrow_ignores_unused_dtype_entries (an invalid TypeError value and a NotImplementedError one, decimal128(10, 2)[pyarrow]).
  • Checked live: DESCRIBE of a 6-column table goes through _read_csv_with_pyarrow() and reads correctly.

🤖 Generated with Claude Code

laughingman7743 and others added 3 commits October 4, 2026 12:35
pandas' PyArrow engine lets pyarrow infer a type for each column and only
then applies dtype=str, so a VARCHAR column whose values look like
numbers came back changed ("007" as "7.0", "1e3" as "1000.0", the text
"nan" as NaN), and without future.infer_string, astype(str) turned NULL
into the string "nan". _read_csv_with_pyarrow() reproduced both.

Read the columns with a string dtype as pyarrow strings, and give the
columns with a NumPy str dtype object values that keep missing values
missing, so these columns have the C engine's values. pandas' own
PyArrow engine, used for options PyAthena does not read itself, keeps
pandas' behavior (pandas-dev/pandas#57666), which the docs now state.

Closes #1062

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A dtype mapping can have keys that are not column names, such as
positions, which pandas ignores; pyarrow's column_types raised TypeError
for them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread pyathena/pandas/result_set.py Outdated
f"f{index}": pa.string() for index, name in enumerate(names) if name in string_columns
}
else:
# pandas ignores dtype keys that are not column names, such as positions.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round one (implementation behavior): base 8bbf8cb5e2d566bfe3dda86273b498168cfe4360, head a3fcbb3e → repaired in 51129926.

Covered: _read_csv_with_pyarrow() (string-column detection, column_types for header and header-less files, the NumPy-str path, both astype() calls, the infer_string=False object conversion, date parsing); its callers _read_csv() and _reads_csv_with_pyarrow(); DefaultPandasTypeConverter dtypes; the C engine as the reference for string columns.

Checked outside the suite (pandas 3.0.6, pyarrow 25.0.1): string columns equal the C engine's with infer_string on and off; a dtype entry for a missing column is ignored by pyarrow; duplicate names both read as text; a header-less .txt keeps "007".

Result: FINDINGS

  1. A dtype key that is not a column name (e.g. {0: str}, which pandas ignores) made pyarrow's column_types raise TypeError: expected bytes, int found. Repaired in 5112992 (only str keys go to column_types) with the dtype_position_key parity case.

Known limit, not changed: for a header-less result with more fields than names, which pandas names from the end, column_types maps names from the start, so the shifted columns keep inferred types. This needs a tab inside a value of a DDL text result.
Validation: lint passed; test_result_set.py 36 passed. Live runs are listed in the PR body (one unnamed failure in the first run, which did not recur in --lf or a full rerun).

Comment thread docs/pandas.md Outdated
With the PyArrow engine, when `keep_default_na` and `na_values` are the defaults and the only pandas.read_csv() options passed to `execute()` are `dtype` as a mapping and `parse_dates` as a list, PyAthena reads the file with `pyarrow.csv`, allowing newlines in quoted values, and converts it to a DataFrame as `pandas.read_csv(engine="pyarrow")` does.
With the PyArrow engine, when `keep_default_na` and `na_values` are the defaults and the only pandas.read_csv() options passed to `execute()` are `dtype` as a mapping and `parse_dates` as a list, PyAthena reads the file with `pyarrow.csv`, allowing newlines in quoted values, and converts it to a DataFrame as `pandas.read_csv(engine="pyarrow")` does, except that columns with a string dtype keep their text and their missing values, as with the C engine.
Otherwise, pandas reads the file, and its PyArrow engine raises an error or returns wrong values when a quoted value containing a newline crosses one of pyarrow's read blocks.
It also changes the values of a column with a string dtype whose values all look like numbers, such as `"007"` to `"7.0"`, and without `future.infer_string`, turns NULL in a column with a string dtype into the string `"nan"`.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, compatibility, operations): base 8bbf8cb5e2d566bfe3dda86273b498168cfe4360, head 487d88ef71399f48639da61af83eca8a9b806822.

Claims checked:

  • "pandas does not expose it" (BUG: pyarrow read_csv engine stripping leading zeros with dtype=str pandas-dev/pandas#57666) was too strong, because the issue is closed upstream. The PR now states that it is reported there and still reproduces with pandas 3.0.6 / pyarrow 25.0.1 (measured).
  • Series.astype(str) turning NaN into "nan" with infer_string=False was measured for StringDtype and object columns; "long-standing" was removed as unverified.
  • PyAthena maps char/varchar/string/array/map/row to str: confirmed in DefaultPandasTypeConverter._dtypes.
  • The docs said the PyArrow engine "changes strings that look like numbers"; it happens only when the column's values all look like numbers (pyarrow infers per column), and the NULL case only in string-dtype columns. The wording was narrowed (487d88e).

Existing callers: string-dtype columns change only where master returned changed values (numeric-looking text, the text "nan" read as NaN, NULL as "nan"). All other columns convert exactly as before, as the parity tests confirm. Results that pandas reads itself are unchanged. No AWS requests change.

Result: CLEAN after the wording corrections.

…yArrow reader

The independent review found that the string-column change:

- applied the NumPy str handling to Arrow string dtypes, whose kind is
  also "U", returning objects instead of the requested dtype;
- typed the first fields of a header-less file while pandas gives the
  names to the last ones, changing other columns' types when a row has
  extra fields;
- left parse_dates columns with a NumPy str dtype as datetimes, while the
  dtype used to apply again after parsing;
- failed on invalid dtype entries for columns not in the result, which
  pandas ignores.

Apply the dtype mapping through _astype_keeping_missing_strings() before
and after parsing dates, limit its NumPy str handling to NumPy dtypes,
count the fields of a header-less file's first line to type the named
ones, and skip dtype values that do not resolve when choosing string
columns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread pyathena/pandas/result_set.py Outdated
return df


def _read_csv_with_pyarrow(source: IOBase, read_csv_kwargs: dict[str, Any]) -> DataFrame:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed): Codex CLI 0.160.0, model gpt-6-astra, codex exec -s read-only --ephemeral, session 01a10518-1ab0-7b41-a3dd-4026904daeb3. Base 8bbf8cb5e2d566bfe3dda86273b498168cfe4360, head 487d88ef71399f48639da61af83eca8a9b806822, detached snapshot (unchanged afterwards). The prompt held the literal diff and conventions only, with no PR text. This was a static review.

Reviewer coverage: the four-file diff, sync/Future/asyncio PandasCursor callers, engine selection, CSV/TSV options, date conversion, tests, docs, and the pandas/pyarrow sources (UNLOAD and the Arrow/Polars result sets do not call the helper).

Verdict: FINDINGS (4 × P2 regressions, 1 × P3)

  1. result_set.py:137: pd.ArrowDtype(pa.string()) / large_string() have kind == "U" too, so they got the NumPy-str object path instead of the requested Arrow dtype.
  2. result_set.py:87: header-less files typed f0… from the start, while names are right-aligned when rows have extra fields (001\t2\t003 with names=["v"] typed f0 as text and left v numeric), changing other columns' output types.
  3. result_set.py:138: a NumPy-str parse_dates column that parses successfully stayed datetime; master re-applied the dtype after dates.
  4. result_set.py:81: pandas_dtype() on every mapping value made an invalid entry for an absent column ({{"unused": "not-a-dtype"}}) fail, while pandas ignores it.
  5. P3 docs/pandas.md:559: the C-engine equivalence needs the parse_dates exception, and the "nan" warning applies only to NumPy str dtypes.
    The reviewer suggested keeping the cast mapping separate from the pyarrow column types.

Author verification: 1, 3, and 4 reproduced (Arrow dtype → object with NaN, pandas gives string[pyarrow] with <NA>; dates → Timestamps, master gives ['2024-01-01', 'None']; TypeError: data type 'not-a-dtype' not understood). 2 was confirmed by the reviewer's example. It was my round-one "known limit", which I had judged rare, but it changes other columns, so it is repaired, not deferred. 5 accepted.

Repair 7bad8fe:

  • _astype_keeping_missing_strings() applies the normalized mapping, and only NumPy (non-extension) str dtypes become object strings with missing values kept. It runs before and after date parsing, so 1 and 3 follow master's dtype-after-dates order; NaT and NULL become NaN rather than master's 'None'/'nan'.
  • Header-less files: the first line's field count gives the offset of the named fields (readline() then seek(0) on the opened stream). _read_csv() passes the helper only an IOBase source, which the PyArrow path always opens.
  • String-column detection skips values that pandas_dtype() rejects; validation stays on the result's columns.
  • Docs: string-dtype columns "are read as text", and the NULL wording names dtype str (NumPy str without infer_string).
  • Tests: parity cases arrow_string_dtypes and tab_separated_numeric_fields, plus test_read_csv_with_pyarrow_string_dtype_after_dates and test_read_csv_with_pyarrow_ignores_unused_dtype_entries (the C engine validates unused entries, so that case compares with pandas' PyArrow engine only). 6 of the new cases fail on 487d88e.

Self-review of the repair: behavior: test_result_set.py + routing tests 54 passed offline; live, DESCRIBE of a 6-column table through the header-less path read correctly, and test_cursor.py -k "pyarrow_engine or show_columns or describe" 8 passed. Claims: the docs and PR text now match the narrowed wording. A narrow independent follow-up on 487d88ef..7bad8fe5 is running.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up 1 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a1052d-3908-7762-a8f3-6bb4fea80ce4. This was a narrow review of 487d88ef71399f48639da61af83eca8a9b806822..7bad8fe55571c5f22ea610c666e1db2b6a351f59, static only. The reviewer marked prior findings 1, 3, and 5 resolved, and 2 and 4 partially resolved.

Verdict: FINDINGS (2 × P2)

  1. result_set.py:123: splitting the first line on the delimiter miscounts a quoted delimiter (007\t"a\tb" counts 3 fields, while pyarrow parses 2), so f1 was typed instead of f0 and v lost "007"; a quoted newline can undercount.
  2. result_set.py:115: pandas_dtype("decimal128(10, 2)[pyarrow]") raises NotImplementedError, which the detection did not catch, so an unused entry still failed.

Author verification: both confirmed (NotImplementedError for that string; the quoted-tab case typed the wrong field).

Repair d09bccd: the first record is read with csv.reader (same " quoting and doubling as the reader) over decoded readline() lines, then seek(0); an empty first record counts as one field. Detection skips values raising TypeError, ValueError, or NotImplementedError. New cases: tab_separated_quoted_tab and tab_separated_quoted_newline, plus the decimal128 entry in the unused-entries test. The quoted-tab case and the unused-entries test fail on 7bad8fe.
Self-review of the repair: lint passed; 58 offline tests passed; live DESCRIBE through the header-less path on S3File read correctly. A full live run at 7bad8fe (with these edits partly in the working tree) gave 360 passed. A clean full run at d09bccd and a narrow Codex follow-up are running.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up 2 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a1053b-6e0a-7943-b4fa-de184b235565. This was a narrow review of 7bad8fe55571c5f22ea610c666e1db2b6a351f59..d09bccd4e9d655791a041ad9cbcd9e13fb108573, static only. The reviewer marked the absent-column dtype issue resolved.

Verdict: FINDINGS (3 × P2), all in the header-less field-count probe:

  1. result_set.py:125: csv.reader limits a field to 131,072 characters, so a long first field that pyarrow accepts now raises _csv.Error.
  2. result_set.py:124: bare-CR records (b"007\r010\r") arrive in one binary readline(), and csv.reader raises "new-line character seen in unquoted field".
  3. result_set.py:124: a UTF-8 BOM before a quoted first field is kept by the decode, so the count disagrees with pyarrow (which strips it), and the wrong field is typed.
    The reviewer also noted that the multiline test put the newline in the last field.

Decision (author), 6bd7711: these three come from counting fields before reading, which was needed only to type header-less fields. Each counting method so far (delimiter split, then csv.reader) has disagreed with pyarrow on some input, so further probing would keep widening this PR. Header-less files are the tab-separated results of DDL statements, so their fields now keep the types pyarrow infers, as on master, and the probe is gone. They keep only the NumPy-str missing-value handling (NULL stays missing without infer_string). Numeric-looking text in DDL results therefore still follows pandas' PyArrow engine; the docs say so ("other than in the tab-separated results of DDL statements").
Also in 6bd7711: the NumPy-str conversion happens once after date parsing, inline, instead of in the separate _astype_keeping_missing_strings() helper. Both astype() calls skip those columns, and date parsing goes through astype("string"), so the result is the same. The parity test compares header-less string columns with pandas' PyArrow engine, except that missing values follow the C engine. The two probe-only cases were removed.
Validation: lint passed; 54 offline tests passed; a clean live run at d09bccd gave 364 passed. A live run at 6bd7711 and a Codex follow-up on d09bccd4..6bd7711b are running.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up 3 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a10548-f947-7573-b0bd-1cfbb7e48175. This covered d09bccd4e9d655791a041ad9cbcd9e13fb108573..6bd7711b3b5b51e3035a7755db53dbc072969291 plus the full diff, static only. The reviewer resolved the long-field, bare-CR, and BOM findings (the probe is gone) and found no defect in the single conversion after dates.

Verdict: FINDINGS (1 × P2)

  • result_set.py:180: in a header-less result with infer_string=False, a field pyarrow infers as float keeps the literal text nan as a float NaN, and the new notna() mask makes it missing like NULL; master returned the string "nan" for it.

Author verification and decision: deferred, no change. Measured on 1.5\nnan\n\n (names=["v"], dtype={"v": str}):

infer_string master / pandas PyArrow engine this PR
True (default) ['1.5', nan, nan] ['1.5', nan, nan]
False ['1.5', 'nan', 'nan'] ['1.5', nan, nan]
With the default setting, master already turns the literal into a missing value: pyarrow's float inference cannot tell it from NULL. With infer_string=False, master returned 'nan' for both the literal and NULL, so it did not keep them apart either. This PR makes infer_string=False match the default. Telling them apart would need Arrow's null bitmap carried through conversion, only for inferred header-less (DDL result) fields, which goes past this issue. Query results with a header read string columns as text and are not affected.

Validation at 6bd7711: tests/pyathena/pandas/ + tests/pyathena/aio/pandas 360 passed live (clean working tree).

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Merge of master (1f6fc7a): master gained #1066 (#1051, duplicate column types) and other PRs. The merge had no conflicts.

Upstream contracts checked: _read_csv_header_as_labels() rewrites names/header/skiprows only for results with duplicate column names. Since #1050, _get_csv_engine() never selects the PyArrow engine for those (_needs_csv_column_name_resolution()), so _read_csv_with_pyarrow() never receives these options; its duplicate-name parity cases call the function directly.
Validation after the merge: just lint passed; 54 offline tests passed; live tests/pyathena/pandas/test_cursor.py -k "pyarrow or duplicate_column" 46 passed. AWS CI reruns on this push.

laughingman7743 and others added 2 commits October 4, 2026 13:42
…type

The follow-up review found that splitting the first line of a
header-less file on the delimiter miscounted a quoted delimiter, which
typed the wrong field as text, and that dtype values pandas_dtype()
rejects with NotImplementedError still failed for absent columns.

Read the first record with the csv module, which follows the same
quoting as the reader, and skip dtype values that raise TypeError,
ValueError, or NotImplementedError when choosing string columns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Typing the named fields of a header-less file needs their positions
before reading, and each way of counting the first record's fields
(delimiter split, the csv module) disagreed with pyarrow on some input:
quoted delimiters, long fields, bare CR line ends, a byte order mark.
Header-less files are the tab-separated results of DDL statements, so
leave their fields with the types pyarrow infers, as before, and keep
only the missing-value handling of NumPy str columns for them.

Also convert NumPy str columns once, after the dates, instead of through
a separate helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@laughingman7743
laughingman7743 marked this pull request as ready for review October 4, 2026 05:14
@laughingman7743
laughingman7743 merged commit 81e2f4f into master Oct 4, 2026
9 checks passed
@laughingman7743
laughingman7743 deleted the fix/1062-pyarrow-string-columns branch October 4, 2026 05:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PandasCursor with engine="pyarrow" changes string values: numeric-looking text, and NULL as 'nan' without infer_string

1 participant