Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/pandas.md
Original file line number Diff line number Diff line change
Expand Up @@ -556,6 +556,8 @@ Common performance options:

With `engine="pyarrow"`, PandasCursor uses the PyArrow engine only when pyarrow is installed, no chunksize is set (explicitly or by `auto_optimize_chunksize`), `quoting` is the default, the result has no columns that need a converter (`boolean`, `decimal`, `varbinary`, `json`, `time with time zone`, and `timestamp with time zone` with the default converter), and the result file is at least `AthenaPandasResultSet.PYARROW_MIN_FILE_SIZE_BYTES` bytes.
Otherwise, it falls back to the C engine.
With the PyArrow engine, when `keep_default_na` and `na_values` are the defaults and the only pandas.read_csv() options passed to `execute()` are `dtype` as a mapping and `parse_dates` as a list, PyAthena reads the file with `pyarrow.csv`, allowing newlines in quoted values, and converts it to a DataFrame as `pandas.read_csv(engine="pyarrow")` does.
Otherwise, pandas reads the file, and its PyArrow engine raises an error or returns wrong values when a quoted value containing a newline crosses one of pyarrow's read blocks.

Apart from PandasCursor's own options such as `engine` and `chunksize`, an option passed here replaces the value PyAthena sets for the same pandas.read_csv() argument.
For example, `dtype` replaces the whole column type mapping, and `parse_dates` replaces the list of date, time, and timestamp columns.
Expand Down
5 changes: 5 additions & 0 deletions docs/polars.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,11 @@ does not require PyArrow as a dependency. CSV results read without the chunksize
PyAthena's own S3FileSystem (fsspec compatible), and other reads use Polars' native S3 access,
so s3fs is also not required.

Options passed to `execute()`, other than PolarsCursor's own options such as `chunksize`, are passed to the Polars read function.
`pl.read_csv()` reads the CSV result with pyarrow when `use_pyarrow=True` is given with `schema_overrides=None`, which replaces PyAthena's column types, and Polars' other conditions for it hold.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round two (claims, compatibility, operations): base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, head e272719123fce03c291f22f015ba00ff04b8eb0d.

Claims checked:

  • pandas builds ParseOptions from delimiter, quote_char, escape_char, and ignore_empty_lines only: confirmed in pandas 3.0.6 arrow_parser_wrapper.py, which also maps on_bad_lines to invalid_row_handler. The PR text now names it.
  • pandas' default NA strings are private: STR_NA_VALUES lives in pandas._libs.parsers and is not exported by pandas.io.parsers, hence the keep_default_na=False condition.
  • Silent wrong values: reproduced 19 of 5,000 with pyarrow.csv directly (newlines_in_values=False), and all were correct with True.
  • BUG: read_csv with pyarrow engine treats newlines in quotes as newline when they should be ignored pandas-dev/pandas#56396 and #59009: open (GitHub search, 2026-10-04).
  • Performance: 0.58 s versus 0.63 s (pandas PyArrow) and 1.49 s (C) on 70 MiB locally; this measures CPU only, not S3 transfer.
  • docs/polars.md said "with use_pyarrow=True and schema_overrides=None" as if those two were sufficient. Polars 1.44.2 also requires n_rows, n_threads, and null_values to be unset and low_memory=False, so the text now names schema_overrides=None and "Polars' other conditions" (e272719). Live, PolarsCursor with those options raises OperationalError: CSV parser got out of sync with chunker on the 2,000-row query, while the default reader returns all rows correctly.
  • docs/pandas.md: the engine conditions match _get_csv_engine(), and the conversion sentence matches the function.

Existing callers: results without multi-line values give frames equal to before (assert_frame_equal(check_exact=True) over every eligible type, infer_string on and off, and 200 random files). Callers that passed other read_csv options or keep_default_na=True with engine="pyarrow" now get the C engine, which is listed as a release-note item. No AWS requests are added (one GET of the result file, as before).

Result: CLEAN after the wording corrections. The duplicate-name interaction from round one stays deferred to #1050.

That read raises an error or returns wrong values when a quoted value containing a newline crosses one of pyarrow's read blocks.
The default Polars CSV reader reads such values.

You can use the PolarsCursor by specifying the `cursor_class`
with the connect method or connection object.

Expand Down
124 changes: 123 additions & 1 deletion pyathena/pandas/result_set.py
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,104 @@ def _no_trunc_date(df: DataFrame) -> DataFrame:
return df


def _read_csv_with_pyarrow(source: str | IOBase, read_csv_kwargs: dict[str, Any]) -> DataFrame:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review round one (implementation behavior): base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, head 5339aff7e2993ebbc72bd0385891a3588bdb1b1d.

Covered: _read_csv_with_pyarrow() compared step by step with pandas 3.0.6 ArrowParserWrapper.read(), arrow_table_to_pandas(), _post_convert_dtypes(), _maybe_convert_string_to_object(), and date_converter(); the _get_csv_engine() conditions and their only caller, _read_csv(); source opening for the PyArrow path (self._fs.open without storage_options, fsspec with them, even None); binary/converter columns (never on this path); .txt results (header-less, names); the auto-chunk and as_pandas() paths (the PyArrow engine is never chunked); sync, async, and aio pandas cursors (shared result set).

Edge cases compared with pandas.read_csv(engine="pyarrow") outside the suite: a header-only file, an all-NULL date column, an all-NULL integer column, and "" versus NULL all gave equal frames; dtype as a single type or None made pandas raise, so the engine condition now requires a mapping and the function assumes one.

Result: FINDINGS

  1. Duplicate column names (SELECT 1 AS x, 2 AS x): master's PyArrow engine silently returned {'x': [2]}, and _read_csv_with_pyarrow() raises AttributeError (OperationalError to the caller) because df[column] is a DataFrame. Deferred to Keep columns with the same name, and key CSV column types by the reader's column labels #1050, which adds a duplicate-name condition to this is_compatible expression that sends such results to the C engine; whichever PR merges second takes both conditions. The PR body says so.

Simplifications already applied before publishing: the reader is a module function (it uses no result-set state), and the non-mapping dtype branches were removed in favor of the engine condition.
Validation: tests/pyathena/pandas/ + tests/pyathena/aio/pandas 327 passed live; the new live test fails on the pandas route with CSV parser got out of sync with chunker.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent review (relayed): Codex CLI 0.160.0, model gpt-6-astra, codex exec -s read-only --ephemeral, session 01a1046c-5f68-7ce1-8882-f37d65fb08a4. Base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, head e272719123fce03c291f22f015ba00ff04b8eb0d, detached snapshot (unchanged afterwards). The prompt held the literal diff and repository conventions plus read-only access to the installed pandas, pyarrow, and Polars sources; it held no PR text or prior findings. This was a static review: no tests, Python, builds, or network access.

Reviewer coverage: the five-file diff, CSV/TSV conversion, engine selection, dtype/date/NULL handling, sync and async callers, stream ownership, fsspec options, tests, and the pandas/Polars docs, compared with pandas 3.0.6, pyarrow 25.0.1, and Polars 1.44.2.

Verdict: FINDINGS (5 × P2 regressions)

  1. result_set.py:135: pandas applies dtype again after parsing dates (_finalize_pandas_output()), so dtype={"d": "string"}, parse_dates=["d"] returns strings on master and datetimes here.
  2. result_set.py:107: with duplicate column names, df[column].dtype hits a DataFrame and raises (OperationalError); pandas skips that lookup when the name has a dtype entry and returns both columns.
  3. result_set.py:73: pandas expands numeric na_values ("-1" also matches -1.0), and a NumPy array of markers raises on or [].
  4. result_set.py:508: sending other options to the C engine breaks PyArrow-only options, e.g. a callable on_bad_lines is rejected by the C engine.
  5. result_set.py:702: the direct fsspec open with storage_options drops pandas' retry with anon=True (pandas/io/common.py).
    The reviewer also judged the DataFrame-equivalence sentence in docs/pandas.md too broad until these are fixed.

Author verification: all five confirmed. Findings 1–4 were run against pandas 3.0.6: findings 1 and 3 reproduce as described, master returns ['x', 'x'] with both values for duplicate names, and the C engine raises on_bad_line can only be a callable function if engine='python' or 'pyarrow'. Finding 5 was confirmed in the pandas source, not run. My round-one record was wrong about duplicate names: it read master's result through to_dict(), which collapses duplicate keys.

Repair in 84e2fca, recorded in the reply below.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repair 84e2fca (old head e272719123fce03c291f22f015ba00ff04b8eb0d → new head 84e2fca36c98d8a6306d1ef3ecfd0a9a591e51c8, same base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1)

  • Findings 3–5 led to a routing change. _get_csv_engine() is back to master. _read_csv() calls _read_csv_with_pyarrow() only when the new _reads_csv_with_pyarrow() holds: keep_default_na is false, na_values is the default ("",) as a list or tuple, and execute() passes only a dtype mapping and/or parse_dates. Every other case goes to pandas.read_csv(engine="pyarrow") as on master, so numeric na_values, keep_default_na, storage_options (with its anonymous retry), and PyArrow-only options such as a callable on_bad_lines behave as before; no case moves to the C engine any more. The storage_options open branch was removed, and the quoting/converters pops for the pandas PyArrow call are restored.
  • Finding 1: the dtype mapping (including the NumPy integer entries) is applied again after date parsing, as _finalize_pandas_output() does.
  • Finding 2: the integer lookup checks column not in dtype first, as _post_convert_dtypes() does, and the string-to-object loop works by position with isetitem().
  • Tests: test_reads_csv_with_pyarrow replaces the engine-selection test; new parity cases for a dtype entry on a date column and for duplicate names fail on e272719 and pass now; the live test no longer has a storage_options case (pandas reads that one).

Self-review of the repair, round one (behavior): parity harness all equal (types, TSV, infer_string off, 200 random files, user dtype/parse_dates); the second astype() mirrors pandas, including its cost; an ndarray na_values is not a list/tuple, so pandas reads it as before. Validation: just lint passed; offline 22 targeted tests passed; tests/pyathena/pandas/ + tests/pyathena/aio/pandas 331 passed live.

Round two (claims): the PR text no longer lists C-engine fallbacks as release-note items. Its only behavior change is the multi-line fix on the default path. docs/pandas.md keeps master's engine conditions and adds when PyAthena reads the file and that pandas' engine still has the limitation otherwise, which also scopes the DataFrame-equivalence sentence the reviewer flagged. The Polars limitation was reported upstream as pola-rs/polars#29717, reproduced on Polars 1.44.2 and main 9ee0dc5b81.

The independent follow-up on this repair is pending.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up 1 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a10488-30bd-78b0-98b8-37d67ec7ed57. This was a full review of 84e2fca36c98d8a6306d1ef3ecfd0a9a591e51c8 (base e49d576ce1e8cb60d74e1fde56de9c72ff3cd8e1, detached snapshot unchanged), static only.

The reviewer found the five earlier findings resolved: dates followed by a final cast; duplicate names via short-circuit and positional conversion; non-default NA settings, other options, and storage_options all delegated to pandas; master's engine selection restored.

Verdict: FINDINGS

  1. P2 result_set.py:106: dtype mapping values reached astype() raw. pandas normalizes them with pandas_dtype(), so dtype={"x": None} gives float64 with NaN in pandas and nullable Int64 here.
  2. P3 result_set.py:124: parse_dates="d" or ("d",) was silently ignored by the helper, while pandas raises TypeError: Only booleans and lists are accepted for the 'parse_dates' parameter.
    The reviewer also noted, as pre-existing (pandas itself fails): duplicate names without a dtype entry, and duplicate date-column names.

Author verification: both reproduced against pandas 3.0.6 (float64 vs Int64; pandas TypeError for a string and a tuple).

Repair fbceb11b: the retained values go through pd.api.types.pandas_dtype(), and _reads_csv_with_pyarrow() requires parse_dates given to execute() to be a list, so other forms reach pandas (and its TypeError) as before. The helper now iterates parse_dates directly. docs/pandas.md says "parse_dates as a list". New cases: parity dtype_none, and the predicate's parse_dates_string and parse_dates_tuple. All 4 fail on 84e2fca and pass now.
Self-review of the repair: behavior: pandas_dtype() is the public form of what _post_convert_dtypes() applies, the parity harness is all equal, and lint passed. Claims: the docs and PR text name the list requirement. No AWS calls change. The narrow independent follow-up on 84e2fca3..fbceb11b is running.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up 2 (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a1049b-7b5b-7670-991a-0c16591db2ae. This was a narrow review of 84e2fca36c98d8a6306d1ef3ecfd0a9a591e51c8..fbceb11b894f3b208554bc53ba86af3e87545119 (detached snapshot unchanged), static only.

Verdict: CLEAN. Both points are resolved, and the reviewer found no new defect.

  • result_set.py:107 normalizes the values as pandas does (None → float64); the regression compares values and dtypes with infer_string on and off.
  • result_set.py:731 admits only a list parse_dates; strings and tuples reach pandas unchanged and raise its TypeError, which _read_csv() wraps in OperationalError. The routing cases and the docs match.

Validation at fbceb11: tests/pyathena/pandas/ + tests/pyathena/aio/pandas 335 passed live.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Merge of master (15399b5): master gained #1044 (_JSONConverter, managed integer dtypes) and other PRs since base e49d576c. The only conflict was in pyathena/pandas/result_set.py, where both sides added a module-level definition at the same place, so both were kept unchanged.

Upstream contracts checked: _read_csv() now sets self._csv_converters before reading and restores json NULLs afterwards. The PyArrow engine is chosen only without converters, so that set is empty on this path, and the _read_csv_with_pyarrow() dispatch is unchanged.
Validation after the merge: just lint passed; the offline pandas result-set and engine/routing tests passed (39); the parity harness is all equal. AWS CI reruns on this push.

"""Read a CSV result with pyarrow as ``pandas.read_csv(engine="pyarrow")`` does.

pandas' PyArrow engine does not set ``newlines_in_values``, so pyarrow
raises or returns wrong values when a quoted value containing a newline
crosses one of its read blocks. This reads the file with that option and
converts the table as pandas does for the options that
``AthenaPandasResultSet._reads_csv_with_pyarrow()`` accepts: NULL-typed
columns become float64, integer columns without a ``dtype`` entry get
NumPy integer types, the ``dtype`` mapping is applied before and after the
``parse_dates`` columns are parsed with ``pandas.to_datetime()``, and
without ``future.infer_string``, strings become objects.

Args:
source: The result file.
read_csv_kwargs: The pandas.read_csv() options built by
``AthenaPandasResultSet._get_csv_read_options()``.

Returns:
The result as a DataFrame.
"""
import pandas as pd
import pyarrow as pa
from pyarrow import csv as pyarrow_csv

header = read_csv_kwargs["header"]
names = read_csv_kwargs["names"]
null_values = list(read_csv_kwargs["na_values"])
table = pyarrow_csv.read_csv(
source,
read_options=pyarrow_csv.ReadOptions(autogenerate_column_names=header is None),
parse_options=pyarrow_csv.ParseOptions(
delimiter=read_csv_kwargs["sep"],
ignore_empty_lines=read_csv_kwargs["skip_blank_lines"],
newlines_in_values=True,
),
convert_options=pyarrow_csv.ConvertOptions(
null_values=null_values, strings_can_be_null="" in null_values
),
)
schema = table.schema
for index, type_ in enumerate(schema.types):
if pa.types.is_null(type_):
schema = schema.set(index, schema.field(index).with_type(pa.float64()))
integer_dtypes = {
pa.int8(): pd.Int8Dtype(),
pa.int16(): pd.Int16Dtype(),
pa.int32(): pd.Int32Dtype(),
pa.int64(): pd.Int64Dtype(),
}
# Integers become nullable dtypes first so that a dtype entry converts them
# without going through float64.
df = table.cast(schema).to_pandas(types_mapper=integer_dtypes.get)
if header is None:
# pandas names the columns beyond the given names by their positions.
df.columns = [str(index) for index in range(len(df.columns) - len(names))] + names

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional review requested by the maintainer (relayed): Claude Code 2.1.287, model claude-fable-5-1, effort high, profile max (first-party claude.ai Max, not Enterprise), session 4a0ff4d9-d1d0-4569-9085-d3f6f96c5219. Snapshot 15399b55355736b8c4bf6711494e61e4f4063b78 (detached, unchanged afterwards), base a57325b1fbc911d928bbe2c7e7054f6bdcb42c92. Tools were limited to Read, Grep, Glob, and git show/diff/log; the pandas, pyarrow, and Polars sources were readable. The prompt held the literal diff only, with no PR text or prior findings. This was a static review. Fable is another Claude model, so this complements the Codex reviews above rather than replacing them.

Reviewer coverage: the helper, the predicate and _read_csv(), _get_csv_read_options(), _get_csv_engine(), _finish_csv_frame(), and _trunc_date(), compared step by step with pandas 3.0.6 (ArrowParserWrapper, arrow_table_to_pandas, _post_convert_dtypes, _maybe_convert_string_to_object, date_converter, _clean_na_values, pandas_dtype); the Polars 1.44.2 use_pyarrow branch; the tests and docs.

Verdict: FINDINGS (all low, no blocker)

  1. Low regression: for a header-less (.txt) result with more fields than names, pandas prefixes "0", "1", … (_adjust_column_names), while df.columns = names raised a length mismatch.
  2. Low test gap: the to_datetime failure fallback was not covered.
  3. Low test design: _pyarrow_read_csv_kwargs() re-implemented _get_csv_read_options(), so a change to the builder would go unnoticed.
  4. Info, docs: newlines_in_values=True reduces pyarrow's multi-threaded parsing, so the reviewer suggested a performance note.
    Pre-existing, as noted by the reviewer: with future.infer_string=False, dtype=str turns NULL into a string in pandas itself; duplicate names without a dtype entry fail in pandas too.

Author verification and repair 2302049:

  1. Reproduced offline: pandas gives ['0', '1', 'col_name'], the helper raised. The reviewer's DESCRIBE scenario does not occur live: on Athena DESCRIBE describes three columns (col_name, data_type, comment), and the helper read it correctly (5 rows). A header-less result whose values contain tabs would still diverge, so the helper now prefixes positional names as pandas does. The new case tab_separated_extra_fields fails on 15399b5.
  2. The fallback already matched pandas. Added the unparsed_dates case.
  3. The parity tests now build their options with the real _get_csv_read_options() on an AthenaPandasResultSet with patched description/output_location, and assert _reads_csv_with_pyarrow().
  4. Not applied. Measured locally on 70 MiB, the helper with newlines_in_values=True takes 0.58 s versus 0.63 s for pandas.read_csv(engine="pyarrow"), so the documented path is not slower than before; the docs state usage facts only.
    Pre-existing note verified: with infer_string=False and dtype=str, pandas' PyArrow engine returns 'nan' strings for NULL (the C engine gives NaN). The helper reproduces this; it is out of scope here.
    Self-review of the repair: behavior: the prefix mirrors _adjust_column_names (PyAthena always passes names for .txt); just lint passed; test_result_set.py 32 passed. Claims: none of the PR text changes except the test list. AWS CI reruns on this push.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent follow-up on the repair (relayed): Codex CLI 0.160.0, model gpt-6-astra, read-only, session 01a104da-610d-7211-b8d2-395579d34123. This was a narrow review of 15399b55355736b8c4bf6711494e61e4f4063b78..2302049431ff2ee9d2b49ba271ae634cc5252a10 (detached snapshot unchanged), static only.

Verdict: CLEAN. All three points are resolved, and the reviewer found no new defect.

  • result_set.py:99 matches pandas' positional prefix; the caller always supplies names for header-less results, and the new case has three fields with one name.
  • test_result_set.py:228 reaches the to_datetime() failure path with "plain" under both string-inference settings.
  • test_result_set.py:190 uses the real eligibility check and option builder; the expected read still goes through pandas' PyArrow engine (engine="pyarrow" comes from the builder), and the separate dtype copies keep the comparison independent.

AWS CI at 2302049: test / run (3.14) passed (test-sqla* are skipped by the path filter, as pandas-only changes).

dtype = dict(read_csv_kwargs["dtype"])
for column in df.columns:
# Integer columns without a dtype entry get NumPy integer types.
if column not in dtype and df[column].dtype in integer_dtypes.values():
dtype[column] = df[column].dtype.numpy_dtype
dtype = {
column: pd.api.types.pandas_dtype(value)
for column, value in dtype.items()
if column in df.columns
}
df = df.astype(dtype)
if not pd.get_option("future.infer_string"):
# Without the string dtype, pandas returns strings, and the string
# categories of categorical columns, as objects.
for index in range(len(df.columns)):
values = df.iloc[:, index]
if values.dtype == "str":
df.isetitem(index, values.astype(object).fillna(None))
elif isinstance(values.dtype, pd.CategoricalDtype) and (
values.dtype.categories.dtype == "str"
):
categories = values.dtype.categories.astype(object)
df.isetitem(
index,
values.astype(pd.CategoricalDtype(categories, ordered=values.dtype.ordered)),
)
for column in read_csv_kwargs["parse_dates"]:
if isinstance(column, int) and column not in df.columns:
column = df.columns[column]
if df[column].dtype.kind in "Mm":
continue
values = df[column].astype("string")
try:
df[column] = pd.to_datetime(values, utc=False)
except (ValueError, TypeError):
# pandas keeps the column as strings if it cannot parse it.
df[column] = values.to_numpy(dtype=object, na_value=float("nan"))
# pandas applies the dtype mapping again after parsing dates.
df = df.astype(dtype)
return df


class _JSONConverter:
"""A json converter for ``pandas.read_csv()`` that keeps NULL from making values floats.

Expand Down Expand Up @@ -329,6 +427,8 @@ class AthenaPandasResultSet(AthenaResultSet):
"time",
"timestamp",
]
# The pandas.read_csv() options given to execute() that _read_csv_with_pyarrow() reads.
_PYARROW_READ_CSV_OPTIONS: ClassVar[frozenset[str]] = frozenset({"dtype", "parse_dates"})

def __init__(
self,
Expand Down Expand Up @@ -696,7 +796,10 @@ def _read_csv(self) -> TextFileReader | DataFrame:
source = self._csv_stream = stack.enter_context(
self._fs.open(self.output_location, mode="rb")
)
result = pd.read_csv(source, **read_csv_kwargs)
if csv_engine == "pyarrow" and self._reads_csv_with_pyarrow():
result = _read_csv_with_pyarrow(source, read_csv_kwargs)
else:
result = pd.read_csv(source, **read_csv_kwargs)
if not isinstance(result, pd.DataFrame):
# The chunk iterator takes ownership of the stream.
stack.pop_all()
Expand All @@ -716,6 +819,25 @@ def _read_csv(self) -> TextFileReader | DataFrame:
_logger.exception(f"Failed to read {self.output_location}.")
raise OperationalError(*e.args) from e

def _reads_csv_with_pyarrow(self) -> bool:
"""Whether ``_read_csv_with_pyarrow()`` reads the CSV result for the PyArrow engine.

It reproduces ``pandas.read_csv(engine="pyarrow")`` for PyAthena's default NA
values and for ``dtype`` as a mapping and ``parse_dates`` as a list given to
``execute()``. With other options, pandas reads the file.

Returns:
True if ``_read_csv_with_pyarrow()`` reads the result.
"""
return (
not self._keep_default_na
and isinstance(self._na_values, (list, tuple))
and list(self._na_values) == [""]
and self._kwargs.keys() <= self._PYARROW_READ_CSV_OPTIONS
and isinstance(self._kwargs.get("dtype", {}), dict)
and isinstance(self._kwargs.get("parse_dates", []), list)
)

def _get_csv_read_options(self, csv_engine: str, chunksize: int | None) -> dict[str, Any]:
"""Build pandas options for Athena CSV or tab-separated results."""
if self.output_location and self.output_location.endswith(".txt"):
Expand Down
61 changes: 60 additions & 1 deletion tests/pyathena/pandas/test_cursor.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,11 @@
from pyathena.model import AthenaQueryExecution
from pyathena.pandas.converter import DefaultPandasTypeConverter
from pyathena.pandas.cursor import PandasCursor
from pyathena.pandas.result_set import AthenaPandasResultSet, PandasDataFrameIterator
from pyathena.pandas.result_set import (
AthenaPandasResultSet,
PandasDataFrameIterator,
_read_csv_with_pyarrow,
)
from pyathena.util import RetryConfig
from tests import ENV
from tests.pyathena.conftest import connect
Expand Down Expand Up @@ -299,6 +303,24 @@ def test_csv_storage_options(self, pandas_cursor, query, expected, binary, with_
# pandas opens and closes the file itself.
assert pandas_cursor.result_set._csv_stream is None

def test_pyarrow_engine_multiline_values_across_blocks(self, pandas_cursor):
# The 2.4 MB result spans several 1 MiB pyarrow read blocks, and its
# values contain a newline and quotes.
with patch(
"pyathena.pandas.result_set._read_csv_with_pyarrow", wraps=_read_csv_with_pyarrow
) as read_csv_with_pyarrow:
pandas_cursor.execute(
"""
SELECT array_join(repeat('x', 600), '') || chr(10)
|| '"' || array_join(repeat('y', 598), '') || '"' AS v
FROM UNNEST(sequence(1, 2000)) AS t(i)
""",
engine="pyarrow",
)
df = pandas_cursor.as_pandas()
read_csv_with_pyarrow.assert_called_once()
assert df["v"].tolist() == ["x" * 600 + '\n"' + "y" * 598 + '"'] * 2000

@pytest.mark.parametrize(
("pandas_cursor", "parquet_engine", "chunksize"),
[
Expand Down Expand Up @@ -1042,6 +1064,43 @@ def test_get_csv_engine_explicit_specification(self):
engine = result_set._get_csv_engine()
assert engine == "c"

@pytest.mark.parametrize(
("keep_default_na", "na_values", "kwargs", "expected"),
[
(False, ("",), {}, True),
(False, ("",), {"dtype": {"a": "str"}, "parse_dates": ["b"]}, True),
(True, ("",), {}, False),
(False, ("", "NA"), {}, False),
(False, np.array([""]), {}, False),
(False, ("",), {"storage_options": None}, False),
(False, ("",), {"on_bad_lines": "skip"}, False),
(False, ("",), {"dtype": "str"}, False),
(False, ("",), {"parse_dates": "d"}, False),
(False, ("",), {"parse_dates": ("d",)}, False),
],
ids=[
"default",
"dtype_parse_dates",
"keep_default_na",
"na_values",
"na_values_array",
"storage_options",
"on_bad_lines",
"single_dtype",
"parse_dates_string",
"parse_dates_tuple",
],
)
def test_reads_csv_with_pyarrow(self, keep_default_na, na_values, kwargs, expected):
# PyAthena reads the CSV result for the PyArrow engine only with the options
# that _read_csv_with_pyarrow() reproduces; pandas reads it otherwise.
with patch("pyathena.pandas.result_set.AthenaResultSet.__init__"):
result_set = AthenaPandasResultSet.__new__(AthenaPandasResultSet)
result_set._keep_default_na = keep_default_na
result_set._na_values = na_values
result_set._kwargs = kwargs
assert result_set._reads_csv_with_pyarrow() is expected

@pytest.mark.parametrize(
("pandas_cursor", "parquet_engine", "chunksize"),
[
Expand Down
Loading
Loading