Problem
Three defects in the pandas, Arrow, and Polars result sets, found by the independent reviews of #1023 and #1029.
All results below were measured on master 9a373fd (2026-10-03). "managed" means managed query result storage, where these cursors read every row through the GetQueryResults fallback (AthenaResultSet._fetch_all_rows()).
1. Duplicate column names lose or mix values
SELECT 1 AS x, 2 AS x; the expected row is (1, 2):
| Cursor |
S3 result file |
managed |
PandasCursor |
[(1, 1)] |
[(1, 1), (2, 2)] |
ArrowCursor |
[(2,)] |
[(1,), (2,)] |
PolarsCursor |
[(1, 1)] |
[(1, 1), (2, 2)] |
- The fallback builds its table with
_rows_to_columnar() (pyathena/result_set.py), which keys the columns by name, so both values are appended to one column and every row becomes two rows.
- On the S3 path, the fetch methods look up values by column name: pandas and Polars rows are dicts keyed by name (
row[1][d[0]], row_dict.get(col)), and Arrow uses RecordBatch.to_pydict(), which keeps one column per name.
2. The fallback drops a data row that looks like the header at the start of a later page
_fetch_all_rows() runs _is_first_row_column_labels() on every GetQueryResults page, while _pre_fetch() checks only the first page.
SELECT CASE WHEN n = 1000 THEN 'a' ELSE CAST(n AS varchar) END AS a FROM UNNEST(sequence(1, 1500)) AS t(n) ORDER BY n returns 1500 rows with the S3 result file but 1499 rows (the 'a' row is missing) on the managed path, for all three cursors.
3. AthenaArrowResultSet fetch after close() raises an incidental TypeError
close() assigns self._batches = [], so a later fetch calls next([]) and raises TypeError: 'list' object is not an iterator instead of returning no rows (pyathena/arrow/result_set.py).
Expected
- Columns with the same name keep their own values in both modes, as
Cursor does.
- Only the first GetQueryResults page can carry the header row.
- Fetching from a closed Arrow result set behaves like the other result sets (no incidental
TypeError).
Problem
Three defects in the pandas, Arrow, and Polars result sets, found by the independent reviews of #1023 and #1029.
All results below were measured on master 9a373fd (2026-10-03). "managed" means managed query result storage, where these cursors read every row through the GetQueryResults fallback (
AthenaResultSet._fetch_all_rows()).1. Duplicate column names lose or mix values
SELECT 1 AS x, 2 AS x; the expected row is(1, 2):PandasCursor[(1, 1)][(1, 1), (2, 2)]ArrowCursor[(2,)][(1,), (2,)]PolarsCursor[(1, 1)][(1, 1), (2, 2)]_rows_to_columnar()(pyathena/result_set.py), which keys the columns by name, so both values are appended to one column and every row becomes two rows.row[1][d[0]],row_dict.get(col)), and Arrow usesRecordBatch.to_pydict(), which keeps one column per name.2. The fallback drops a data row that looks like the header at the start of a later page
_fetch_all_rows()runs_is_first_row_column_labels()on every GetQueryResults page, while_pre_fetch()checks only the first page.SELECT CASE WHEN n = 1000 THEN 'a' ELSE CAST(n AS varchar) END AS a FROM UNNEST(sequence(1, 1500)) AS t(n) ORDER BY nreturns 1500 rows with the S3 result file but 1499 rows (the'a'row is missing) on the managed path, for all three cursors.3.
AthenaArrowResultSetfetch afterclose()raises an incidentalTypeErrorclose()assignsself._batches = [], so a later fetch callsnext([])and raisesTypeError: 'list' object is not an iteratorinstead of returning no rows (pyathena/arrow/result_set.py).Expected
Cursordoes.TypeError).