Problem
When columns with the same name have different Athena types, PandasCursor and ArrowCursor fail to read the S3 CSV result file, because the type maps they pass to the CSV reader are keyed by column name, and the reader applies an entry to every column with that name.
SELECT 1 AS x, 'a1' AS x, CAST('12:34:56' AS TIME) AS x, measured on Athena on 2026-10-04 with #1050 applied (8e5dda6):
| Cursor |
S3 result file |
managed (GetQueryResults) |
PandasCursor |
OperationalError: Unable to parse string "12:34:56.000" at position 0 |
[(1, 'a1', time(12, 34, 56))] |
ArrowCursor |
ValueError: time data '1' does not match format '%H:%M:%S.%f' |
[(1, 'a1', time(12, 34, 56))] |
PolarsCursor |
[(1, 'a1', time(12, 34, 56))] |
[(1, 'a1', time(12, 34, 56))] |
- pandas.
pandas.read_csv(dtype={"x": ...}) applies the dtype to x and to the renamed x.1, x.2 (checked with pandas 3.0.6). The integer dtype of the first x is applied to the varchar and time columns.
- Arrow.
pyarrow.csv.ConvertOptions(column_types={"x": ...}) types every column named x, and AthenaArrowResultSet.converters keeps one converter per name, so the converter of the last x is applied to all of them.
#1032 (PR #1050) fixed columns with the same name and the same type. It keeps converters and parse_dates per column under pandas' and Polars' renamed names, but it does not change how the pandas dtype and pyarrow column_types maps match these columns.
Before #1050, the same query also failed for PolarsCursor on the S3 path and for all three cursors on the managed path.
Expected
Columns with the same name and different types convert by their own types when read from the S3 result file, as they do through GetQueryResults and in Cursor.
Environment
Problem
When columns with the same name have different Athena types,
PandasCursorandArrowCursorfail to read the S3 CSV result file, because the type maps they pass to the CSV reader are keyed by column name, and the reader applies an entry to every column with that name.SELECT 1 AS x, 'a1' AS x, CAST('12:34:56' AS TIME) AS x, measured on Athena on 2026-10-04 with #1050 applied (8e5dda6):PandasCursorOperationalError: Unable to parse string "12:34:56.000" at position 0[(1, 'a1', time(12, 34, 56))]ArrowCursorValueError: time data '1' does not match format '%H:%M:%S.%f'[(1, 'a1', time(12, 34, 56))]PolarsCursor[(1, 'a1', time(12, 34, 56))][(1, 'a1', time(12, 34, 56))]pandas.read_csv(dtype={"x": ...})applies the dtype toxand to the renamedx.1,x.2(checked with pandas 3.0.6). Theintegerdtype of the firstxis applied to the varchar and time columns.pyarrow.csv.ConvertOptions(column_types={"x": ...})types every column namedx, andAthenaArrowResultSet.converterskeeps one converter per name, so the converter of the lastxis applied to all of them.#1032 (PR #1050) fixed columns with the same name and the same type. It keeps
convertersandparse_datesper column under pandas' and Polars' renamed names, but it does not change how the pandasdtypeand pyarrowcolumn_typesmaps match these columns.Before #1050, the same query also failed for
PolarsCursoron the S3 path and for all three cursors on the managed path.Expected
Columns with the same name and different types convert by their own types when read from the S3 result file, as they do through GetQueryResults and in
Cursor.Environment
uv.lock.DUPLICATE_COLUMN_NAME).