Problem
With managed query result storage, the pandas, Arrow, and Polars cursors read every row through GetQueryResults and convert the values with DefaultTypeConverter (AthenaResultSet._fetch_all_rows()), not with the converter passed to connect(converter=...) or cursor(converter=...).
A custom converter therefore applies to S3 result files but not to managed results.
After #1005 (Polars) and #1023 (Arrow), the fetch methods do not run the cursor's converters on these already converted values, except for json columns in Arrow and Polars, which #1010 keeps as text and decodes with the cursor's json converter. Pandas never ran them on this path.
S3FSCursor uses its own converter on this path since #1010.
Expected: a custom converter gives the same values with managed storage as with an S3 result file.
Reproduction
Measured on Athena with #1010 applied (on master after #1023):
from pyathena import connect
from pyathena.arrow.converter import DefaultArrowTypeConverter
from pyathena.arrow.cursor import ArrowCursor
converter = DefaultArrowTypeConverter()
converter.set("varchar", lambda v: v.upper() if v is not None else v)
for kwargs in ({}, {"work_group": "<managed work group>", "s3_staging_dir": ""}):
cursor = connect(cursor_class=ArrowCursor, converter=converter, **kwargs).cursor()
cursor.execute("SELECT 1 AS id, 'abc' AS v")
print(cursor.fetchall())
# S3 result file: [(1, 'ABC')]
# managed: [(1, 'abc')]
PandasCursor (DefaultPandasTypeConverter) and PolarsCursor (DefaultPolarsTypeConverter) give the same results.
Before #1023, ArrowCursor ran its converters again on the managed values, so this example returned 'ABC', but the default converters raised TypeError for TIME, VARBINARY, and JSON objects.
Environment
Proposed fix (optional)
The cursor converters expect the text of the CSV result file, while the managed path builds the table or DataFrame from converted values.
One option is to keep the columns whose types the cursor's converter maps as text on the managed path and convert them in the fetch methods, as #1010 does for json. This changes the as_arrow() / as_polars() / as_pandas() column types for those types on the managed path, so the approach needs agreement first.
Found during the review of #1010 (#934).
Problem
With managed query result storage, the pandas, Arrow, and Polars cursors read every row through
GetQueryResultsand convert the values withDefaultTypeConverter(AthenaResultSet._fetch_all_rows()), not with the converter passed toconnect(converter=...)orcursor(converter=...).A custom converter therefore applies to S3 result files but not to managed results.
After #1005 (Polars) and #1023 (Arrow), the fetch methods do not run the cursor's converters on these already converted values, except for
jsoncolumns in Arrow and Polars, which #1010 keeps as text and decodes with the cursor'sjsonconverter. Pandas never ran them on this path.S3FSCursoruses its own converter on this path since #1010.Expected: a custom converter gives the same values with managed storage as with an S3 result file.
Reproduction
Measured on Athena with #1010 applied (on master after #1023):
PandasCursor(DefaultPandasTypeConverter) andPolarsCursor(DefaultPolarsTypeConverter) give the same results.Before #1023,
ArrowCursorran its converters again on the managed values, so this example returned'ABC', but the default converters raisedTypeErrorfor TIME, VARBINARY, and JSON objects.Environment
Proposed fix (optional)
The cursor converters expect the text of the CSV result file, while the managed path builds the table or DataFrame from converted values.
One option is to keep the columns whose types the cursor's converter maps as text on the managed path and convert them in the fetch methods, as #1010 does for
json. This changes theas_arrow()/as_polars()/as_pandas()column types for those types on the managed path, so the approach needs agreement first.Found during the review of #1010 (#934).