Skip to content

Custom converters do not apply to pandas, Arrow, and Polars results with managed query result storage #1028

Description

@laughingman7743

Problem

With managed query result storage, the pandas, Arrow, and Polars cursors read every row through GetQueryResults and convert the values with DefaultTypeConverter (AthenaResultSet._fetch_all_rows()), not with the converter passed to connect(converter=...) or cursor(converter=...).
A custom converter therefore applies to S3 result files but not to managed results.
After #1005 (Polars) and #1023 (Arrow), the fetch methods do not run the cursor's converters on these already converted values, except for json columns in Arrow and Polars, which #1010 keeps as text and decodes with the cursor's json converter. Pandas never ran them on this path.
S3FSCursor uses its own converter on this path since #1010.

Expected: a custom converter gives the same values with managed storage as with an S3 result file.

Reproduction

Measured on Athena with #1010 applied (on master after #1023):

from pyathena import connect
from pyathena.arrow.converter import DefaultArrowTypeConverter
from pyathena.arrow.cursor import ArrowCursor

converter = DefaultArrowTypeConverter()
converter.set("varchar", lambda v: v.upper() if v is not None else v)

for kwargs in ({}, {"work_group": "<managed work group>", "s3_staging_dir": ""}):
    cursor = connect(cursor_class=ArrowCursor, converter=converter, **kwargs).cursor()
    cursor.execute("SELECT 1 AS id, 'abc' AS v")
    print(cursor.fetchall())
# S3 result file: [(1, 'ABC')]
# managed:        [(1, 'abc')]

PandasCursor (DefaultPandasTypeConverter) and PolarsCursor (DefaultPolarsTypeConverter) give the same results.
Before #1023, ArrowCursor ran its converters again on the managed values, so this example returned 'ABC', but the default converters raised TypeError for TIME, VARBINARY, and JSON objects.

Environment

Proposed fix (optional)

The cursor converters expect the text of the CSV result file, while the managed path builds the table or DataFrame from converted values.
One option is to keep the columns whose types the cursor's converter maps as text on the managed path and convert them in the fetch methods, as #1010 does for json. This changes the as_arrow() / as_polars() / as_pandas() column types for those types on the managed path, so the approach needs agreement first.

Found during the review of #1010 (#934).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions