Skip to content

PolarsCursor with unload=True and chunksize returns the UNLOAD statement's metadata and None rows #1007

Description

@laughingman7743

Problem

With PolarsCursor(unload=True, chunksize=N), description and the fetch methods keep the metadata of the UNLOAD statement itself instead of the columns in the Parquet files, so every fetched row has the wrong shape and is filled with None.

In the non-chunked path, AthenaPolarsResultSet._as_polars() replaces _metadata with the Parquet schema after reading the files. The chunked path (_create_dataframe_iterator() → _iter_parquet_chunks()) never does, and __init__ caches the column names and converters from the UNLOAD statement's metadata (pyathena/polars/result_set.py, _create_dataframe_iterator() and _column_names_cache).
iterrows() then looks up those names in the Parquet chunk with row_dict.get(col), which returns None.

The existing chunked-UNLOAD tests only check row counts.

Reproduction

Measured on master a18ebda (2026-10-03):

from pyathena import connect
from pyathena.polars.cursor import PolarsCursor

cursor = connect(cursor_class=PolarsCursor).cursor(unload=True, chunksize=10)
cursor.execute("SELECT 1 AS a, 'x' AS b")
cursor.description  # [('rows', 'bigint', None, None, 19, 0, 'UNKNOWN')]
cursor.fetchone()   # (None,)

cursor = connect(cursor_class=PolarsCursor).cursor(unload=True)
cursor.execute("SELECT 1 AS a, 'x' AS b")
cursor.description  # [('a', 'integer', ...), ('b', 'varchar', ...)]
cursor.fetchone()   # (1, 'x')

Expected

With chunksize, description and the fetched rows match the non-chunked UNLOAD result.

Found by the independent review of #1005 (#935).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions