Problem
ArrowCursor drops the rows of a single-column CSV result whose value is SQL NULL.
Athena writes such a row as an empty line, and the pyarrow CSV reader skips empty lines (ignore_empty_lines=not binary_columns in pyathena/arrow/result_set.py:325, so only results with a varbinary column keep them).
The rows are lost for every column type, from fetchall() and from as_arrow().
PandasCursor, PolarsCursor, and S3FSCursor return the NULL rows.
Expected: [(None,), (1,), (2,)] and a 3-row table.
Reproduction
Measured on Athena (S3 result file, not managed storage) on master 16aef64:
from pyathena import connect
from pyathena.arrow.cursor import ArrowCursor
cursor = connect(cursor_class=ArrowCursor).cursor()
cursor.execute("SELECT x FROM (VALUES 1, NULL, 2) AS t(x) ORDER BY x NULLS FIRST")
cursor.fetchall() # [(1,), (2,)]
cursor.execute("SELECT CAST(NULL AS VARCHAR) AS v")
cursor.fetchall() # []
cursor.as_arrow().num_rows # 0
The same queries return 3 rows and 1 row with PandasCursor, PolarsCursor, and S3FSCursor.
Environment
- PyAthena master 16aef64, Python 3.13.1, pyarrow 25.0.1,
ArrowCursor (sync; AsyncArrowCursor and AioArrowCursor share the result set).
Found during the review of #1010 (#934).
Problem
ArrowCursordrops the rows of a single-column CSV result whose value is SQL NULL.Athena writes such a row as an empty line, and the pyarrow CSV reader skips empty lines (
ignore_empty_lines=not binary_columnsinpyathena/arrow/result_set.py:325, so only results with avarbinarycolumn keep them).The rows are lost for every column type, from
fetchall()and fromas_arrow().PandasCursor,PolarsCursor, andS3FSCursorreturn the NULL rows.Expected:
[(None,), (1,), (2,)]and a 3-row table.Reproduction
Measured on Athena (S3 result file, not managed storage) on master 16aef64:
The same queries return 3 rows and 1 row with
PandasCursor,PolarsCursor, andS3FSCursor.Environment
ArrowCursor(sync;AsyncArrowCursorandAioArrowCursorshare the result set).Found during the review of #1010 (#934).