Skip to content

ArrowCursor drops NULL rows of single-column CSV results #1026

Description

@laughingman7743

Problem

ArrowCursor drops the rows of a single-column CSV result whose value is SQL NULL.
Athena writes such a row as an empty line, and the pyarrow CSV reader skips empty lines (ignore_empty_lines=not binary_columns in pyathena/arrow/result_set.py:325, so only results with a varbinary column keep them).
The rows are lost for every column type, from fetchall() and from as_arrow().
PandasCursor, PolarsCursor, and S3FSCursor return the NULL rows.

Expected: [(None,), (1,), (2,)] and a 3-row table.

Reproduction

Measured on Athena (S3 result file, not managed storage) on master 16aef64:

from pyathena import connect
from pyathena.arrow.cursor import ArrowCursor

cursor = connect(cursor_class=ArrowCursor).cursor()
cursor.execute("SELECT x FROM (VALUES 1, NULL, 2) AS t(x) ORDER BY x NULLS FIRST")
cursor.fetchall()  # [(1,), (2,)]

cursor.execute("SELECT CAST(NULL AS VARCHAR) AS v")
cursor.fetchall()  # []
cursor.as_arrow().num_rows  # 0

The same queries return 3 rows and 1 row with PandasCursor, PolarsCursor, and S3FSCursor.

Environment

  • PyAthena master 16aef64, Python 3.13.1, pyarrow 25.0.1, ArrowCursor (sync; AsyncArrowCursor and AioArrowCursor share the result set).

Found during the review of #1010 (#934).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions