Problem
ArrowCursor fails to read a CSV result when a quoted value containing a newline crosses a block boundary of the pyarrow CSV reader.
AthenaArrowResultSet._read_csv() builds csv.ParseOptions without newlines_in_values (pyathena/arrow/result_set.py, the .csv branch), so pyarrow splits blocks at any newline and its parser gets out of sync with the chunker.
The default block_size is 128 MiB (AthenaArrowResultSet.DEFAULT_BLOCK_SIZE), so the error occurs with CSV results larger than that or with a smaller block_size passed to execute().
AsyncArrowCursor and AioArrowCursor share the result set.
Expected: the rows are read, as with the default block size for a small result.
Reproduction
Measured on Athena (S3 result file) with pyarrow 25.0.1:
from pyathena import connect
from pyathena.arrow.cursor import ArrowCursor
sql = """
SELECT array_join(repeat('x', 600), '') || chr(10) || array_join(repeat('y', 600), '') AS v
FROM UNNEST(sequence(1, 20)) AS t(i)
"""
cursor = connect(cursor_class=ArrowCursor).cursor()
cursor.execute(sql)
cursor.as_arrow().num_rows # 20
cursor.execute(sql, block_size=1024)
# OperationalError: CSV parser got out of sync with chunker. This can mean the data file
# contains cell values spanning multiple lines; please consider enabling the option
# 'newlines_in_values'.
Offline, with the same reader options and block_size=1024, newlines_in_values=True reads 50 rows of 300-byte two-line values correctly, while the current options raise the error above.
A single value larger than block_size still fails with newlines_in_values=True (straddling object straddles two block boundaries).
Environment
- PyAthena master (de8cc52), Python 3.13.1, pyarrow 25.0.1,
ArrowCursor.
Found during the review of #1031 (#1026).
Problem
ArrowCursorfails to read a CSV result when a quoted value containing a newline crosses a block boundary of the pyarrow CSV reader.AthenaArrowResultSet._read_csv()buildscsv.ParseOptionswithoutnewlines_in_values(pyathena/arrow/result_set.py, the.csvbranch), so pyarrow splits blocks at any newline and its parser gets out of sync with the chunker.The default
block_sizeis 128 MiB (AthenaArrowResultSet.DEFAULT_BLOCK_SIZE), so the error occurs with CSV results larger than that or with a smallerblock_sizepassed toexecute().AsyncArrowCursorandAioArrowCursorshare the result set.Expected: the rows are read, as with the default block size for a small result.
Reproduction
Measured on Athena (S3 result file) with pyarrow 25.0.1:
Offline, with the same reader options and
block_size=1024,newlines_in_values=Truereads 50 rows of 300-byte two-line values correctly, while the current options raise the error above.A single value larger than
block_sizestill fails withnewlines_in_values=True(straddling object straddles two block boundaries).Environment
ArrowCursor.Found during the review of #1031 (#1026).