Problem
Since #1044 (not released yet), PandasCursor can emit DtypeWarning: Columns (...) have mixed types. Specify dtype option on import or set low_memory=False. for a json column with numeric (or boolean) values and NULLs, when the result is large enough that pandas' C parser converts it in several internal blocks.
_JSONConverter (pyathena/pandas/result_set.py :184 on master d989767) returns a sentinel object for NULL, so that the column stays object and restore() (:220) can put None back. With low_memory=True (pandas' default), the C parser infers a dtype per internal block. A block without NULL gets int64, and a block with the sentinel gets object. pandas then warns when it joins the blocks. Before #1044, NULL was converted to None, the blocks were int64 and float64, and they joined to float64 without a warning.
The values are correct: the column is object, with ints and None. Only the warning is new, and users cannot silence it through PyAthena without filtering warnings.
Reproduction
Measured with Athena, through PandasCursor on master d989767:
cursor = connect(..., cursor_class=PandasCursor).cursor() # engine="auto" or "c"
df = cursor.execute("""
SELECT n, CASE WHEN n > 400000 THEN NULL ELSE CAST(n AS JSON) END AS j
FROM (
SELECT a * 50000 + b AS n
FROM UNNEST(sequence(0, 8)) AS s(a) CROSS JOIN UNNEST(sequence(1, 50000)) AS t(b)
)
WHERE n <= 400010
ORDER BY n
""").as_pandas()
# DtypeWarning: Columns (1: j) have mixed types. Specify dtype option on import or set low_memory=False.
# df["j"].dtype == object, 10 None values
The same query with only the j column did not warn. Offline, pandas.read_csv(..., converters={"j": _JSONConverter(conv)}, engine="c") on a two-column CSV with 400,000 numeric j values followed by 10 empty ones warns. With the plain converter, which returns None, it gives float64 and no warning.
Environment
PyAthena master d989767, pandas 3.0.6, Python 3.13.1, PandasCursor with engine="auto" and engine="c".
Proposed fix (optional)
Keep every block object, so that pandas never sees mixed block dtypes. One option is an object dtype entry for json columns that have a _JSONConverter, if pandas applies it together with the converter. Another is to suppress DtypeWarning for those columns only, around read_csv(). Check both against the cost limits of #1044 (no per-value allocation, no extra column scans), and validate with the reproduction above plus an offline multi-block test.
Problem
Since #1044 (not released yet),
PandasCursorcan emitDtypeWarning: Columns (...) have mixed types. Specify dtype option on import or set low_memory=False.for a json column with numeric (or boolean) values and NULLs, when the result is large enough that pandas' C parser converts it in several internal blocks._JSONConverter(pyathena/pandas/result_set.py:184 on master d989767) returns a sentinel object for NULL, so that the column stays object andrestore()(:220) can put None back. Withlow_memory=True(pandas' default), the C parser infers a dtype per internal block. A block without NULL gets int64, and a block with the sentinel gets object. pandas then warns when it joins the blocks. Before #1044, NULL was converted to None, the blocks were int64 and float64, and they joined to float64 without a warning.The values are correct: the column is object, with ints and None. Only the warning is new, and users cannot silence it through PyAthena without filtering warnings.
Reproduction
Measured with Athena, through
PandasCursoron master d989767:The same query with only the
jcolumn did not warn. Offline,pandas.read_csv(..., converters={"j": _JSONConverter(conv)}, engine="c")on a two-column CSV with 400,000 numericjvalues followed by 10 empty ones warns. With the plain converter, which returns None, it gives float64 and no warning.Environment
PyAthena master d989767, pandas 3.0.6, Python 3.13.1,
PandasCursorwithengine="auto"andengine="c".Proposed fix (optional)
Keep every block object, so that pandas never sees mixed block dtypes. One option is an
objectdtype entry for json columns that have a_JSONConverter, if pandas applies it together with the converter. Another is to suppressDtypeWarningfor those columns only, aroundread_csv(). Check both against the cost limits of #1044 (no per-value allocation, no extra column scans), and validate with the reproduction above plus an offline multi-block test.