Skip to content

PandasCursor warns DtypeWarning for large json columns with NULL #1078

Description

@laughingman7743

Problem

Since #1044 (not released yet), PandasCursor can emit DtypeWarning: Columns (...) have mixed types. Specify dtype option on import or set low_memory=False. for a json column with numeric (or boolean) values and NULLs, when the result is large enough that pandas' C parser converts it in several internal blocks.

_JSONConverter (pyathena/pandas/result_set.py :184 on master d989767) returns a sentinel object for NULL, so that the column stays object and restore() (:220) can put None back. With low_memory=True (pandas' default), the C parser infers a dtype per internal block. A block without NULL gets int64, and a block with the sentinel gets object. pandas then warns when it joins the blocks. Before #1044, NULL was converted to None, the blocks were int64 and float64, and they joined to float64 without a warning.

The values are correct: the column is object, with ints and None. Only the warning is new, and users cannot silence it through PyAthena without filtering warnings.

Reproduction

Measured with Athena, through PandasCursor on master d989767:

cursor = connect(..., cursor_class=PandasCursor).cursor()  # engine="auto" or "c"
df = cursor.execute("""
SELECT n, CASE WHEN n > 400000 THEN NULL ELSE CAST(n AS JSON) END AS j
FROM (
    SELECT a * 50000 + b AS n
    FROM UNNEST(sequence(0, 8)) AS s(a) CROSS JOIN UNNEST(sequence(1, 50000)) AS t(b)
)
WHERE n <= 400010
ORDER BY n
""").as_pandas()
# DtypeWarning: Columns (1: j) have mixed types. Specify dtype option on import or set low_memory=False.
# df["j"].dtype == object, 10 None values

The same query with only the j column did not warn. Offline, pandas.read_csv(..., converters={"j": _JSONConverter(conv)}, engine="c") on a two-column CSV with 400,000 numeric j values followed by 10 empty ones warns. With the plain converter, which returns None, it gives float64 and no warning.

Environment

PyAthena master d989767, pandas 3.0.6, Python 3.13.1, PandasCursor with engine="auto" and engine="c".

Proposed fix (optional)

Keep every block object, so that pandas never sees mixed block dtypes. One option is an object dtype entry for json columns that have a _JSONConverter, if pandas applies it together with the converter. Another is to suppress DtypeWarning for those columns only, around read_csv(). Check both against the cost limits of #1044 (no per-value allocation, no extra column scans), and validate with the reproduction above plus an offline multi-block test.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions