Skip to content

Complex-value parsing loses values (typed array(json), native arrays) and PandasCursor NULL TIME differs by result storage #1041

Description

@laughingman7743

Problem

Three defects in result value conversion, found by the independent reviews of #1033.
Measured on 2026-10-04 on master 9a373fd and on the #1033 branch.

1. A typed array(json) value whose first element is not a string can fail to parse

TypedValueConverter._convert_typed_array() (pyathena/parser.py) decides whether to try json.loads() from the first few characters of the value (a quote in value[1:10], or a leading [{, [null, [[).
Otherwise it uses the native-format splitter, which splits a JSON string element at its commas.

cursor.execute(
    """SELECT ARRAY[json_parse('123456789'), json_parse('"a,b"')] AS v""",
    result_set_type_hints={"v": "array(json)"},
)
cursor.fetchone()  # JSONDecodeError: Unterminated string starting at: line 1 column 1 (char 0)

Without the type hint, the same value converts to [123456789, 'a,b'].
On master, the same array with the string first also failed; #1033 fixes that order (its JSON elements are re-encoded with json.dumps()), but not this one.
The same check applies to arrays nested in maps and rows.

2. PandasCursor returns NaT for a NULL TIME from a result file and None from GetQueryResults

SELECT * FROM (VALUES CAST('12:34:56' AS TIME), NULL) AS t(v):

Result storage fetchall() as_pandas()["v"]
S3 result file [(time(12, 34, 56),), (NaT,)] [time(12, 34, 56), NaT]
managed (GetQueryResults) [(time(12, 34, 56),), (None,)] [time(12, 34, 56), None]

The result-file path parses time columns with parse_dates and _trunc_date()'s .dt.time, which keeps NULL as NaT; the fallback builds the column from DefaultTypeConverter values.
After #1033, time with time zone columns use a converter in both modes and return None.

3. Native-format arrays lose empty and space-only strings and split elements at commas

Measured on master abc99b0 (2026-10-04). The text is the same from GetQueryResults and from the CSV result file, and the typed (array(varchar)) and untyped conversions agree:

Query value Athena text Converted
ARRAY['x', '', 'y'] [x, , y] ['x', 'y']
ARRAY['x', ''] [x, ] ['x']
ARRAY['x', ' ', 'y'] [x, , y] ['x', 'y']
ARRAY['a,b', 'c'] [a,b, c] ['a', 'b', 'c']
ARRAY[''] [] []
ARRAY['null', 'NULL'] [null, NULL] [None, None]
MAP(ARRAY['k', ''], ARRAY['', 'v']) {=v, k=} {'': 'v', 'k': ''}
CAST(ROW('', 'x') AS ROW(a varchar, b varchar)) {a=, b=x} {'a': '', 'b': 'x'}

The array splitter (_split_array_items() and the native branch of _convert_typed_array() in pyathena/parser.py) splits at every comma and drops empty and whitespace-only items.
Athena separates elements with ", ", so splitting at ", " and keeping empty items would restore the first four rows.
ARRAY[''] versus ARRAY[], the text 'null' versus NULL, and elements that contain ", " cannot be told apart from the text.
Maps and rows already keep empty strings.

Expected

  1. A typed array(json) value converts as JSON regardless of the type of its first element.
  2. PandasCursor represents a NULL TIME the same way in both result storage modes.
  3. Native-format arrays keep empty and whitespace-only elements and do not split elements at a comma that is not followed by a space, where the text allows it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions