Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions docs/api/sqlalchemy.rst
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,9 @@ Compilers
.. autoclass:: pyathena.sqlalchemy.compiler.AthenaTypeCompiler
:members:

.. autoclass:: pyathena.sqlalchemy.compiler.AthenaDMLTypeCompiler
:members:

.. autoclass:: pyathena.sqlalchemy.compiler.AthenaStatementCompiler
:members:

Expand Down
29 changes: 23 additions & 6 deletions docs/sqlalchemy.md
Original file line number Diff line number Diff line change
Expand Up @@ -737,6 +737,25 @@ engine_arrow = create_engine(
)
```

## DDL and CAST types

Athena parses DDL statements such as `CREATE TABLE` with Hive type syntax, and queries with Trino type syntax.
`CREATE TABLE` column types, and types compiled with `TypeEngine.compile()`, use the Hive syntax.
`CAST` target types use the Trino syntax.

| SQLAlchemy type | Table DDL | CAST |
|---|---|---|
| `Integer`, `INTEGER` | `INT` | `INTEGER` |
| `String`, `Text`, `CLOB` | `STRING` | `VARCHAR` |
| `LargeBinary`, `BINARY`, `VARBINARY` | `BINARY` | `VARBINARY` |
| `JSON` | Raises `CompileError` | `JSON` |
| `AthenaStruct` | `STRUCT<name:type, ...>` | `ROW(name type, ...)` |
| `AthenaMap` | `MAP<key, value>` | `MAP(key, value)` |
| `AthenaArray`, `ARRAY` | `ARRAY<item>` | `ARRAY(item)` |

Complex types apply the same syntax to their nested types.
An `AthenaStruct` without fields raises `CompileError` in both.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round two finding (repaired in 684eccd4fa1d5e801e94c66cceed31d4b9d84b94). The sentence "...raises CompileError in both, because Athena has no empty STRUCT type" carried a rationale aside, which user docs leave out. It now states only the behavior.


## Floating-point types

| SQLAlchemy type | Table DDL | CAST |
Expand Down Expand Up @@ -824,10 +843,8 @@ CREATE TABLE users (
)
```

`CREATE TABLE` renders `AthenaStruct` columns with Hive `STRUCT<name:type, ...>` syntax at every nesting depth.
That includes top-level columns, fields of a STRUCT, STRUCT values inside MAP, and STRUCT values inside ARRAY.
Integer fields, and integer MAP keys and values, use `INT` in that DDL.
`CAST` and other SQL expressions keep `ROW(...)`, `MAP(...)`, and `ARRAY(...)`, and spell integers as `INTEGER`.
`CREATE TABLE` renders `AthenaStruct` columns with Hive `STRUCT<name:type, ...>` syntax at every nesting depth, and `CAST` renders them as `ROW(name type, ...)`.
See [DDL and CAST types](#ddl-and-cast-types).

#### Querying STRUCT data

Expand Down Expand Up @@ -1122,7 +1139,7 @@ An outer `TypeDecorator` retains its result processor as well as native ARRAY or
Raw `text()` queries and direct DB API queries retain the cursor's existing conversion behavior described below; they do not receive this projection automatically.

Compared with earlier releases, reflected ARRAY columns are no longer reported as `String`.
ARRAY DDL now renders integer elements as `INT` and row elements as `STRUCT<...>`, which Athena requires for nested DDL types.
ARRAY DDL now renders integer elements as `INT` and row elements as `STRUCT<...>`.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round two finding (repaired in 684eccd4fa1d5e801e94c66cceed31d4b9d84b94). This pre-existing sentence, outside the original diff, said INT is "required for nested DDL types". #886 measured that CREATE EXTERNAL TABLE ... (a MAP<STRING, INTEGER>) parses, so the requirement was false for INT. Removed the clause; the section otherwise matches the new type syntax section.

Code that inspects reflected types or compares compiled SQL strings should account for these changes.

#### Basic Usage
Expand Down Expand Up @@ -1367,7 +1384,7 @@ Athena's JSON type support has specific limitations:

- **JSON objects and arrays are supported** - `CAST('...' AS JSON)` accepts an object or a top-level array such as `[1, 2, 3]`
- **Arrays within objects are supported** - JSON objects can contain arrays as property values
- **DML only** - JSON type is supported for SELECT queries but not in CREATE TABLE statements
- **DML only** - JSON type is supported for SELECT queries but not in CREATE TABLE statements; compiling `CREATE TABLE` with a `JSON` column raises `CompileError`

```python
# Supported: JSON object with nested array
Expand Down
31 changes: 28 additions & 3 deletions pyathena/sqlalchemy/array.py
Original file line number Diff line number Diff line change
Expand Up @@ -235,6 +235,27 @@ def decorator_impl(self, type_: types.TypeDecorator[Any]) -> TypeEngine[Any]:
return variant
return type_.load_dialect_impl(self.dialect)

def dialect_type(self, type_: TypeEngine[Any]) -> TypeEngine[Any]:
"""Resolve the type this dialect uses for a SQLAlchemy type.

Takes the Athena variant from ``with_variant()`` and the implementation
of a TypeDecorator until neither applies.

Args:
type_: The declared type.

Returns:
The resolved type.
"""
while True:
variant = self.variant(type_)
if variant is not None:
type_ = variant
elif isinstance(type_, types.TypeDecorator):
type_ = self.decorator_impl(type_)
else:
return type_

@staticmethod
def has_unknown_element(type_: TypeEngine[Any]) -> bool:
if isinstance(type_, sqltypes.ARRAY):
Expand Down Expand Up @@ -579,7 +600,9 @@ def process(self, expression, **kw):
):
raise exc.CompileError("An ARRAY slice assignment requires a non-NULL array")
rhs = compiler.process(value, **kw)
rhs_type = compiler._complex_dml_type(expression.value_type, require_precision=True)
rhs_type = compiler._dml_type_compiler.process_element(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex finding 3 (repaired in 857014657720868939b8132c6dd89fa62a585315): this RHS type now goes through process_element(), so NullType (including one reached through with_variant) raises "Bound ARRAY values require an explicit element type" as on base. Covered by test_array_assignment_rejects_unknown_value_type.

expression.value_type, require_precision=True
)
rhs = f"CAST({rhs} AS {rhs_type})"
if final_slice:
# Reject SQL expressions that evaluate to NULL without issuing a second statement.
Expand Down Expand Up @@ -626,7 +649,7 @@ def _rebuild(self, array, array_type, path, rhs, **kw):
array_type = self._type_inspector.array_type(array_type)
if array_type is None:
raise exc.CompileError("Partial ARRAY updates require an ARRAY column type")
array_sql_type = compiler._complex_dml_type(array_type)
array_sql_type = compiler._dml_type_compiler.process_element(array_type)
array = f"coalesce({array}, CAST(ARRAY[] AS {array_sql_type}))"
bound = path[0]
if isinstance(bound, Slice):
Expand All @@ -635,7 +658,9 @@ def _rebuild(self, array, array_type, path, rhs, **kw):

def _prefix_and_padding(self, array, start, array_type):
prefix = f"slice({array}, 1, least({start} - 1, cardinality({array})))"
element_type = self.compiler._complex_dml_type(_ArrayTypeInspector.item_type(array_type))
element_type = self.compiler._dml_type_compiler.process_element(
_ArrayTypeInspector.item_type(array_type)
)
padding = (
f"repeat(CAST(NULL AS {element_type}), "
f"CAST(greatest({start} - 1 - cardinality({array}), 0) AS INTEGER))"
Expand Down
Loading
Loading