Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
55 commits
Select commit Hold shift + click to select a range
db61309
protocol: add v2 protocol design
tkilias Aug 14, 2026
12e29d2
protocol: add run batch correlation
tkilias Aug 14, 2026
695d2b1
protocol: mark completed run groups
tkilias Aug 17, 2026
95841aa
protocol: define high-level payload contracts
tkilias Aug 17, 2026
486798b
protocol: remove legacy metadata fields
tkilias Aug 17, 2026
35c21be
Use Draft-07 v2 JSON schemas
tkilias Aug 18, 2026
90b51c4
Reorganize v2 protocol design docs
tkilias Aug 18, 2026
ab08878
docs(v2): define high-level Exasol-to-Arrow type mapping
tkilias Aug 24, 2026
a4822d6
docs(v2): expose Exasol column type parameters
tkilias Aug 24, 2026
d615498
docs(v2): name Exasol type enum
tkilias Aug 26, 2026
a53378c
docs(v2): centralize schemas and examples with CI validation
tkilias Aug 26, 2026
ffedcd5
Merge origin/main into documentation/design_v2_protocol
tkilias Aug 26, 2026
cf85adb
docs(v2): map decimals to sized Arrow types
tkilias Aug 26, 2026
f458fb1
docs(v2): streamline type mapping metadata
tkilias Aug 26, 2026
834ff7f
docs: consolidate v2 data stream references
tkilias Aug 27, 2026
5a08777
docs: place call lifecycle under low-level protocol
tkilias Aug 27, 2026
b83f88f
docs: consolidate low-level data stream references
tkilias Aug 27, 2026
2742461
some cleanup
tkilias Aug 27, 2026
158f8c1
docs: use fixed-width year-month interval storage
tkilias Aug 27, 2026
a584dfa
sort files into sub directories
tkilias Aug 27, 2026
9368b05
docs: address v2 protocol review comments
tkilias Aug 28, 2026
8ee11b9
Add Arrow schema types and variadic buffers
tkilias Sep 1, 2026
4bc2774
Merge origin/main into documentation/design_v2_protocol
tkilias Sep 10, 2026
ce2ddfa
Extract shared column schema definition
tkilias Sep 10, 2026
b793157
Make extracted column schema compatible with validator
tkilias Sep 10, 2026
c7b68c7
Add schema loader diagnostics
tkilias Sep 10, 2026
2c0393e
Fix external schema loader dispatch
tkilias Sep 10, 2026
e272597
Inline column schemas and test supported types
tkilias Sep 11, 2026
c934ed0
protocol: separate control messages from record batches
tkilias Sep 14, 2026
999502b
protocol: model control message variants
tkilias Sep 14, 2026
4141d89
docs: describe record batch and inline buffer encoding
tkilias Sep 14, 2026
8504839
docs: define Arrow schema traversal order
tkilias Sep 14, 2026
2af4983
docs: clarify Next flow control semantics
tkilias Sep 15, 2026
6be0359
docs: clarify correlation columns in record batches
tkilias Sep 17, 2026
38dbd77
docs: link Arrow run-end encoding
tkilias Sep 17, 2026
f7d5be9
docs: specify range_run extension encoding
tkilias Sep 17, 2026
ddca562
docs: clarify group and row ID encoding examples
tkilias Sep 17, 2026
244859a
protocol: rename record batch metadata
tkilias Sep 17, 2026
5c66c01
docs: consolidate metadata lifecycle rules
tkilias Sep 17, 2026
8d34f2f
docs: extract group and row correlation
tkilias Sep 17, 2026
ae1c7a0
docs: clarify CHAR padding semantics
tkilias Sep 19, 2026
f7c0fc7
docs: clarify asynchronous call closure
tkilias Sep 20, 2026
4c403b6
docs: clarify row_id default handling
tkilias Sep 20, 2026
2cbccef
Update doc/design/v2/protocol/high_level/type_mapping.md
tkilias Sep 20, 2026
fb359ed
protocol: advertise server capabilities
tkilias Sep 21, 2026
7f1dbc4
protocol: advertise high-level protocol identity
tkilias Sep 21, 2026
9aa7d1d
docs: clarify script metadata timing
tkilias Sep 22, 2026
e44be61
docs: define frame length prefix
tkilias Sep 22, 2026
a59f7b1
docs: add bidirectional data scheduling paths
tkilias Sep 22, 2026
591e445
allow column metadata in call metadata
tkilias Sep 23, 2026
b7a8ec4
docs: add protocol cleanup call and quiz
tkilias Sep 23, 2026
36a9b14
docs: move protocol quiz to dedicated branch
tkilias Sep 23, 2026
d21b6da
Merge origin/main into protocol design
tkilias Sep 26, 2026
8073495
#70: Format v2 protocol test
tkilias Sep 26, 2026
5bfa9b0
#70: Remove obsolete schema validator
tkilias Sep 26, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions doc/changes/unreleased.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ n/a

## Features / Enhancements

- #70: Added the v2 UDF protocol design and validation artifacts
- #61: Added stacktrace-aware exceptions and assertions to v2
- #32: Added clang tidy to v2
- #36: Added developer guide for clang-tidy and clang-format
Expand All @@ -19,6 +20,7 @@ n/a

## Internal

* #70: Removed the standalone JSON schema validation script and its dedicated CI workflow
* #60: Added Mull mutation testing workflow and report-viewing documentation
for v2; targets without generated mutants now produce warnings instead of
failing the workflow, and separated Linux EventFd code and factory-based
Expand Down
17 changes: 17 additions & 0 deletions doc/design/v2/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# UDF Protocol v2 Design Documents

This directory contains the split protocol design for the new UDF protocol.

- [protocol/design_draft.md](protocol/design_draft.md) is the umbrella design draft.
- [protocol/low_level/protocol.md](protocol/low_level/protocol.md) describes the wire-level rules, generic call lifecycle, and control
stream.
- [protocol/high_level/calls.md](protocol/high_level/calls.md) describes `Run`, Function operations,
`get_connection`, `get_script`, and DB/UDFRunner scheduling policy.
- [protocol/high_level/payloads.md](protocol/high_level/payloads.md) defines their named string and JSON
payload contracts.
- [protocol/high_level/type_mapping.md](protocol/high_level/type_mapping.md) is the high-level Exasol-to-Arrow
column conversion contract, including physical types, parameters, and extension metadata.

Mermaid sources and rendered SVGs use matching names and scopes so the textual and visual material stays aligned.
The low-level schema defines reusable Arrow-compatible physical type capabilities. Exasol type selection and
logical/extension metadata are defined by `high_level/type_mapping.md`.
248 changes: 248 additions & 0 deletions doc/design/v2/protocol/design_draft.md

Large diffs are not rendered by default.

15 changes: 15 additions & 0 deletions doc/design/v2/protocol/high_level/call_model.mmd
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
flowchart TD
Call["Call\n(CallOpen / first peer call-scoped message / CallClose)"]
Call -->|opened by DB, or by UDFRunner only when nested| Opener{{"DB opens top-level calls"}}
Call -->|optional, at most one| DataStream["Bidirectional Data Stream\n(independent schema per direction)"]
Call -->|optional| Nested["Nested Call(s)\n(started by nested CallOpen traffic)"]

Run["Run (DB-opened)"] -.->|instance of| Call
Function["Function operation (DB-opened)"] -.->|instance of| Call
Cleanup["cleanup (DB-opened)"] -.->|instance of| Call
GetConnection["get_connection (UDFRunner-opened)"] -.->|instance of| Call
GetScript["get_script (UDFRunner-opened)"] -.->|instance of| Call

Run --> RunData["Input/Output Data Stream"]
Function --> FunctionKinds["default output columns / virtual schema / import SQL / export SQL"]
Cleanup --> CleanupSemantics["No payload or data stream; release retained resources"]
1 change: 1 addition & 0 deletions doc/design/v2/protocol/high_level/call_model.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
149 changes: 149 additions & 0 deletions doc/design/v2/protocol/high_level/calls.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
# UDF Protocol v2: High-Level Calls

This document describes the high-level protocol calls built on top of the generic call and data-stream mechanisms.

## Scope

This document covers:

- `Run` and Function operations
- `cleanup`
- `get_connection` and `get_script`
- which calls carry data streams
- representative message sequences
- DB/UDFRunner scheduling policy

Related diagrams:

- [call_model.svg](call_model.svg)
- [nested_calls.svg](nested_calls.svg)
- [run_sequence.svg](run_sequence.svg)
- [endpoint_scheduling.svg](endpoint_scheduling.svg)

## Call Families

### `DB`-opened calls
Comment thread
tkilias marked this conversation as resolved.

| Call | Data stream | Notes |
| --- | --- | --- |
| `Run` | Yes, bidirectional | Each direction carries group and row correlation in the data itself. |
| Function operation | No | One of `default_output_columns`, `virtual_schema_adapter`, `generate_sql_for_import_spec`, or `generate_sql_for_export_spec`. |
| `cleanup` | No | DB-opened between calls to let UDFRunner release resources retained from completed calls. |

### Nested `UDFRunner`-opened calls

| Call | Data stream | Notes |
| --- | --- | --- |
| `get_connection` | No | Returns connection information. |
| `get_script` | No | Returns script content. |

These `UDFRunner`-opened calls are ordinary nested calls, not a separate callback transport. `UDFRunner` does not open
top-level calls while idle; it can open them only while handling an active DB call.

## Call-Specific Result Payloads

Each high-level call defines the names and bodies of its own result payloads, carried by `Payloads(...)`. A result
payload may be sent while the call remains active or together with `CloseCall` in the same composite
`StreamMessage`. `CloseCall` is unilateral: the sender does not send further messages on the stream, and the receiver
does not acknowledge it or send further messages on that stream. Messages already in transit may arrive afterward and
are ignored. Calls that have no result payload may still close normally.

## Typical Semantics

### `Run`

- opened by `DB`
- uses `call_metadata` and `column_metadata`, sent before any call, between calls, or with the opening message
- may carry the first input batch together with the opening message
- may stay active while nested calls such as `get_script` or `get_connection` execute
- group and row correlation belong in the data, not in `Next(...)`

See [Group and Row Correlation](correlation.md).

### Function Operations

- opened by `DB` with one of the Function operation names
- each call uses `call_metadata` and `column_metadata`, sent before any call, between calls, or with the opening message
- has no attached data stream in the current model
- has the operation-specific request and result payloads defined in
[payloads.md](payloads.md)

### `cleanup`

- opened by `DB` between calls, after one or more preceding calls have completed
- has no attached data stream, request payload, or normal result payload
- gives `UDFRunner` an explicit opportunity to release resources retained from completed calls or nested calls
- must complete cleanup before sending the normal `CloseCall`; `CloseCall` with `Error` reports cleanup failure
- does not close the connection or prevent later calls after a normal close
- may be repeated periodically by `DB`

The v1 implementation contains an `MT_CLEANUP` message type, but uses it as a cleanup/termination response signal
while processing older request types rather than as a standalone DB-issued operation. The v2 `cleanup` call keeps the
name for compatibility while defining an explicit between-call lifecycle.

### `get_connection`

- opened by `UDFRunner`
- returns `Payloads(connection_information)`
- may still carry additional named payload traffic while active

### `get_script`

- opened by `UDFRunner`
- returns `Payloads(script)`
- may still carry additional named payload traffic while active

## Payload Contracts

The complete call-metadata, script-metadata, Function, and nested-call payload contracts are defined in
Comment thread
tkilias marked this conversation as resolved.
[payloads.md](payloads.md). `StringPayload` is used directly for
scalar strings; JSON is used only where a payload has structured fields.

## Representative Sequences

The current design keeps the high-level sequences intentionally simple:

- nested callback-style calls execute while a parent `Run` or Function call remains active
- `Run` combines `OpenCall`, `call_metadata`, input schema announcement, and the first input batch when practical
- `cleanup` is sequenced between calls: the previous call closes, `DB` opens `cleanup`, and the next call starts only
after the cleanup call closes
- normal completion of the call's data stream is marked by unilateral `CloseCall`; late in-flight stream messages are ignored

See [nested_calls.svg](nested_calls.svg) and
[run_sequence.svg](run_sequence.svg).

## Scheduling Policy

The source material implies the following DB/UDFRunner deadlock-avoidance rules. These rules govern high-level
call orchestration and do not alter the generic Client/Server stream rules in the low-level protocol.

### `UDFRunner`

1. run socket handling and user-code execution as independently wakeable activities
2. wait for either DB socket activity or user-code activity; do not block solely on socket receive
3. queue record batches produced by user code until the first-batch rule or a usable `Next(...)` byte budget permits emission
4. use `Next(...)` byte budgets to bound data in flight; do not impose a message-count limit
5. send regular `KeepAlive` messages so `DB` can continue housekeeping

### `DB`

1. prioritize nested-call responses before data-stream work
2. receive and process record batches from `UDFRunner` as a distinct data-stream event; the first batch may accompany
call-opening/control traffic
3. queue DB-produced input batches until the first-batch rule or a usable `Next(...)` byte budget permits emission
4. if nothing is ready to send, block waiting for new incoming messages
5. monitor peer liveness and terminate unhealthy sessions when needed
6. optionally open `cleanup` between calls and wait for its `CloseCall` before starting the next call

See [endpoint_scheduling.svg](endpoint_scheduling.svg).

## Forward-Looking Ideas Still Open
<>
- `ExecuteScript` and `execute_query` call shapes and data streams
- whether `UDFRunner` may open its own pquery-style call to `DB`
Comment thread
tkilias marked this conversation as resolved.
- whether table-prefetch-like declarations should be added for future call setup
Comment thread
tkilias marked this conversation as resolved.

## Relationship To Other Docs

- low-level protocol lives in [../low_level/protocol.md](../low_level/protocol.md)
- high-level payload contracts live in [payloads.md](payloads.md)
149 changes: 149 additions & 0 deletions doc/design/v2/protocol/high_level/correlation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
# UDF Protocol v2: Group and Row Correlation

This document defines how group and row correlation columns are represented in `Run` record batches.

## Correlation Columns

Each `Run` direction may combine multiple logical groups in one record batch. Its `DataSchema` sets
`has_group_id` and `has_row_id` according to the correlation columns present. When both are set, they form an
ordered reserved prefix before user data columns:

| Position | Column | Purpose |
| --- | --- | --- |
| `0` | Group ID | Identifies the logical input group. |
| `1` | Row ID | Identifies the input row to which an output row maps. |
| `2+` | User data | Input or output columns defined by the call. |

The supported layouts are:

```text
DataSchema {
has_group_id: true,
has_row_id: true,
fields: [group_id: uint64, row_id: uint64, group_id: utf8, value: utf8]
}

DataRecordBatchMetadata {
length: 3,
is_end_of_group: true,
nodes: ...,
buffers: ...
}

record batch {
group_id: [7, 7, 8],
row_id: [1, 2, 1],
group_id: ["a", "b", "c"],
value: ["left", "right", "only"]
}
```

The third column is user data and deliberately has the same name as the reserved group ID column. Correlation
columns are identified by their positions—the first column is the group ID and the second is the row ID—not by
field names. `DataRecordBatchMetadata` contains only batch metadata; the column buffer bytes are transported
separately according to `buffer_transport`.

With only a row ID, the row ID is the first column:

```text
DataSchema {
has_group_id: false,
has_row_id: true,
fields: [row_id: uint64, value: utf8]
}

record batch {
row_id: [1, 2],
value: ["left", "right"]
}
```

A batch without either correlation column has no reserved prefix:

```text
DataSchema {
has_group_id: false,
has_row_id: false,
fields: [value: utf8]
}

record batch {
value: ["left", "right"]
}
```

## Group and Row ID Semantics

Groups may span multiple rows. Group IDs and row IDs are independent opaque identifiers: row IDs may restart at any
value when the group changes or continue across groups. Equal or consecutive row-ID values in different groups do
not imply any relationship between those groups. The meaningful association is the `(group_id, row_id)` pair at one
record position.

In the `UDFRunner` to `DB` direction, both the group ID and row ID are references to the inbound `(group_id, row_id)`
pair from `DB` to `UDFRunner`; they are not newly allocated output identifiers. This applies to both `RETURNS` and
`EMITS` UDFs.

`DataRecordBatchMetadata.is_end_of_group` marks whether the group identified by the batch's final row is complete. It
is defined only for a `Run` direction whose schema sets `has_group_id` to `true`:

- `true` means no later batch in that direction contains the trailing group;
- `false` means the trailing group continues in a later batch;
- a change in group ID still delimits each non-trailing group within the same batch; and
- an empty batch does not complete a group.

This is a group-boundary marker, not an end-of-stream marker. Generic stream-completion semantics remain a low-level
open question and are not encoded in `DataRecordBatchMetadata`.

## Record Batch Mappings

For `DB` to `UDFRunner`, group IDs are usually repeated because a group contains multiple rows, while row IDs are
usually ordered and often consecutive. The following input contains two groups in one record batch. The groups may
have been produced as separate group-sized segments, for example by different threads, and then fused into this
batch:

```text
record batch: group_id row_id
10 101
10 102
20 103
20 104

group_id: [10, 10, 20, 20] -> RunEndEncoded(run_ends=[2, 4], values=[10, 20])
row_id: [101, 102, 103, 104] -> range_run(run_ends=[2, 4], values=[101, 103])
```

The `range_run` run ends align with the group ends, even though the row IDs continue from 102 to 103. The
consecutive values enable compression within each group-sized run, but row IDs 102 and 103 have no semantic
relationship merely because they are adjacent or belong to different groups. Fusion does not change the correlation
values.

For `UDFRunner` to `DB`, the output references the inbound rows by the same pair:

```text
inbound rows: (group_id, row_id) = [(10, 101), (10, 102), (20, 103)]

RETURNS output (one output per input row):
group_id: [10, 10, 20]
row_id: [101, 102, 103] -> range_run(run_ends=[2, 3], values=[101, 103])

EMITS output (multiple outputs for input row (10, 101)):
group_id: [10, 10, 10, 20]
row_id: [101, 101, 101, 103] -> RunEndEncoded(run_ends=[3, 4], values=[101, 103])
```

For `RETURNS`, `UDFRunner` usually produces exactly one output row per input row or group, so row-ID references are
usually ordered and group IDs are usually repeated. For `EMITS`, a group usually contains multiple output rows and
multiple output rows usually reference the same input row, so both group IDs and row-ID references are usually
repeated.

## Preferred Encodings

| Column and direction | Preferred encoding | Compatible fallback |
| --- | --- | --- |
| Group ID, either direction | [`RunEndEncoded`](https://arrow.apache.org/docs/format/Columnar.html#run-end-encoded-layout) over unsigned 64-bit IDs when groups contain repeated rows. | Plain unsigned 64-bit. |
| Row ID, `DB` to `UDFRunner` | [`exasol.udf.range_run`](range_run.md) extension array. | Plain unsigned 64-bit. |
| Row ID, `UDFRunner` to `DB`, `RETURNS` UDF | [`exasol.udf.range_run`](range_run.md) extension array. | Plain unsigned 64-bit. |
| Row ID, `UDFRunner` to `DB`, `EMITS` UDF | [`RunEndEncoded`](https://arrow.apache.org/docs/format/Columnar.html#run-end-encoded-layout) over unsigned 64-bit IDs. | Plain unsigned 64-bit. |

Plain unsigned 64-bit encoding remains the compatible fallback when the preferred encoding is unavailable or not
beneficial. The detailed `exasol.udf.range_run` storage encoding is specified in [range_run.md](range_run.md).
25 changes: 25 additions & 0 deletions doc/design/v2/protocol/high_level/endpoint_scheduling.mmd
Comment thread
tkilias marked this conversation as resolved.
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
stateDiagram-v2
state "UDFRunner" as UDF {
[*] --> UdfIdle
UdfIdle --> UdfInitializing: Accept
UdfInitializing --> UdfActive: Capabilities
UdfActive --> UdfWaitingForActivity: Wait for DB or user code
UdfWaitingForActivity --> UdfOutputReady: User-code output ready
UdfOutputReady --> UdfSendingData: First batch or Next credit
UdfSendingData --> UdfActive: Transfer done
UdfWaitingForActivity --> UdfActive: Socket or non-output user-code event
}

state "DB" as DB {
[*] --> DbIdle
DbIdle --> DbInitializing: Connect
DbInitializing --> DbActive: Capabilities
DbActive --> DbHandlingNestedCall: Call
DbHandlingNestedCall --> DbActive: Reply
DbActive --> DbWaitingForUdfMessage: Wait
DbWaitingForUdfMessage --> DbActive: Control message
DbWaitingForUdfMessage --> DbReceivingData: RecordBatch received
DbReceivingData --> DbActive: RecordBatch processed
DbActive --> DbSendingData: Data available
DbSendingData --> DbActive: Transfer done
}
1 change: 1 addition & 0 deletions doc/design/v2/protocol/high_level/endpoint_scheduling.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
14 changes: 14 additions & 0 deletions doc/design/v2/protocol/high_level/examples/call_metadata.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"database_name": "EXASOL",
"database_version": "8.0",
"session_id": "42",
"statement_id": 1,
"node_count": 1,
"node_id": 0,
"vm_id": "7",
"maximal_memory_limit": "1073741824",
"script_schema": "SYS",
"input_iter_type": "EXACTLY_ONCE",
"output_iter_type": "EXACTLY_ONCE",
"single_call_mode": false
}
Loading
Loading