Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
129 changes: 127 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,13 @@ on:
push:
branches: [ main, master ]
pull_request:
workflow_dispatch:
inputs:
run_gpu:
description: Run GPU conformance on the self-hosted opencl-gpu runner
type: boolean
required: false
default: false

jobs:
rust:
Expand Down Expand Up @@ -56,6 +63,9 @@ jobs:
done < <(find "$dir" -name Cargo.toml -type f)
done

- name: Install numerical test dependencies
run: sudo apt-get update && sudo apt-get install -y m4

- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
Expand Down Expand Up @@ -87,5 +97,120 @@ jobs:
cargo test
fi

# Note: OpenCL-related tests are not run in CI by default.
# They require GPU drivers and an OpenCL ICD on the runner.
- name: Native lint
run: cargo clippy --all-targets --features "${{ matrix.features }}" -- -D warnings

- name: Release numerical conformance
run: cargo test --release --features complex --test numerics --target-dir target/host-release

opencl-cpu:
name: OpenCL CPU numerical conformance
runs-on: ubuntu-latest
env:
CC: /usr/bin/cc
CXX: /usr/bin/c++
M4: /usr/bin/m4
CARGO_TARGET_X86_64_UNKNOWN_LINUX_GNU_LINKER: /usr/bin/cc
HA_NDARRAY_OPENCL_DEVICE: CPU
POCL_CACHE_DIR: /tmp/ha-ndarray-pocl
steps:
- uses: actions/checkout@v4
- name: Bootstrap local path dependencies
shell: bash
run: |
set -euo pipefail

declare -A seen
queue=("$PWD")
seen["$PWD"]=1

while ((${#queue[@]})); do
dir="${queue[0]}"
queue=("${queue[@]:1}")

while IFS= read -r cargo; do
while IFS= read -r dep; do
[[ -n "$dep" ]] || continue
dep_dir="$(cd "$dir/.." && pwd)/$dep"

if [[ ! -d "$dep_dir" ]]; then
echo "Cloning dependency repo: $dep"
git clone --depth 1 "https://github.com/TinyChain-Inc/$dep.git" "$dep_dir"
fi

if [[ -z "${seen[$dep_dir]+x}" ]]; then
queue+=("$dep_dir")
seen["$dep_dir"]=1
fi
done < <(
grep -hoE 'path[[:space:]]*=[[:space:]]*"\.\./[^"]+"' "$cargo" \
| sed -E 's/.*"\.\.\/([^"]+)".*/\1/' \
| sort -u || true
)
done < <(find "$dir" -name Cargo.toml -type f)
done

- name: Install CPU OpenCL and oracle build dependencies
run: sudo apt-get update && sudo apt-get install -y pocl-opencl-icd ocl-icd-opencl-dev clinfo build-essential m4
- uses: dtolnay/rust-toolchain@stable
with:
components: clippy, rustfmt
- uses: Swatinem/rust-cache@v2
- run: clinfo -l
- run: cargo test --all-targets --features opencl,complex --target-dir target/opencl-tests
- run: cargo test --release --features opencl,complex --test numerics --target-dir target/opencl-tests
- run: cargo test --doc --all-features --target-dir target/opencl-tests
- run: cargo clippy --all-targets --all-features --target-dir target/opencl-tests -- -D warnings

opencl-gpu:
name: OpenCL GPU numerical conformance
if: github.event_name == 'workflow_dispatch' && inputs.run_gpu
runs-on: [self-hosted, linux, x64, opencl-gpu]
env:
CC: /usr/bin/cc
CXX: /usr/bin/c++
M4: /usr/bin/m4
CARGO_TARGET_X86_64_UNKNOWN_LINUX_GNU_LINKER: /usr/bin/cc
HA_NDARRAY_OPENCL_DEVICE: GPU
steps:
- uses: actions/checkout@v4
- name: Bootstrap local path dependencies
shell: bash
run: |
set -euo pipefail

declare -A seen
queue=("$PWD")
seen["$PWD"]=1

while ((${#queue[@]})); do
dir="${queue[0]}"
queue=("${queue[@]:1}")

while IFS= read -r cargo; do
while IFS= read -r dep; do
[[ -n "$dep" ]] || continue
dep_dir="$(cd "$dir/.." && pwd)/$dep"

if [[ ! -d "$dep_dir" ]]; then
echo "Cloning dependency repo: $dep"
git clone --depth 1 "https://github.com/TinyChain-Inc/$dep.git" "$dep_dir"
fi

if [[ -z "${seen[$dep_dir]+x}" ]]; then
queue+=("$dep_dir")
seen["$dep_dir"]=1
fi
done < <(
grep -hoE 'path[[:space:]]*=[[:space:]]*"\.\./[^"]+"' "$cargo" \
| sed -E 's/.*"\.\.\/([^"]+)".*/\1/' \
| sort -u || true
)
done < <(find "$dir" -name Cargo.toml -type f)
done

- uses: dtolnay/rust-toolchain@stable
- name: Require installed OpenCL runtime and oracle tools
run: clinfo -l && m4 --version
- run: cargo test --all-targets --features opencl,complex --target-dir target/opencl-gpu-tests
- run: cargo test --release --features opencl,complex --test numerics --target-dir target/opencl-gpu-tests
19 changes: 15 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,15 +27,26 @@
device capacity before allocating buffers or enqueueing commands, keep
transfers bounded, and propagate saturation to the caller; do not hide it with
unbounded host queues or an implicit CPU/GPU fallback.
- Device class selection is fixed before admission. If the selected class is
absent, fail unless bootstrap explicitly configured a fallback; never search
other classes opportunistically, and never reroute because a device is busy.
- Select the platform automatically based on workload size within user-configured
constraints. Respect the permitted platform type, enabled backends, and configured
OpenCL device class. An array's incoming platform is not a permanent execution pin.
- At a scheduling boundary, use one selected platform for operation construction,
prerequisite transforms, and the returned array. Axis reductions select using the
input element count, which represents the work to consume, not the output size.
- Choose the OpenCL device class before admission within the configured constraints.
If it is absent, fail unless bootstrap explicitly configured a fallback; never
broaden the allowed classes opportunistically or reroute because a device is busy.
Workload-based selection is normal scheduling, not recovery from a capability,
compilation, execution, or capacity failure.
- Run `cargo fmt` and `cargo clippy` before pushing.

## Testing Guidelines
- Framework: Rust `#[test]` with `cargo test`; integration tests live in `tests/*.rs`.
- Name tests descriptively (e.g., `#[test] fn transpose_concat_validates_dims()`), assert both values and shapes.
- For GPU-specific logic, guard with feature flags and provide host fallbacks when possible.
- Guard GPU-specific tests with feature flags. Use explicit backend adapters for
kernel conformance and separate tests for automatic workload-based selection;
do not change production scheduling to pin conformance tests to a backend.
A required but unavailable device must fail its test job, not fall back to host.
- Keep tests deterministic and fast; seed randomness when used.

## Commit & Pull Request Guidelines
Expand Down
5 changes: 4 additions & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ crate-type = ["cdylib", "rlib"]
all = ["complex", "freqfs", "opencl", "stream"]
complex = ["num-complex", "rustfft"]
debug_crash = [] # enable this to panic rather than retuning an error when debugging
freqfs = ["freqfs/stream", "stream"]
freqfs = ["dep:freqfs", "stream"]
opencl = ["memoize", "ocl"]
stream = ["destream", "futures"]
wasm = ["wasm-bindgen"]
Expand All @@ -41,3 +41,6 @@ safecast = "0.2"
smallvec = "1.13"
transpose = "0.2"
wasm-bindgen = { version = "0.2", optional = true }

[dev-dependencies]
rug = { version = "1", default-features = false, features = ["float", "complex", "integer", "rational"] }
189 changes: 189 additions & 0 deletions NUMERICS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,189 @@
# Numerical behavior

ha-ndarray owns numerical policy for callers and execution backends. The host
and OpenCL implementations obey the same scalar rules. Storage codecs do not
change this contract.

## Supported operations

| Family | Integers | f32/f64 | Complex32/Complex64 |
| --- | --- | --- | --- |
| Storage, geometry, copying, conditional selection | Yes | Yes | Yes |
| Add/subtract/multiply/divide/power, including scalar operands | Yes | Yes | Yes |
| Remainder, ordering, min/max | Yes | Yes | No |
| Array rounding | No | Yes | No |
| Components, argument, conjugation, conjugate transpose | No | No | Yes |
| Absolute value | Same dtype | Same dtype | Corresponding real dtype |
| exp/ln/log and nine trigonometric functions | No | Yes | Yes |
| Equality, boolean operations | u8 output | u8 output | u8 output |
| is_nan/is_inf | No public numeric predicate trait | u8 output | u8 output |
| Cast | All supported destinations | All supported destinations | All supported destinations |
| Sum/product, matrix products, diagonal | Yes | Yes | Yes |
| FFT/IFFT | No | No | Host only |
| Uniform/normal random constructors | No | f32 output only | No |

Integers are i8/i16/i32/i64 and u8/u16/u32/u64. Complex numbers require the complex
feature. Unsupported trait combinations do not compile; unavailable execution
capabilities return Error::Unsupported. OpenCL FFT does not silently copy to a
host implementation: select host execution explicitly.

Transforms and copies preserve payload values. Conditional selection uses nonzero
u8 conditions; unselected branches are not guaranteed to avoid evaluation or
errors. Shapes/dtypes must match where required by the existing API; broadcasting
and casting remain explicit.

## Scalar rules

Integer addition, subtraction, multiplication, absolute value, and nonnegative
powers wrap modulo 2^width. Signed results reinterpret these bits as two's
complement. The absolute value of signed MIN remains MIN. Division truncates
toward zero. Integer division/remainder by zero return zero. MIN/-1 wraps to
MIN; MIN%-1 is zero. Nonzero remainder has the dividend's sign.

Negative integer powers return 1 for base 1, the parity result for base -1, and
zero otherwise, including base zero. 0^0 is 1. Exponents retain their full width,
without floating conversion. Reductions, matrix accumulation, and range
arithmetic use these same rules.

Real floats use nearest-even basic arithmetic and gradual underflow. Preserve
subnormal inputs/outputs, signed zeros, and infinities. Native execution threads
must retain the standard floating-point environment. Overflow and invalid
arithmetic produce IEEE values rather than domain errors or clamping. Nonzero
division by signed zero yields signed infinity; 0/0 yields NaN. NaN payload/sign
are not portable.

Remainder follows Rust % / OpenCL fmod, not Euclidean or IEEE remainder.
The scalar Real::round integer helper is the identity; array rounding remains
float-only. Finite x % infinity is x; infinity % y and x % 0 are NaN. Round uses nearest
integer with ties away from zero and preserves zero sign; abs clears the sign.

ln(±0) is negative infinity; ln of negative real inputs is NaN. Log(x,b) means
ln(x)/ln(b), including exceptional results. Power follows real powf semantics:
x^±0 is 1 even for NaN x, and 1^y is 1 even for NaN y. Negative finite bases with
nonintegral finite exponents produce NaN. Signed-zero/infinity results depend
on exponent sign and odd-integer parity. Other NaN inputs propagate NaN.

Comparisons use IEEE unordered NaN semantics. Min/max propagate any NaN;
among equal signed zeros min selects -0 and max selects +0. Reduction identities
must not replace valid infinities with finite bounds.

Complex operations follow num-complex definitions, exceptional cases, and
principal branches. Branch-cut sides depend on imaginary signed zero. Complex
power with zero exponent is 1+0i. Complex logarithm uses log of the magnitude and
atan2 of the components; complex-base log divides these logarithms.

A number is false exactly when it equals zero. Complex zero requires both
components to be zero. NaN is truthy. Boolean operations and predicates return
exactly u8 0 or 1. Complex is_nan/is_inf inspect either component independently.

## Cast compatibility

Casts preserve number-general's CastFrom<Number> pipeline, not direct C casts:

- f32/Complex32 first convert the real component to i32/u32 for integer
destinations; f64/Complex64 use i64/u64. This conversion truncates, saturates
at the intermediate width, and maps NaN to zero, before narrowing.
- Signed-to-unsigned conversion first reinterprets at the source width.
- Unsigned-to-signed conversion first uses the corresponding signed width,
except u8 first becomes i16.
- Integers through 32 bits first become f32 for float/complex destinations;
64-bit integers first become f64.
- Complex-to-real takes the real component; real-to-complex supplies zero
imaginary component. Complex-to-complex converts both components.

Thus -1i8 cast to u64 is 255, u32::MAX cast to i64 is -1, and 16_777_217i32 cast
to f64 is 16_777_216. Regression fixtures pin these compatibility rules.

## Accuracy and aggregates

Basic real arithmetic and individual conversions specified above are correctly
rounded. Finite real transcendental results have an 8-ULP bound. Finite complex
results use componentwise absolute error <=32u*max(1,abs(reference)), where u is
half machine epsilon. Check classification, signed zero, and branches separately.

Elementwise nodes retain separate rounding. Aggregates may reassociate or use
contraction, without bitwise reproducibility. For finite intermediates without
overflow/underflow, use gamma(k)=ku/(1-ku), k=8N for real reductions/dot products
and k=32N for complex equivalents/FFTs. Scale by sums of absolute terms for
sums/dots, the absolute exact product for products, and the input absolute sum
for FFTs. These bounds do not apply for ku>=1 or intermediate range violations.
Extreme-input aggregate classifications can differ with evaluation order.

FFT/IFFT operate independently on each last-axis batch and are unnormalized:
ifft(fft(x)) approximates N*x. Fourier point reads remain unsupported. This
upgrade does not introduce empty-shape support to operations lacking it.

Uniform random constructors return f32 values in [0,1); normal constructors
sample the standard normal distribution with finite Box-Muller inputs. Sequences
need not match across backends or independent evaluations. Range constructors
retain their existing step calculation and apply the scalar rules above.

## Backend requirements and compatibility

Execution platforms are selected automatically from workload size within the
platform type, enabled backends, and user-configured device constraints. Select
once at a scheduling boundary and use that platform consistently for prerequisite
transforms, the operation, and its returned array. Axis reductions use the input
element count. Automatic scheduling does not permit fallback after a numerical
capability or execution failure. Conformance adapters explicitly select the backend
under test independently of the public array scheduler.

OpenCL checks numerical capabilities on the selected device: f64 support,
subnormals, infinities/NaNs, nearest rounding, and correctly rounded f32 division.
Unsupported paths report operation, dtype, device, and required capability.
Compilation is device-specific. Relaxed-math/flush-to-zero options are disabled,
and elementwise contraction is disabled.

Compatibility changes include wide integer powers and zero remainders,
scalar division by zero following the same dtype rules as array division,
NaN-propagating min/max, complex predicates, OpenCL casts/unary operations,
and inverse FFT dispatch. The previously unreadable complex re()/im() outputs
now correctly declare the corresponding real dtype. Other ordinary signatures
and persistent formats are unchanged.

## Conformance references and validation

The shared suite uses explicitly selected host/OpenCL buffers. Development-only
MPFR/MPC directed rounding encloses individual results; composed logarithm and
DFT references propagate bounds through every intermediate operation. Reference
precision starts at 256 bits and doubles through 4096. A correctly rounded
reference is accepted only when both endpoints round directly to the same f32
or f64 value. Unresolved references fail with operation/input diagnostics.
NaNs, infinities, signed zeros, and branch conventions are checked separately.
Integer expectations and finite algebraic sums, products, and dot products use
exact Integer/Rational arithmetic. Number-general cast fixtures remain separate.

Aggregate checks cover f32/f64 and both complex widths, reduction lengths
1, 7, 8, 9, 63, 64, 65, and 129, transformed inputs, multiple axes, and both
keepdims settings. Batched matrix cases include tile boundaries and padding.
Forward and inverse FFTs are checked independently for lengths 1, 3, 8, and 17,
with three distinct complex batches. Each batch has its own error bound;
round-trip bounds include propagated forward error plus inverse error. Reference
rounding never enlarges the specified aggregate tolerances.

Synthetic capability tests cover missing flags, failed queries, complex component
precision, and cast input/output/intermediate precision. These supplement actual
device execution; production still validates the selected queue device before
compilation. Mandatory CPU-OpenCL CI and the manually dispatched `opencl-gpu`
runner execute the same cases. Missing required hardware fails the selected job.

Implementation and native/CPU-OpenCL validation are separate from GPU approval.
Actual GPU conformance remains a mandatory pending gate: CPU OpenCL results do
not establish GPU conformance. A future CubeCL implementation must satisfy the
same contract and shared cases before being advertised as supported.

Validation before the scheduling correction below used PoCL 6.0+debian (OpenCL 3.0, LLVM 18.1.8)
on `cpu-skylake-avx512-AMD Ryzen 7 7840HS w/ Radeon 780M Graphics`, explicitly
selected as a CPU device. The device name does not indicate GPU execution.

| Validation | Debug | Release |
|---|---:|---:|
| Native host, complex enabled, all targets | 57 passed | 57 passed |
| CPU-OpenCL, complex enabled, all targets | 91 passed | 91 passed |

All-feature compilation, Clippy with warnings denied, doctests (no examples),
formatting, and diff checks passed. Actual GPU execution has not been validated.

The axis-reduction scheduler subsequently restored automatic workload-based
selection. Backend conformance also checks explicit reduction adapters so small
OpenCL cases remain device-executed even when the general scheduler chooses host.
Loading
Loading