diff --git a/AGENTS.md b/AGENTS.md index ddbd418..f8c471b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,9 +35,9 @@ removals. A JOSS paper + citation remain a post-2.0 follow-up. | --- | --- | | `config.py` | `CameraConfig` (frozen dataclass of detector params) and `SensorType` enum. Pure data + validation + temperature scaling. No randomness. | | `noise.py` | The physics. Pure functions: `CameraConfig` + exposure + temperature + seeded `Generator` → electrons/ADU. This is where noise models and the opt-in reusable `DetectorWorkspace` live. | -| `backend.py` | Optional NumPy/CuPy array and RNG boundary, explicit host conversion, and backend convolution. NumPy is the reference/default. Kept local rather than `aocore.Backend`: getframes draws per-frame noise from a device-native (cuRAND) stream, which `aocore.Backend.random` (host-side `Generator`) does not provide. | +| `backend.py` | Optional NumPy/CuPy array and RNG boundary, explicit host conversion, and backend convolution. NumPy is the reference/default. Owns the stack's `device` vocabulary (`"cpu"`, `"gpu"`, `"gpu:N"`, `"auto"`; parsed by `_parse_device`, resolved by `get_backend`) and `precision` vocabulary (`resolve_precision`: `"single"`/`"double"`, aliases `"float32"`/`"float64"`). A GPU `ArrayBackend` carries a `device_id`; `activate()` makes it current. Kept local rather than `aocore.Backend`: getframes draws per-frame noise from a device-native (cuRAND) stream, which `aocore.Backend.random` (host-side `Generator`) does not provide. | | `frame.py` | `Frame` container: a NumPy array (ADU) plus metadata; array-like; optional FITS export. | -| `camera.py` | `Camera`, the main user-facing object. Orchestrates config + scene + noise into `Frame`s. Holds the RNG and high-level methods (`dark_frame`, `dark_series`, reset-correlated `nondestructive_series`, `correlated_double_sample`, `expose`, `observe`, `*_series`, `master_*`). Reset-correlated readout has one core, the private `_ramp_reads`, which walks a ramp on an arbitrary per-read interval pattern; `nondestructive_series` drives it uniformly and `correlated_double_sample` drives it as pedestal-then-signal. Put new ramp-readout modes there rather than in a second loop — the interval-scaled bias/settling/avalanche terms are easy to get subtly wrong twice. | +| `camera.py` | `Camera`, the main user-facing object. Orchestrates config + scene + noise into `Frame`s. Holds the RNG and high-level methods (`dark_frame`, `dark_series`, reset-correlated `nondestructive_series`, `correlated_double_sample`, `expose`, `observe`, `*_series`, `master_*`). Reset-correlated readout has one core, the private `_ramp_reads`, which walks a ramp on an arbitrary per-read interval pattern; `nondestructive_series` drives it uniformly and `correlated_double_sample` drives it as pedestal-then-signal. Put new ramp-readout modes there rather than in a second loop — the interval-scaled bias/settling/avalanche terms are easy to get subtly wrong twice. Every public method that touches device arrays is decorated with `backend._on_device`, which runs it (each step, for generator methods) with the camera's CUDA device current; decorate new ones too, or `device="gpu:N"` silently allocates on the wrong card. | | `calibrate.py` | Master-frame builders (`combine`) and `calibrate` reduction — the raw → reduced → truth loop (phase 1.1). | | `observation.py` | `Observation` / `ObservationTruth` / `Pointing`: the time-series driver, jitter/drift/dither, per-frame truth (phase 1.2). | | `spectral.py` | Opt-in spectral mode: `QE`, `SED` (relative or absolute via `from_flux_density`), `Spectrum`, `SpectralBandpass`, effective-QE folding, transmission-product helpers (`product`, `from_file`/`from_product`), optional `astropy.units` coercion. | @@ -109,9 +109,18 @@ indices and takes a background/threshold/window — a different contract from offset-centred `_radial_grid` in `analysis/apertures.py`, and `backend.py` (see the Architecture table). A bug in an aocore primitive is fixed in aocore, never worked around here. `tests/test_conformance.py` runs the `aocore.conformance` -checks that apply to a detector package (image-plane centring, unit flux); the -OPD-driven ones (tilt, slopes, wind, Zernikes, RMS) do not apply because -getframes PSFs are analytic, not images of an OPD map. +image-builder checks that apply to a detector package, for every PSF model: +`check_point_source_centring` (also on the vignetting map), +`check_point_source_flux`, and `check_edge_flux_loss` (light off a detector edge +is lost, not renormalized). They need `aocore>=0.1.3`, pinned in the `dev` +extra only; the runtime pin stays `>=0.1.2`. The OPD-driven checks (tilt, +slopes, wind, Zernikes, RMS) do not apply because getframes PSFs are analytic, +not images of an OPD map. It also pins the section 8 +software vocabulary: every `device=` takes `"cpu"`/`"gpu"`/`"gpu:N"`/`"auto"` +and every working-precision argument takes `precision="single"|"double"` +(`Camera`, `Scene.photon_rate_map`/`photoelectron_rate_map`, and the `noise` +functions with a `float_dtype`). The dataset `dtype` is a host *storage* type, +not a working precision, so it stays a NumPy dtype. ## Adding a camera preset diff --git a/CHANGELOG.md b/CHANGELOG.md index 4d207b7..b7d5061 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,60 @@ to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [Unreleased] +## [2.4.0] - 2026-10-07 + +### Added + +- **`device="gpu:N"` and `device="auto"`.** Every `device` argument (`Camera`, + `get_backend`, and the CLI's `[camera]` table) now speaks the AO stack's + vocabulary (aocore CONVENTIONS 8.1): `"cpu"`, `"gpu"`, `"gpu:N"` for CUDA + device `N`, and `"auto"`, which picks the GPU when CuPy is installed and sees + a device and the CPU otherwise. A `"gpu:N"` beyond the devices CuPy sees + raises a `ValueError` naming the count. The old spellings (`"numpy"`, + `"cuda"`, `"cupy"`) still work, now case-insensitively. A GPU camera is + pinned to its card: the fixed-pattern maps and the cuRAND streams are created + on it and every camera method runs with it current, so a `"gpu:1"` camera + works whichever device is current at the call, and `with_config` keeps it. + New: `Camera.device_id`, `ArrayBackend.device_id`, `ArrayBackend.spec` + (`"cpu"` or `"gpu:N"`) and `ArrayBackend.activate()` (the device context, + for calling the low-level `noise` functions on another card). +- **`precision="single"` / `"double"`.** The working precision takes the + shared names (aocore CONVENTIONS 8.2), with `"float32"`/`"float64"` kept as + aliases: on `Camera`, as a new `precision` keyword on + `Scene.photon_rate_map`/`photoelectron_rate_map` (beside `dtype`), and on the + `noise` functions that take a `float_dtype` (`simulate_frame`, + `fixed_pattern_maps`, `dark_signal_map`, `photo_signal_map`). A `dtype` and a + `precision` that disagree raise `ValueError`. `getframes.resolve_precision` + maps any of these names to the NumPy dtype. `Camera.precision` still reports + the dtype name (`"float32"`/`"float64"`) whichever spelling was passed. + `dataset.pairs(dtype=...)` is unchanged: it is the host *storage* type of the + finished arrays, not a working precision. +- The CLI's `[camera]` table takes a `device` key. +- **Conformance tests** for the device and precision vocabulary, including that + each precision name selects the same dtype as in aocore. +- **Edge-flux conformance for every PSF model.** `tests/test_conformance.py` + now uses aocore 0.1.3's image-builder checks (`check_point_source_centring`, + `check_point_source_flux`) instead of feeding analytic PSFs through the + OPD-driven checks with a dummy OPD, and adds `check_edge_flux_loss` for + Gaussian, Moffat, elliptical Gaussian, Airy and array PSFs, guarding the + 2.3.0 fix. The `dev` extra pins `aocore>=0.1.3,<0.2`; the runtime + requirement is unchanged. + +### Changed + +- An unknown `device` string now raises `ValueError` listing the accepted words + (`'cpu', 'gpu', 'gpu:N' ... or 'auto'`), and a non-string `device` a + `TypeError`. `device="gpu"` with CuPy installed but no CUDA device raises + `RuntimeError` at construction rather than failing at the first frame. +- `Camera.__repr__` shows the GPU number (`device='gpu:0'`). +- The noise functions' `float_dtype` default is now `None` (still float64). + +### Fixed + +- `dataset.pairs` with a GPU camera failed on an implicit CuPy-to-NumPy + conversion; it now copies each frame to the host through `to_numpy`, as do + the CLI's `.npy`/`.npz` writers now that the CLI can select a GPU. + ## [2.3.0] - 2026-10-07 ### Fixed @@ -592,7 +646,8 @@ together in 1.0. - Documentation, runnable examples, and CI (lint, type-check, test matrix, PyPI release via Trusted Publishing). -[Unreleased]: https://github.com/jacotay7/getframes/compare/2.3.0...HEAD +[Unreleased]: https://github.com/jacotay7/getframes/compare/2.4.0...HEAD +[2.4.0]: https://github.com/jacotay7/getframes/compare/2.3.0...2.4.0 [2.3.0]: https://github.com/jacotay7/getframes/compare/2.2.0...2.3.0 [2.2.0]: https://github.com/jacotay7/getframes/compare/2.1.1...2.2.0 [2.1.1]: https://github.com/jacotay7/getframes/compare/2.1.0...2.1.1 diff --git a/README.md b/README.md index 1a45c4f..7917f73 100644 --- a/README.md +++ b/README.md @@ -60,7 +60,7 @@ frame = cam.with_config(resolution=(256, 256)).observe(scene, exposure=300.0, se import cupy as cp # and the same path on a GPU -cam = gf.Camera.from_preset("andor_ocam2k", device="gpu", precision="float32") +cam = gf.Camera.from_preset("andor_ocam2k", device="gpu", precision="single") rate = cp.full(cam.resolution, 2.0e6, dtype=cp.float32) # photons/s/pixel frame = cam.expose(rate, exposure=1.0e-3, seed=0) # CuPy ADU, no host copy ``` @@ -140,8 +140,8 @@ for the methodology. - **Scale & datasets** — a float32 fast path, vectorised multi-source rendering, a streaming raw+truth `dataset` generator and a `getframes` CLI; see **[Scale & datasets](https://jacotay7.github.io/getframes/guides/datasets/)**. -- **GPU-optional** — every camera takes `device="gpu"` (CuPy) and keeps the - detector path and truth arrays device-resident. CPU and GPU have independent +- **GPU-optional** — every camera takes `device="gpu"`, `"gpu:N"` or `"auto"` + (CuPy) and keeps the detector path and truth arrays device-resident. CPU and GPU have independent RNG streams, so a `seed` repeats exactly on a fixed backend while parity across backends means matching statistics, not identical pixels. - **Reproducible and typed** — all randomness flows through a camera-owned seeded diff --git a/docs/guides/datasets.md b/docs/guides/datasets.md index 2b73251..a67ed72 100644 --- a/docs/guides/datasets.md +++ b/docs/guides/datasets.md @@ -8,15 +8,17 @@ unchanged and remains the default. ## The float32 fast path -Pass `precision="float32"` when you build a `Camera` to run the whole signal chain +Pass `precision="single"` when you build a `Camera` to run the whole signal chain — and each frame's ground truth — in single precision. That halves the memory of the per-pixel arrays, which matters for large detectors and when you are generating -thousands of frames: +thousands of frames. `"single"`/`"double"` is the vocabulary shared across the AO +stack; `"float32"`/`"float64"` are accepted aliases, and `cam.precision` reports +the NumPy dtype name (`"float32"`) whichever spelling you passed: ```python import getframes as gf -cam = gf.Camera.from_preset("zwo_asi2600mm", precision="float32") +cam = gf.Camera.from_preset("zwo_asi2600mm", precision="single") frame = cam.expose(photon_rate=200.0, exposure=30.0, seed=0) frame.dtype # uint32 — the digitised ADU stay exact integers @@ -28,13 +30,20 @@ way. Persistent PRNU/DSNU, amplifier gain/offset, structured-bias, and per-pixel read-noise maps use the selected precision too, so a `float32` camera does not keep hidden double-precision detector-sized coefficients. The result matches the `float64` path statistically and to single-precision tolerance for deterministic -maps. If you call the scene or noise layers directly, the same control is a `dtype` -/ `float_dtype` argument: +maps. If you call the scene or noise layers directly, the same control is a +`precision` keyword, or the older `dtype` / `float_dtype` argument (give one, or +both only when they agree): ```python -rate_map = scene.photon_rate_map(dtype="float32") # f32 photons/s/pixel map +rate_map = scene.photon_rate_map(precision="single") # f32 photons/s/pixel map +rate_map = scene.photon_rate_map(dtype="float32") # the same map ``` +The `dtype` of [`pairs`][getframes.dataset.pairs] is different: it is the +*storage* type the finished `raw`/`truth` arrays are cast to on the host, not a +working precision, so it stays a NumPy dtype (the camera's `precision` sets how +they are computed). + ## Vectorised catalog rendering A [`Catalog`][getframes.scene.sources.Catalog] of many stars no longer loops in @@ -73,7 +82,7 @@ re-iterable source of random fields: ```python import getframes as gf -cam = gf.Camera.from_preset("zwo_asi2600mm", precision="float32") +cam = gf.Camera.from_preset("zwo_asi2600mm", precision="single") scenes = gf.dataset.random_star_fields(n=10_000, shape=cam.resolution, seed=0) ds = gf.dataset.pairs(camera=cam, scenes=scenes, exposure=60.0, dtype="float32", seed=1) @@ -102,7 +111,8 @@ A `generate` config names a preset (or an inline camera) and a frame spec: [camera] preset = "andor_ikon_m934" default_temperature_c = -60.0 -precision = "float32" +precision = "single" # or "double"; "float32"/"float64" also accepted +device = "auto" # cpu | gpu | gpu:N | auto (GPU when CuPy sees one) [frame] type = "dark" # dark | bias | flat | light @@ -117,7 +127,7 @@ requested `shape`: ```toml [camera] preset = "zwo_asi2600mm" -precision = "float32" +precision = "single" [dataset] n = 1000 diff --git a/docs/guides/gpu.md b/docs/guides/gpu.md index e78d8c4..a76232b 100644 --- a/docs/guides/gpu.md +++ b/docs/guides/gpu.md @@ -14,7 +14,7 @@ import getframes as gf camera = gf.Camera.from_preset( "andor_ocam2k", device="gpu", - precision="float32", + precision="single", ) photon_rate = cp.full(camera.resolution, 2.0e6, dtype=cp.float32) frame = camera.expose(photon_rate, exposure=1.0e-3, seed=0) @@ -33,6 +33,48 @@ binning, truth, and ADU digitisation. Wavelength-resolved truth on device. Static fixed-pattern maps are constructed in the camera's working precision and cached once, so reuse the same `Camera` in a frame loop. +## Choosing a device + +`device` takes the words shared across the AO stack: + +| `device` | Runs on | +| --- | --- | +| `"cpu"` (default) | NumPy, the reference implementation | +| `"gpu"` | CuPy on the CUDA device current when the camera is built | +| `"gpu:N"` | CuPy on CUDA device `N` | +| `"auto"` | the current CUDA device when CuPy is installed and sees one, else the CPU | + +Matching is case-insensitive, and the older spellings still work: `"numpy"` for +`"cpu"`, `"cuda"`/`"cupy"` for `"gpu"` (also `"cuda:N"`). An unknown word, or a +`"gpu:N"` beyond the devices CuPy sees, raises `ValueError` naming the device +count; `"gpu"` without CuPy raises `ImportError`. `"auto"` never raises for a +missing GPU, so one script runs on a laptop and a CUDA workstation alike: + +```python +camera = gf.Camera.from_preset("andor_ocam2k", device="auto", precision="single") +camera.device # "gpu" or "cpu" +camera.device_id # CUDA device number, or None on the CPU +``` + +A GPU camera is pinned to one card. Its fixed-pattern maps and its cuRAND +streams are created on that device, and every camera method runs with it +current, so `device="gpu:1"` works whichever device is current when you call +it; `Camera.with_config` keeps the same card. The low-level +[`getframes.noise`](../reference.md) functions run on the current device; to use +them on another card, enter the backend's device context: + +```python +backend = gf.get_backend("gpu:1") +with backend.activate(): + ... # noise.simulate_frame(..., backend=backend) +``` + +`precision` is `"single"` (float32) or `"double"` (float64, the default), with +`"float32"`/`"float64"` accepted as aliases; see +[Scale & datasets](datasets.md#the-float32-fast-path). + +## Host transfers + `frame.data` is the zero-copy device interface. `np.asarray(frame)`, `Frame.stats()`, and `Frame.to_fits()` are explicit host-facing operations and copy GPU data. Use `getframes.to_numpy()` when a named host boundary is clearer. diff --git a/docs/reference.md b/docs/reference.md index dead207..eea1da8 100644 --- a/docs/reference.md +++ b/docs/reference.md @@ -24,6 +24,8 @@ ::: getframes.backend.get_backend +::: getframes.backend.resolve_precision + ::: getframes.backend.get_array_module ::: getframes.backend.to_numpy diff --git a/pyproject.toml b/pyproject.toml index edf7eed..4eb6eec 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -45,6 +45,8 @@ dependencies = [ [project.optional-dependencies] dev = [ + # The test suite runs aocore's image-builder conformance checks (0.1.3+). + "aocore>=0.1.3,<0.2", "pytest>=7.0", "pytest-cov>=4.0", "ruff>=0.6", diff --git a/src/getframes/__about__.py b/src/getframes/__about__.py index f3d6acb..e613ae3 100644 --- a/src/getframes/__about__.py +++ b/src/getframes/__about__.py @@ -1,4 +1,4 @@ # SPDX-License-Identifier: MIT """Single source of truth for the package version.""" -__version__ = "2.3.0" +__version__ = "2.4.0" diff --git a/src/getframes/__init__.py b/src/getframes/__init__.py index f1f5efa..e6df6b8 100644 --- a/src/getframes/__init__.py +++ b/src/getframes/__init__.py @@ -21,7 +21,7 @@ nondestructive_stack_statistics, ramp_photon_transfer, ) -from .backend import ArrayBackend, get_array_module, get_backend, to_numpy +from .backend import ArrayBackend, get_array_module, get_backend, resolve_precision, to_numpy from .calibrate import calibrate, combine from .camera import Camera from .config import CameraConfig, SensorType @@ -108,5 +108,6 @@ "load_preset", "nondestructive_stack_statistics", "ramp_photon_transfer", + "resolve_precision", "to_numpy", ] diff --git a/src/getframes/backend.py b/src/getframes/backend.py index fa399ac..6608601 100644 --- a/src/getframes/backend.py +++ b/src/getframes/backend.py @@ -10,10 +10,34 @@ from __future__ import annotations +import contextlib +from collections.abc import Callable, Generator, Iterator from dataclasses import dataclass -from typing import Any, cast +from functools import wraps +from inspect import isgeneratorfunction +from typing import Any, ParamSpec, TypeVar, cast import numpy as np +from numpy.typing import DTypeLike + +_P = ParamSpec("_P") +_R = TypeVar("_R") + +_CPU_NAMES = frozenset({"cpu", "numpy"}) +_GPU_NAMES = frozenset({"gpu", "cuda", "cupy"}) +_DEVICE_HELP = ( + "expected 'cpu', 'gpu', 'gpu:N' (CUDA device N) or 'auto' " + "(aliases: 'numpy' for 'cpu'; 'cuda'/'cupy' for 'gpu')" +) + +# Working-precision vocabulary (aocore CONVENTIONS 8.2): the canonical names are +# "single"/"double"; "float32"/"float64" are aliases. +_PRECISIONS: dict[str, type[np.floating[Any]]] = { + "single": np.float32, + "double": np.float64, + "float32": np.float32, + "float64": np.float64, +} def _cupy_seed(seed: Any) -> int | None: @@ -66,29 +90,67 @@ def seed(self, seed: Any) -> None: @dataclass(frozen=True) class ArrayBackend: - """Array namespace and RNG factory for one detector execution device.""" + """Array namespace and RNG factory for one detector execution device. + + Attributes + ---------- + xp: + The array module, :mod:`numpy` or :mod:`cupy`. + device: + Device kind, ``"cpu"`` or ``"gpu"``. + device_id: + CUDA device number the GPU backend allocates on and launches kernels on + (``None`` for the CPU backend). :func:`get_backend` always fills it in for + a GPU backend, so a camera stays on one card whatever device is current + when it is later called. + """ xp: Any device: str + device_id: int | None = None @property def is_cpu(self) -> bool: """Whether arrays live in host NumPy storage.""" return self.device == "cpu" + @property + def spec(self) -> str: + """The ``device`` string that selects this backend again (``"cpu"`` or ``"gpu:N"``).""" + if self.is_cpu or self.device_id is None: + return self.device + return f"{self.device}:{self.device_id}" + + def activate(self) -> contextlib.AbstractContextManager[Any]: + """Context that makes this backend's CUDA device current. + + A no-op on the CPU backend. Camera methods enter it themselves; wrap direct + calls to the low-level :mod:`getframes.noise` functions in it when using a + GPU other than the current one. + """ + if self.is_cpu or self.device_id is None: + return contextlib.nullcontext() + return cast(contextlib.AbstractContextManager[Any], self.xp.cuda.Device(self.device_id)) + def asarray(self, value: Any, *, dtype: Any | None = None) -> Any: """Convert ``value`` to an array on this backend.""" - return self.xp.asarray(value, dtype=dtype) + if self.is_cpu: + return self.xp.asarray(value, dtype=dtype) + with self.activate(): + return self.xp.asarray(value, dtype=dtype) def default_rng(self, seed: Any = None, *, float_dtype: Any = np.float64) -> Any: - """Create a backend-native random generator.""" + """Create a backend-native random generator (on this backend's device).""" if self.is_cpu: return self.xp.random.default_rng(seed) # CuPy's Generator construction initializes device-side state and is much # slower than RandomState for the per-exposure seed contract. RandomState # still owns an independent, backend-native cuRAND stream and exposes all - # distributions used by the detector chain. - return _CuPyGenerator(self.xp.random.RandomState(_cupy_seed(seed)), self.xp, float_dtype) + # distributions used by the detector chain. The cuRAND state is created on + # the selected device, so it must be built inside the device context. + with self.activate(): + state = self.xp.random.RandomState(_cupy_seed(seed)) + return _CuPyGenerator(state, self.xp, float_dtype) def convolve(self, array: Any, kernel: Any) -> Any: """Convolve with constant-zero boundary conditions on this backend.""" @@ -98,7 +160,8 @@ def convolve(self, array: Any, kernel: Any) -> Any: return ndimage.convolve(array, kernel, mode="constant", cval=0.0) from cupyx.scipy import ndimage # pragma: no cover - optional CUDA dependency - return ndimage.convolve(array, kernel, mode="constant", cval=0.0) + with self.activate(): + return ndimage.convolve(array, kernel, mode="constant", cval=0.0) def scalar(self, value: Any) -> float: """Transfer one scalar to the host for validation or metadata.""" @@ -115,36 +178,214 @@ def to_numpy(self, value: Any) -> np.ndarray[Any, Any]: _CPU_BACKEND = ArrayBackend(np, "cpu") -def get_backend(device: str = "cpu") -> ArrayBackend: - """Return the backend for ``device`` (``"cpu"`` or ``"gpu"``). +def _import_cupy() -> Any: + """Import CuPy lazily (raises :class:`ImportError` when it is not installed).""" + import cupy + + return cupy + + +def _gpu_device_count(cupy: Any) -> int: + """Number of CUDA devices CuPy can use (``0`` when the runtime is unusable).""" + try: + return int(cupy.cuda.runtime.getDeviceCount()) + except Exception: # pragma: no cover - depends on the local CUDA driver/runtime + return 0 + + +def _parse_device(device: str) -> tuple[str, int | None]: + """Split a device string into ``("cpu" | "gpu" | "auto", index or None)``.""" + if not isinstance(device, str): + raise TypeError(f"device must be a string, got {type(device).__name__}; {_DEVICE_HELP}.") + name = device.strip().lower() + base, sep, index = name.partition(":") + if base == "auto" and not sep: + return "auto", None + if base in _CPU_NAMES and not sep: + return "cpu", None + if base in _GPU_NAMES: + if not sep: + return "gpu", None + if index.isdigit(): + return "gpu", int(index) + raise ValueError(f"unknown device {device!r}; {_DEVICE_HELP}.") + + +def _gpu_backend(device: str, index: int | None) -> ArrayBackend: + try: + cupy = _import_cupy() + except ImportError as exc: + raise ImportError( + f"device={device!r} requires CuPy; install getframes[gpu], " + "or use device='auto' to fall back to the CPU." + ) from exc + count = _gpu_device_count(cupy) + if count < 1: + raise RuntimeError( + f"device={device!r} needs a CUDA device, but CuPy sees none; " + "use device='auto' to fall back to the CPU." + ) + if index is None: + index = int(cupy.cuda.runtime.getDevice()) + elif index >= count: + raise ValueError( + f"device={device!r} asks for CUDA device {index}, but CuPy sees {count} " + f"device(s) (numbered 0-{count - 1}; CUDA_VISIBLE_DEVICES controls the numbering)." + ) + return ArrayBackend(cupy, "gpu", index) - CuPy is an optional dependency and is imported lazily only for ``"gpu"``. - ``"cuda"`` and ``"cupy"`` are accepted aliases. + +def get_backend(device: str = "cpu") -> ArrayBackend: + """Return the backend for ``device``. + + Parameters + ---------- + device: + ``"cpu"`` (NumPy, the reference), ``"gpu"`` (CuPy on the current CUDA + device), ``"gpu:N"`` (CuPy on CUDA device ``N``), or ``"auto"`` (the + current CUDA device when CuPy is installed and sees one, else the CPU). + Matching is case-insensitive; ``"numpy"`` is an alias of ``"cpu"`` and + ``"cuda"``/``"cupy"`` of ``"gpu"`` (including ``"cuda:N"``). + + Raises + ------ + ValueError + For an unknown device string, or a GPU number CuPy does not see. + ImportError + For ``"gpu"``/``"gpu:N"`` without CuPy installed. + RuntimeError + For ``"gpu"``/``"gpu:N"`` when CuPy is installed but sees no CUDA device. + + Notes + ----- + CuPy is an optional dependency and is imported lazily, only for a GPU (or + ``"auto"``) device. """ - name = str(device).lower() - if name in {"cpu", "numpy"}: + kind, index = _parse_device(device) + if kind == "cpu": return _CPU_BACKEND - if name in {"gpu", "cuda", "cupy"}: + if kind == "auto": + try: + cupy = _import_cupy() + except ImportError: + return _CPU_BACKEND + if _gpu_device_count(cupy) < 1: + return _CPU_BACKEND + return _gpu_backend(device, index) + + +def resolve_precision(precision: str | DTypeLike) -> np.dtype[Any]: + """Return the floating-point working dtype named by ``precision``. + + Parameters + ---------- + precision: + ``"single"`` or ``"double"`` (the shared AO-stack vocabulary), their + aliases ``"float32"``/``"float64"`` (case-insensitive), or a NumPy + ``float32``/``float64`` dtype. + + Raises + ------ + ValueError + For any other precision. + """ + if isinstance(precision, str): + selected = _PRECISIONS.get(precision.strip().lower()) + if selected is not None: + return np.dtype(selected) + elif precision is not None: # np.dtype(None) would silently mean float64 try: - import cupy - except ImportError as exc: # pragma: no cover - depends on optional install - raise ImportError(f"device={device!r} requires CuPy; install getframes[gpu]") from exc - return ArrayBackend(cupy, "gpu") - raise ValueError(f"unknown device {device!r}; expected 'cpu' or 'gpu'.") + dtype = np.dtype(precision) + except TypeError: + dtype = None + if dtype is not None and dtype in (np.dtype(np.float32), np.dtype(np.float64)): + return dtype + raise ValueError( + f"precision must be 'single' or 'double' (aliases 'float32'/'float64'), got {precision!r}." + ) + + +def _working_dtype( + dtype: DTypeLike | None, precision: str | None, *, name: str = "dtype" +) -> np.dtype[Any]: + """Merge a legacy ``dtype`` argument with the ``precision`` vocabulary. + + ``None`` for both gives ``float64``. Both may be given only when they agree. + """ + from_dtype = None if dtype is None else np.dtype(dtype) + if precision is None: + return np.dtype(np.float64) if from_dtype is None else from_dtype + from_precision = resolve_precision(precision) + if from_dtype is not None and from_dtype != from_precision: + raise ValueError( + f"{name}={from_dtype.name!r} conflicts with precision={precision!r}; pass one of them." + ) + return from_precision + + +def _on_device(method: Callable[_P, _R]) -> Callable[_P, _R]: + """Run a method of an object with a ``_backend`` inside that backend's device context. + + Generator methods enter the context for every step, so lazily produced frames + are also computed on the object's device. + """ + if isgeneratorfunction(method): + + @wraps(method) + def generator_wrapper(*args: _P.args, **kwargs: _P.kwargs) -> Any: + backend: ArrayBackend = args[0]._backend # type: ignore[attr-defined] + iterator = cast(Generator[Any, None, Any], method(*args, **kwargs)) + return _stepped_on(backend, iterator) + + return cast(Callable[_P, _R], generator_wrapper) + + @wraps(method) + def wrapper(*args: _P.args, **kwargs: _P.kwargs) -> _R: + backend: ArrayBackend = args[0]._backend # type: ignore[attr-defined] + with backend.activate(): + return method(*args, **kwargs) + + return wrapper + + +def _stepped_on(backend: ArrayBackend, iterator: Generator[Any, None, Any]) -> Iterator[Any]: + """Advance ``iterator`` one item at a time inside ``backend``'s device context.""" + while True: + with backend.activate(): + try: + item = next(iterator) + except StopIteration: + return + yield item + + +def _array_device(value: Any) -> contextlib.AbstractContextManager[Any]: + """Context that makes ``value``'s CUDA device current (a no-op for host arrays).""" + if type(value).__module__.split(".", 1)[0] == "cupy": + return cast(contextlib.AbstractContextManager[Any], value.device) + return contextlib.nullcontext() def get_array_module(value: Any) -> Any: """Return NumPy or CuPy for an existing array without copying it.""" module = type(value).__module__.split(".", 1)[0] if module == "cupy": - return get_backend("gpu").xp + return _import_cupy() return np def to_numpy(value: Any) -> np.ndarray[Any, Any]: """Return ``value`` in host NumPy storage, copying device arrays explicitly.""" module = type(value).__module__.split(".", 1)[0] - return get_backend("gpu" if module == "cupy" else "cpu").to_numpy(value) + if module == "cupy": + return cast(np.ndarray[Any, Any], _import_cupy().asnumpy(value)) + return np.asarray(value) -__all__ = ["ArrayBackend", "get_array_module", "get_backend", "to_numpy"] +__all__ = [ + "ArrayBackend", + "get_array_module", + "get_backend", + "resolve_precision", + "to_numpy", +] diff --git a/src/getframes/camera.py b/src/getframes/camera.py index 1087350..35f0dda 100644 --- a/src/getframes/camera.py +++ b/src/getframes/camera.py @@ -4,12 +4,12 @@ from __future__ import annotations from dataclasses import replace -from typing import TYPE_CHECKING, Any, ClassVar +from typing import TYPE_CHECKING, Any import numpy as np from . import noise -from .backend import ArrayBackend, get_backend +from .backend import ArrayBackend, _on_device, get_backend, resolve_precision from .config import CameraConfig from .frame import Frame, FrameTruth from .observation import Observation, ObservationTruth, Pointing @@ -52,23 +52,25 @@ class Camera: Optional seed for this camera's internal random generator, giving reproducible output across calls when no per-call seed is supplied. precision: - Working floating-point precision of the signal chain: ``"float64"`` (the - exact default) or ``"float32"`` for the memory-light fast path — half the + Working floating-point precision of the signal chain: ``"double"`` (the + exact default) or ``"single"`` for the memory-light fast path — half the per-pixel memory, useful for large detectors and bulk dataset generation. - The digitised ADU stay integer either way; only the floating-point arrays - (including each frame's ground truth) change. + ``"float64"`` and ``"float32"`` are aliases (case-insensitive). The + digitised ADU stay integer either way; only the floating-point arrays + (including each frame's ground truth) change. :attr:`precision` reports + the NumPy dtype name (``"float64"`` or ``"float32"``) whichever spelling + was passed. device: Execution device for detector arrays and random sampling: ``"cpu"`` - (NumPy, default) or ``"gpu"`` (optional CuPy). GPU frames keep their ADU - and truth arrays on device; CPU and GPU seeds are reproducible within a - backend but intentionally do not produce identical random samples. + (NumPy, default), ``"gpu"`` (optional CuPy, on the current CUDA device), + ``"gpu:N"`` (CUDA device ``N``), or ``"auto"`` (the GPU when CuPy and a + CUDA device are usable, else the CPU); see + :func:`~getframes.backend.get_backend`. Fixed-pattern maps, the cuRAND + streams, and every frame live on the selected device. GPU frames keep + their ADU and truth arrays on device; CPU and GPU seeds are reproducible + within a backend but intentionally do not produce identical random samples. """ - _PRECISIONS: ClassVar[dict[str, type[np.floating[Any]]]] = { - "float32": np.float32, - "float64": np.float64, - } - def __init__( self, config: CameraConfig, @@ -80,18 +82,15 @@ def __init__( ) -> None: if not isinstance(config, CameraConfig): raise TypeError("config must be a CameraConfig instance.") - if precision not in self._PRECISIONS: - raise ValueError( - f"precision must be one of {sorted(self._PRECISIONS)}, got {precision!r}." - ) + float_dtype = resolve_precision(precision) self.config = config self.default_temperature_c = ( default_temperature_c if default_temperature_c is not None else config.dark_current_ref_temp_c ) - self.precision = precision - self._float_dtype = self._PRECISIONS[precision] + self.precision = float_dtype.name + self._float_dtype = float_dtype.type self._backend: ArrayBackend = get_backend(device) self._rng = self._backend.default_rng(seed, float_dtype=self._float_dtype) self._seeded_rng = ( @@ -99,9 +98,10 @@ def __init__( if self._backend.is_cpu else self._backend.default_rng(0, float_dtype=self._float_dtype) ) - self._fixed_patterns = noise.fixed_pattern_maps( - config, backend=self._backend, float_dtype=self._float_dtype - ) + with self._backend.activate(): + self._fixed_patterns = noise.fixed_pattern_maps( + config, backend=self._backend, float_dtype=self._float_dtype + ) self._dark_signal_cache_key: tuple[float, float] | None = None self._dark_signal_cache: Any | None = None @@ -146,16 +146,21 @@ def sensor_type(self) -> str: @property def device(self) -> str: - """Execution device (``"cpu"`` or ``"gpu"``).""" + """Execution device kind (``"cpu"`` or ``"gpu"``); see :attr:`device_id`.""" return self._backend.device + @property + def device_id(self) -> int | None: + """CUDA device number of a GPU camera (``None`` on the CPU).""" + return self._backend.device_id + def with_config(self, **changes: Any) -> Camera: - """Return a new camera with configuration fields overridden.""" + """Return a new camera with configuration fields overridden (same device).""" return Camera( self.config.replace(**changes), default_temperature_c=self.default_temperature_c, precision=self.precision, - device=self.device, + device=self._backend.spec, ) # ------------------------------------------------------------------ @@ -254,6 +259,7 @@ def _crop_to_roi(self, value: Any, binning: int = 1) -> Any: return value return value[self._binned_roi_slices(binning)] + @_on_device def dark_frame( self, exposure: float, @@ -294,6 +300,7 @@ def dark_frame( data = self._crop_to_roi(data) return Frame(data=data, metadata=self._metadata("dark", exposure, temp, seed)) + @_on_device def dark_series( self, exposure: float, @@ -312,6 +319,7 @@ def dark_series( frame.metadata["frame_index"] = i yield frame + @_on_device def nondestructive_series( self, photon_rate: PhotonRate, @@ -601,6 +609,7 @@ def _ramp_reads( ) yield Frame(data=data, metadata=metadata, truth=truth) + @_on_device def dark_nondestructive_series( self, read_interval: float, @@ -625,6 +634,7 @@ def dark_nondestructive_series( include_truth=include_truth, ) + @_on_device def correlated_double_sample( self, photon_rate: PhotonRate, @@ -757,6 +767,7 @@ def correlated_double_sample( ) return Frame(data=data, metadata=metadata, truth=truth) + @_on_device def expose( self, photon_rate: PhotonRate, @@ -949,6 +960,7 @@ def _expose_into( metadata["binning_mode"] = binning_mode return Frame(data=result.adu, metadata=metadata, truth=truth) + @_on_device def expose_spectral( self, photon_rate_cube: NDArray[np.floating[Any]], @@ -1044,6 +1056,7 @@ def _fold_spectral_cube( qe = self._backend.asarray(self.config.qe_curve(wavelengths_host), dtype=self._float_dtype) return cube, wavelengths_host, xp.tensordot(qe, cube, axes=(0, 0)), xp.sum(cube, axis=0) + @_on_device def correlated_double_sample_spectral( self, photon_rate_cube: NDArray[np.floating[Any]], @@ -1105,6 +1118,7 @@ def _binned_extra(self, extra_electrons: PhotonRate, binning: int) -> Any: return extra * float(binning * binning) return noise.block_sum(extra, binning) + @_on_device def flat_frame( self, photon_rate: PhotonRate, @@ -1132,6 +1146,7 @@ def flat_frame( frame.metadata["frame_type"] = "flat" return frame + @_on_device def bias_frame( self, temperature: float | None = None, @@ -1143,6 +1158,7 @@ def bias_frame( frame.metadata["frame_type"] = "bias" return frame + @_on_device def observe( self, scene: Scene, @@ -1210,6 +1226,7 @@ def _tag_science(frame: Frame, scene: Scene, spectral: bool) -> None: if scene.wcs is not None: frame.metadata.update(scene.wcs.header_cards()) + @_on_device def expose_series( self, photon_rate: PhotonRate, @@ -1249,6 +1266,7 @@ def expose_series( frame.metadata["frame_index"] = i yield frame + @_on_device def observe_series( self, scene: Scene, @@ -1417,6 +1435,7 @@ def _source_names(sources: Sequence[Source]) -> list[str]: # ------------------------------------------------------------------ # Calibration masters # ------------------------------------------------------------------ + @_on_device def master_bias( self, n_frames: int, @@ -1431,6 +1450,7 @@ def master_bias( frames = (self.bias_frame(temperature, seed=s) for s in self._series_seeds(seed, n_frames)) return combine(frames, method=method) + @_on_device def master_dark( self, exposure: float, @@ -1449,6 +1469,7 @@ def master_dark( return combine(self.dark_series(exposure, n_frames, temperature, seed=seed), method=method) + @_on_device def master_flat( self, photon_rate: PhotonRate, @@ -1517,5 +1538,5 @@ def __repr__(self) -> str: h, w = self.resolution return ( f"Camera(name={self.config.name!r}, sensor={self.config.sensor_type.value!r}, " - f"resolution={h}x{w}, device={self.device!r})" + f"resolution={h}x{w}, device={self._backend.spec!r})" ) diff --git a/src/getframes/cli.py b/src/getframes/cli.py index ab8f793..81c0a75 100644 --- a/src/getframes/cli.py +++ b/src/getframes/cli.py @@ -23,6 +23,7 @@ import numpy as np from .__about__ import __version__ +from .backend import to_numpy from .camera import Camera from .config import CameraConfig from .frame import Frame @@ -34,7 +35,7 @@ import tomli as tomllib # CLI-only keys in a [camera] table that are not CameraConfig fields. -_CAMERA_META_KEYS = ("preset", "default_temperature_c", "precision") +_CAMERA_META_KEYS = ("preset", "default_temperature_c", "precision", "device") def _load_toml(path: str) -> dict[str, Any]: @@ -46,8 +47,10 @@ def _camera_from_config(cfg: dict[str, Any]) -> Camera: """Build a :class:`Camera` from a config's ``[camera]`` table. Either ``preset = ""`` (with optional overrides) or an inline - :class:`~getframes.config.CameraConfig` table. ``default_temperature_c`` and - ``precision`` are camera-construction options, not config fields. + :class:`~getframes.config.CameraConfig` table. ``default_temperature_c``, + ``precision`` (``"double"``/``"single"``) and ``device`` (``"cpu"``, + ``"gpu"``, ``"gpu:N"`` or ``"auto"``) are camera-construction options, not + config fields. """ cam_cfg = dict(cfg.get("camera", {})) kwargs: dict[str, Any] = {} @@ -55,6 +58,8 @@ def _camera_from_config(cfg: dict[str, Any]) -> Camera: kwargs["default_temperature_c"] = float(cam_cfg["default_temperature_c"]) if "precision" in cam_cfg: kwargs["precision"] = str(cam_cfg["precision"]) + if "device" in cam_cfg: + kwargs["device"] = str(cam_cfg["device"]) if "preset" in cam_cfg: config = load_preset(str(cam_cfg["preset"])) @@ -72,11 +77,11 @@ def _write_frame(frame: Frame, path: str) -> None: if suffix in (".fits", ".fit"): frame.to_fits(path, overwrite=True) elif suffix == ".npy": - np.save(path, np.asarray(frame.data)) + np.save(path, to_numpy(frame.data)) elif suffix == ".npz": - raw = np.asarray(frame.data) + raw = to_numpy(frame.data) if frame.truth is not None: - np.savez(path, raw=raw, truth=np.asarray(frame.truth.mean_electrons)) + np.savez(path, raw=raw, truth=to_numpy(frame.truth.mean_electrons)) else: np.savez(path, raw=raw) else: @@ -210,7 +215,7 @@ def main(argv: list[str] | None = None) -> int: args = parser.parse_args(argv) try: exit_code: int = args.func(args) - except (ValueError, KeyError, FileNotFoundError) as exc: + except (ValueError, KeyError, FileNotFoundError, ImportError, RuntimeError) as exc: parser.error(str(exc)) return exit_code diff --git a/src/getframes/dataset.py b/src/getframes/dataset.py index dee7705..e4d2169 100644 --- a/src/getframes/dataset.py +++ b/src/getframes/dataset.py @@ -28,6 +28,7 @@ import numpy as np +from .backend import to_numpy from .scene import Bandpass, GaussianPSF, PointSource, Scene, Sky, Telescope if TYPE_CHECKING: @@ -217,8 +218,8 @@ def __iter__(self) -> Iterator[Pair]: ) assert frame.truth is not None # include_truth=True yield { - "raw": np.asarray(frame.data, dtype=self.dtype), - "truth": np.asarray(frame.truth.mean_electrons, dtype=self.dtype), + "raw": np.asarray(to_numpy(frame.data), dtype=self.dtype), + "truth": np.asarray(to_numpy(frame.truth.mean_electrons), dtype=self.dtype), } def to_npz(self, directory: str, *, prefix: str = "pair", compress: bool = False) -> list[str]: diff --git a/src/getframes/frame.py b/src/getframes/frame.py index 35866a4..66f8bd0 100644 --- a/src/getframes/frame.py +++ b/src/getframes/frame.py @@ -9,7 +9,7 @@ import numpy as np from aocore import block_sum -from .backend import get_array_module, to_numpy +from .backend import _array_device, get_array_module, to_numpy if TYPE_CHECKING: from numpy.typing import NDArray @@ -126,12 +126,13 @@ def binned(self, factor: int, *, method: str = "sum") -> Frame: raise ValueError( f"frame shape {data.shape} is not divisible by binning factor {factor}." ) - if factor == 1: - binned = data.copy() - else: - binned = block_sum(data, factor) - if method == "mean": - binned = binned / (factor * factor) + with _array_device(data): + if factor == 1: + binned = data.copy() + else: + binned = block_sum(data, factor) + if method == "mean": + binned = binned / (factor * factor) metadata = dict(self.metadata) metadata["binning"] = int(metadata.get("binning", 1)) * factor return Frame(data=binned, metadata=metadata, truth=None) diff --git a/src/getframes/noise.py b/src/getframes/noise.py index 1193a11..0e409b6 100644 --- a/src/getframes/noise.py +++ b/src/getframes/noise.py @@ -39,7 +39,7 @@ import numpy as np from aocore import block_sum -from .backend import ArrayBackend, get_backend +from .backend import ArrayBackend, _working_dtype, get_backend if TYPE_CHECKING: from numpy.typing import DTypeLike, NDArray @@ -108,6 +108,8 @@ def __init__(self) -> None: def _device_index(backend: ArrayBackend) -> int | None: if backend.is_cpu: return None + if backend.device_id is not None: + return backend.device_id return int(backend.xp.cuda.runtime.getDevice()) @contextmanager @@ -302,9 +304,16 @@ def fixed_pattern_maps( config: CameraConfig, *, backend: ArrayBackend | None = None, - float_dtype: DTypeLike = DEFAULT_FLOAT_DTYPE, + float_dtype: DTypeLike | None = None, + precision: str | None = None, ) -> FixedPatternMaps: - """Build all repeatable per-pixel detector maps once on the selected device.""" + """Build all repeatable per-pixel detector maps once on the selected device. + + The working precision is ``float_dtype`` or, equivalently, ``precision`` + (``"single"``/``"double"``, aliases ``"float32"``/``"float64"``); both default + to ``float64`` and may be combined only when they agree. + """ + float_dtype = _working_dtype(float_dtype, precision, name="float_dtype") resolved = backend or get_backend() xp = resolved.xp shape = config.resolution @@ -447,10 +456,11 @@ def dark_signal_map( config: CameraConfig, exposure_s: float, temperature_c: float, - float_dtype: DTypeLike = DEFAULT_FLOAT_DTYPE, + float_dtype: DTypeLike | None = None, *, backend: ArrayBackend | None = None, fixed_patterns: FixedPatternMaps | None = None, + precision: str | None = None, ) -> Any: """Per-pixel *mean* dark signal in electrons, including fixed-pattern structure. @@ -462,8 +472,10 @@ def dark_signal_map( exposure-scaled and dark-removable. ``float_dtype`` selects the working precision (``float64`` exact default, or - ``float32`` for the memory-light fast path). + ``float32`` for the memory-light fast path); ``precision`` (``"single"`` / + ``"double"``) is the equivalent named spelling. """ + float_dtype = _working_dtype(float_dtype, precision, name="float_dtype") resolved = backend or get_backend() xp = resolved.xp height, width = config.resolution @@ -504,11 +516,12 @@ def photo_signal_map( exposure_s: float, background_photon_rate: PhotonRate, quantum_efficiency: float | None = None, - float_dtype: DTypeLike = DEFAULT_FLOAT_DTYPE, + float_dtype: DTypeLike | None = None, *, backend: ArrayBackend | None = None, fixed_patterns: FixedPatternMaps | None = None, out: Any | None = None, + precision: str | None = None, ) -> Any: """Per-pixel *mean* photo-generated signal in electrons (noise-free). @@ -524,8 +537,10 @@ def photo_signal_map( ``quantum_efficiency`` overrides ``config.quantum_efficiency`` when given. The spectral path uses this with a pre-multiplied (already-photoelectron) map and ``quantum_efficiency = 1.0``. ``float_dtype`` selects the working precision - (``float64`` default, or ``float32`` for the memory-light fast path). + (``float64`` default, or ``float32`` for the memory-light fast path); + ``precision`` (``"single"`` / ``"double"``) is the equivalent named spelling. """ + float_dtype = _working_dtype(float_dtype, precision, name="float_dtype") resolved = backend or get_backend() xp = resolved.xp height, width = config.resolution @@ -1227,7 +1242,8 @@ def simulate_frame( binning_mode: str = "digital", rng: Any | None = None, seed: int | None = None, - float_dtype: DTypeLike = DEFAULT_FLOAT_DTYPE, + float_dtype: DTypeLike | None = None, + precision: str | None = None, backend: ArrayBackend | None = None, fixed_patterns: FixedPatternMaps | None = None, _dark_signal: Any | None = None, @@ -1281,6 +1297,10 @@ def simulate_frame( exact default) or ``float32`` for the memory-light fast path used for large detectors and bulk dataset generation. The digitised ADU stay integer regardless; only the floating-point signal chain and the truth arrays change. + precision: + The same choice by name: ``"double"`` or ``"single"`` (aliases + ``"float64"``/``"float32"``). Give either this or ``float_dtype``, or both + only when they agree. workspace: Optional reusable :class:`DetectorWorkspace`. Scratch arrays are private and never escape in the returned result. A workspace is sequential-use @@ -1291,6 +1311,7 @@ def simulate_frame( owns its lifetime and must not reuse it while a consumer still needs the frame. """ + float_dtype = _working_dtype(float_dtype, precision, name="float_dtype") resolved = backend or get_backend() if exposure_s < 0: raise ValueError("exposure_s must be non-negative.") diff --git a/src/getframes/scene/scene.py b/src/getframes/scene/scene.py index 88a6744..1dceff2 100644 --- a/src/getframes/scene/scene.py +++ b/src/getframes/scene/scene.py @@ -10,6 +10,7 @@ import numpy as np from numpy.typing import NDArray +from ..backend import _working_dtype from .optics import Telescope from .psf import PSF from .sources import RenderContext, Sky, Source @@ -116,7 +117,9 @@ def photon_rate_map( self, time_s: float | None = None, offset_xy: tuple[float, float] = (0.0, 0.0), - dtype: DTypeLike = np.float64, + dtype: DTypeLike | None = None, + *, + precision: str | None = None, ) -> NDArray[np.float64]: """Render the sources through the PSF into a photons/s/pixel map. @@ -135,8 +138,13 @@ def photon_rate_map( dtype: Output (and working) floating-point dtype. ``float64`` is the exact default; ``float32`` halves the map's memory for the fast path. + precision: + The same choice by name: ``"double"`` or ``"single"`` (aliases + ``"float64"``/``"float32"``). Give either this or ``dtype``, or both + only when they agree. """ - return self._render(lambda _sed: 1.0, time_s, offset_xy, dtype) + working = _working_dtype(dtype, precision) + return self._render(lambda _sed: 1.0, time_s, offset_xy, working) def sky_photon_rate(self) -> float: """Uniform sky background in photons/s/pixel (``0`` if no sky is set).""" @@ -160,7 +168,9 @@ def photoelectron_rate_map( qe_curve: QE, time_s: float | None = None, offset_xy: tuple[float, float] = (0.0, 0.0), - dtype: DTypeLike = np.float64, + dtype: DTypeLike | None = None, + *, + precision: str | None = None, ) -> NDArray[np.float64]: """Render sources to a *photoelectron*-rate map (e-/s/pixel) in spectral mode. @@ -169,14 +179,18 @@ def photoelectron_rate_map( detector ``qe_curve`` with the band's spectral response). The result is already in photoelectrons, so the camera applies a unit QE downstream. - ``time_s`` and ``offset_xy`` behave as in :meth:`photon_rate_map`. + ``time_s``, ``offset_xy``, ``dtype`` and ``precision`` behave as in + :meth:`photon_rate_map`. Requires a band with a spectral response (see :attr:`is_spectral_capable`). """ band = self.optics.band if band is None or band.response is None: raise ValueError("photoelectron_rate_map requires a band with a spectral response.") - return self._render(lambda sed: band.effective_qe(qe_curve, sed), time_s, offset_xy, dtype) + working = _working_dtype(dtype, precision) + return self._render( + lambda sed: band.effective_qe(qe_curve, sed), time_s, offset_xy, working + ) def sky_electron_rate(self, qe_curve: QE) -> float: """Uniform sky background in photoelectrons/s/pixel for spectral mode.""" diff --git a/tests/test_backend.py b/tests/test_backend.py new file mode 100644 index 0000000..8df8420 --- /dev/null +++ b/tests/test_backend.py @@ -0,0 +1,242 @@ +# SPDX-License-Identifier: MIT +"""Device and precision vocabulary (aocore CONVENTIONS 8.1-8.2), without a GPU. + +CuPy is replaced by a small fake module, or hidden entirely, so these run on any +CI machine; the real-GPU twins live in ``test_gpu.py``. +""" + +from __future__ import annotations + +import sys +from types import SimpleNamespace +from typing import Any + +import numpy as np +import pytest + +import getframes as gf +from getframes import backend, noise + + +class _FakeDevice: + """Stand-in for ``cupy.cuda.Device`` that records the current device.""" + + def __init__(self, fake: _FakeCuPy, index: int) -> None: + self._fake = fake + self._index = index + self._previous: list[int] = [] + + def __enter__(self) -> _FakeDevice: + self._previous.append(self._fake.current) + self._fake.current = self._index + return self + + def __exit__(self, *exc: object) -> None: + self._fake.current = self._previous.pop() + + +class _FakeCuPy: + """Just enough of CuPy for device resolution and RNG construction.""" + + __name__ = "cupy" + + def __init__(self, count: int) -> None: + self.current = 0 + self.rng_devices: list[int] = [] + self.cuda = SimpleNamespace( + runtime=SimpleNamespace( + getDeviceCount=lambda: count, + getDevice=lambda: self.current, + ), + Device=lambda index: _FakeDevice(self, index), + ) + self.random = SimpleNamespace(RandomState=self._random_state) + + def _random_state(self, seed: Any) -> object: + self.rng_devices.append(self.current) + return object() + + +@pytest.fixture +def fake_cupy(monkeypatch: pytest.MonkeyPatch) -> _FakeCuPy: + fake = _FakeCuPy(count=2) + monkeypatch.setattr(backend, "_import_cupy", lambda: fake) + return fake + + +@pytest.fixture +def no_cupy(monkeypatch: pytest.MonkeyPatch) -> None: + # A ``None`` entry makes ``import cupy`` raise ImportError, installed or not. + monkeypatch.setitem(sys.modules, "cupy", None) + + +@pytest.mark.parametrize("device", ["cpu", "CPU", "numpy", " cpu "]) +def test_cpu_spellings_select_numpy(device: str) -> None: + selected = gf.get_backend(device) + assert selected.xp is np + assert selected.device == "cpu" + assert selected.device_id is None + assert selected.spec == "cpu" + + +@pytest.mark.parametrize( + "device", + ["tpu", "gpu:x", "gpu:-1", "gpu:1.0", "gpu:", "gpu:0:1", "cpu:0", "auto:0", "", "gpu 0"], +) +def test_unknown_device_strings_raise(device: str) -> None: + with pytest.raises(ValueError, match="expected 'cpu', 'gpu', 'gpu:N'"): + gf.get_backend(device) + + +def test_non_string_device_raises_type_error() -> None: + with pytest.raises(TypeError, match="device must be a string"): + gf.get_backend(0) # type: ignore[arg-type] + + +@pytest.mark.parametrize("device", ["gpu:0", "gpu:1", "cuda:1", "cupy:1", "GPU:1"]) +def test_gpu_number_selects_that_device(fake_cupy: _FakeCuPy, device: str) -> None: + selected = gf.get_backend(device) + assert selected.device == "gpu" + assert selected.device_id == int(device[-1]) + assert selected.spec == f"gpu:{device[-1]}" + + +def test_plain_gpu_pins_the_current_device(fake_cupy: _FakeCuPy) -> None: + fake_cupy.current = 1 + assert gf.get_backend("gpu").device_id == 1 + + +@pytest.mark.parametrize("device", ["gpu:2", "gpu:7"]) +def test_gpu_number_out_of_range_names_the_device_count(fake_cupy: _FakeCuPy, device: str) -> None: + with pytest.raises(ValueError, match=r"CuPy sees 2 device\(s\) \(numbered 0-1"): + gf.get_backend(device) + + +def test_gpu_rng_is_created_on_the_selected_device(fake_cupy: _FakeCuPy) -> None: + # The cuRAND state must be built with the camera's device current, not + # whichever device happens to be current at the call. + selected = gf.get_backend("gpu:1") + selected.default_rng(5) + assert fake_cupy.rng_devices == [1] + assert fake_cupy.current == 0 + + +def test_auto_uses_the_gpu_when_cupy_sees_one(fake_cupy: _FakeCuPy) -> None: + selected = gf.get_backend("auto") + assert selected.device == "gpu" + assert selected.device_id == 0 + + +def test_auto_falls_back_to_cpu_without_a_cuda_device(monkeypatch: pytest.MonkeyPatch) -> None: + monkeypatch.setattr(backend, "_import_cupy", lambda: _FakeCuPy(count=0)) + assert gf.get_backend("auto").is_cpu + with pytest.raises(RuntimeError, match="sees none"): + gf.get_backend("gpu") + + +@pytest.mark.usefixtures("no_cupy") +def test_auto_falls_back_to_cpu_when_cupy_is_missing() -> None: + assert gf.get_backend("auto").is_cpu + camera = gf.Camera.from_preset("generic_cmos", device="auto").with_config(resolution=(8, 8)) + assert camera.device == "cpu" + assert camera.device_id is None + frame = camera.expose(10.0, 0.1, seed=1) + assert isinstance(frame.data, np.ndarray) + + +@pytest.mark.usefixtures("no_cupy") +@pytest.mark.parametrize("device", ["gpu", "gpu:0", "cuda"]) +def test_gpu_without_cupy_points_at_the_extra_and_auto(device: str) -> None: + with pytest.raises(ImportError, match=r"getframes\[gpu\].*device='auto'"): + gf.get_backend(device) + + +def test_camera_reports_and_preserves_its_device() -> None: + camera = gf.Camera.from_preset("generic_cmos", device="numpy") + assert camera.device == "cpu" + assert camera.device_id is None + assert camera.with_config(resolution=(8, 8)).device == "cpu" + assert "device='cpu'" in repr(camera) + + +# ---------------------------------------------------------------- precision + + +@pytest.mark.parametrize( + ("precision", "expected"), + [ + ("single", np.float32), + ("double", np.float64), + ("float32", np.float32), + ("float64", np.float64), + ("Single", np.float32), + ("DOUBLE", np.float64), + (np.float32, np.float32), + (np.dtype(np.float64), np.float64), + ], +) +def test_precision_names_and_aliases(precision: Any, expected: type) -> None: + assert gf.resolve_precision(precision) == np.dtype(expected) + + +@pytest.mark.parametrize("precision", ["half", "float16", "complex64", "int32", np.int32, None]) +def test_unknown_precision_raises(precision: Any) -> None: + with pytest.raises(ValueError, match="precision must be 'single' or 'double'"): + gf.resolve_precision(precision) + + +@pytest.mark.parametrize( + ("precision", "name"), + [("single", "float32"), ("double", "float64"), ("float32", "float32"), ("float64", "float64")], +) +def test_camera_precision_vocabulary(precision: str, name: str) -> None: + camera = gf.Camera.from_preset("generic_cmos", precision=precision).with_config( + resolution=(8, 8) + ) + assert camera.precision == name + frame = camera.expose(10.0, 0.1, seed=1) + assert frame.truth is not None + assert frame.truth.mean_electrons.dtype == np.dtype(name) + + +def test_single_and_float32_cameras_are_the_same_camera() -> None: + single = gf.Camera.from_preset("generic_cmos", precision="single").with_config( + resolution=(8, 8) + ) + float32 = gf.Camera.from_preset("generic_cmos", precision="float32").with_config( + resolution=(8, 8) + ) + np.testing.assert_array_equal( + single.expose(50.0, 0.2, seed=3).data, float32.expose(50.0, 0.2, seed=3).data + ) + + +def test_scene_rate_maps_take_precision() -> None: + scene = gf.Scene( + shape=(16, 16), + optics=gf.Telescope(1.0, 0.5, band=gf.Bandpass.johnson("V")), + psf=gf.GaussianPSF(1.0), + sources=[gf.PointSource(x=8.0, y=8.0, magnitude=15.0)], + ) + assert scene.photon_rate_map().dtype == np.float64 + assert scene.photon_rate_map(precision="single").dtype == np.float32 + assert scene.photon_rate_map(dtype=np.float32, precision="float32").dtype == np.float32 + np.testing.assert_array_equal( + scene.photon_rate_map(precision="double"), scene.photon_rate_map(dtype=np.float64) + ) + with pytest.raises(ValueError, match="conflicts with precision"): + scene.photon_rate_map(dtype=np.float64, precision="single") + + +def test_noise_layer_takes_precision() -> None: + config = gf.load_preset("generic_cmos").replace(resolution=(8, 8)) + result = noise.simulate_frame(config, 10.0, 0.1, temperature_c=20.0, seed=0, precision="single") + assert result.mean_photoelectrons.dtype == np.float32 + maps = noise.fixed_pattern_maps(config.replace(prnu=0.01), precision="single") + assert maps.prnu_multiplier.dtype == np.float32 + dark = noise.dark_signal_map(config, 1.0, 20.0, precision="single") + assert dark.dtype == np.float32 + photo = noise.photo_signal_map(config, 10.0, 1.0, 0.0, precision="double") + assert photo.dtype == np.float64 + with pytest.raises(ValueError, match="float_dtype='float32' conflicts"): + noise.dark_signal_map(config, 1.0, 20.0, np.float32, precision="double") diff --git a/tests/test_cli.py b/tests/test_cli.py index a050511..5337a37 100644 --- a/tests/test_cli.py +++ b/tests/test_cli.py @@ -1,6 +1,8 @@ # SPDX-License-Identifier: MIT """Tests for the phase 1.6 ``getframes`` command-line interface.""" +import sys + import numpy as np import pytest @@ -91,3 +93,30 @@ def test_bad_output_extension_errors(tmp_path): ) with pytest.raises(SystemExit): cli.main(["generate", config, "-o", str(tmp_path / "frame.txt")]) + + +def test_camera_table_takes_device_and_precision(tmp_path, monkeypatch): + # The [camera] table speaks the shared vocabulary; "auto" runs on the CPU + # when CuPy is unavailable (hidden here so the test is machine-independent). + monkeypatch.setitem(sys.modules, "cupy", None) + config = _write( + tmp_path / "f.toml", + '[camera]\npreset = "generic_cmos"\nprecision = "single"\ndevice = "auto"\n\n' + '[frame]\ntype = "dark"\nexposure_s = 1.0\nseed = 0\n', + ) + camera = cli._camera_from_config(cli._load_toml(config)) + assert camera.device == "cpu" + assert camera.precision == "float32" + out = tmp_path / "dark.npz" + assert cli.main(["generate", config, "-o", str(out)]) == 0 + assert np.load(out)["raw"].shape == camera.resolution + + +@pytest.mark.parametrize("device", ["tpu", "gpu:x"]) +def test_camera_table_rejects_unknown_device(tmp_path, device): + config = _write( + tmp_path / "f.toml", + f'[camera]\npreset = "generic_cmos"\ndevice = "{device}"\n\n[frame]\ntype = "dark"\n', + ) + with pytest.raises(SystemExit): + cli.main(["generate", config]) diff --git a/tests/test_conformance.py b/tests/test_conformance.py index 0b072bf..25e00fb 100644 --- a/tests/test_conformance.py +++ b/tests/test_conformance.py @@ -4,13 +4,17 @@ getframes owns detectors, not wavefronts: its PSFs are analytic profiles (or a user kernel), not images of an OPD map. So the OPD-driven checks (tilt direction, slope sign, wind motion, Zernike basis, RMS) do not apply. What does apply is the -flat-wavefront case: where the optical axis sits in the image plane (rule 1.3) and -that a normalized point-source image carries unit flux (rule 3.3). The callables -below take the checks' OPD argument for its shape only. +image-builder case, through aocore's render checks: where the optical axis sits +in the image plane (rule 1.3), that a normalized point-source image carries unit +flux, and that light falling off a detector edge is lost rather than +renormalized back onto it (rule 3.3). + +Section 8 (software) applies in full: the ``device`` and ``precision`` words. """ from __future__ import annotations +import sys from typing import Any import aocore @@ -18,6 +22,8 @@ import pytest from aocore import conformance +import getframes as gf +from getframes import backend from getframes.scene import ( AiryPSF, ArrayPSF, @@ -43,34 +49,54 @@ def _psfs() -> list[PSF]: ] -def _on_axis_image(psf: PSF, shape: tuple[int, ...]) -> np.ndarray[Any, Any]: - """A unit-flux point source on the optical axis ``(n - 1) / 2`` (rule 1.3).""" +def _point_source_image( + psf: PSF, shape: tuple[int, int], position: tuple[float, float] = (0.0, 0.0) +) -> np.ndarray[Any, Any]: + """A unit-flux point source ``position = (y, x)`` pixels from the window centre. + + Pixel ``i`` is at ``i - (n - 1) / 2`` from the centre (rule 1.2), so the + optical axis is the pixel-index position ``(n - 1) / 2`` (rule 1.3). + """ height, width = shape image = np.zeros((height, width)) - psf.add_source(image, (width - 1) / 2.0, (height - 1) / 2.0, 1.0, PLATE_SCALE_ARCSEC) + x = (width - 1) / 2.0 + position[1] + y = (height - 1) / 2.0 + position[0] + psf.add_source(image, x, y, 1.0, PLATE_SCALE_ARCSEC) return image @pytest.mark.parametrize("psf", _psfs(), ids=lambda psf: type(psf).__name__) def test_point_source_image_carries_unit_flux(psf: PSF) -> None: # Rule 3.3: the window holds all of the light, so the image sums to the flux. - conformance.check_unit_flux(lambda opd: _on_axis_image(psf, opd.shape), pupil_shape=(96, 96)) + conformance.check_point_source_flux(lambda shape: _point_source_image(psf, shape)) @pytest.mark.parametrize("psf", _psfs(), ids=lambda psf: type(psf).__name__) -@pytest.mark.parametrize("shape", [(64, 64), (65, 65), (64, 65)]) -def test_pixel_coordinates_put_the_axis_on_pixel_centres(psf: PSF, shape: tuple[int, int]) -> None: +def test_pixel_coordinates_put_the_axis_on_pixel_centres(psf: PSF) -> None: # Rules 1.2-1.3: getframes positions are pixel-centre indices, so a source at # ``(n - 1) / 2`` images to the window centre, between pixels for even n. - conformance.check_image_centring(lambda opd: _on_axis_image(psf, opd.shape), pupil_shape=shape) + conformance.check_point_source_centring( + lambda shape: _point_source_image(psf, shape), + shapes=[(64, 64), (65, 65), (64, 65)], + ) -@pytest.mark.parametrize("shape", [(64, 64), (65, 65), (64, 65)]) -def test_vignetting_is_centred_on_the_optical_axis(shape: tuple[int, int]) -> None: +@pytest.mark.parametrize("psf", _psfs(), ids=lambda psf: type(psf).__name__) +def test_light_off_the_detector_edge_is_lost(psf: PSF) -> None: + # Rule 3.3: a source on the frame edge deposits only the light that lands on + # the detector (about half), never its whole flux renormalized into the frame. + conformance.check_edge_flux_loss( + lambda shape, position: _point_source_image(psf, shape, position) + ) + + +def test_vignetting_is_centred_on_the_optical_axis() -> None: # Rule 1.3: the field-dependent illumination falls off about ``(n - 1) / 2``. + # The pattern is centro-symmetric, so its centroid locates the axis just as + # an on-axis point source's does. vignetting = Vignetting(strength=0.4, power=2.0) - conformance.check_image_centring( - lambda opd: vignetting.illumination_map(opd.shape), pupil_shape=shape + conformance.check_point_source_centring( + vignetting.illumination_map, shapes=[(64, 64), (65, 65), (64, 65)] ) @@ -81,3 +107,43 @@ def test_arcsecond_constant_is_the_shared_one() -> None: assert psf.ARCSEC_TO_RAD is aocore.ARCSEC_TO_RAD assert thermal.ARCSEC_TO_RAD is aocore.ARCSEC_TO_RAD + + +@pytest.mark.parametrize("precision", ["single", "double", "float32", "float64"]) +def test_precision_vocabulary_is_the_shared_one(precision: str) -> None: + # Rule 8.2: "single"/"double", with "float32"/"float64" as aliases, name the + # same working dtype here as in aocore. + expected = aocore.get_backend("cpu", precision).real_dtype + assert gf.resolve_precision(precision) == expected + camera = gf.Camera.from_preset("generic_cmos", precision=precision).with_config( + resolution=(8, 8) + ) + frame = camera.expose(10.0, 0.1, seed=0) + assert frame.truth is not None + assert frame.truth.mean_electrons.dtype == expected + + +@pytest.mark.parametrize( + ("device", "kind", "index"), + [ + ("cpu", "cpu", None), + ("gpu", "gpu", None), + ("gpu:0", "gpu", 0), + ("gpu:3", "gpu", 3), + ("auto", "auto", None), + ], +) +def test_device_vocabulary_is_the_shared_one(device: str, kind: str, index: int | None) -> None: + # Rule 8.1: "cpu", "gpu", "gpu:N" and "auto" are all understood. + assert backend._parse_device(device) == (kind, index) + + +def test_auto_device_runs_without_a_gpu_and_to_numpy_is_the_host_boundary( + monkeypatch: pytest.MonkeyPatch, +) -> None: + # Rule 8.1: "auto" falls back to the CPU when no GPU is usable, and + # ``to_numpy`` hands back host storage. + monkeypatch.setitem(sys.modules, "cupy", None) + camera = gf.Camera.from_preset("generic_cmos", device="auto").with_config(resolution=(8, 8)) + assert camera.device == "cpu" + assert isinstance(gf.to_numpy(camera.expose(10.0, 0.1, seed=0).data), np.ndarray) diff --git a/tests/test_gpu.py b/tests/test_gpu.py index 978a16c..d0ead1f 100644 --- a/tests/test_gpu.py +++ b/tests/test_gpu.py @@ -220,5 +220,62 @@ def test_gpu_spectral_exposure_preserves_device_cube_truth() -> None: def test_unknown_device_is_actionable() -> None: - with pytest.raises(ValueError, match="expected 'cpu' or 'gpu'"): + with pytest.raises(ValueError, match="expected 'cpu', 'gpu', 'gpu:N'"): gf.Camera.from_preset("generic_cmos", device="tpu") + + +def test_gpu_number_selects_that_device_for_maps_rng_and_frames() -> None: + cupy = _cupy() + camera = gf.Camera.from_preset("generic_cmos", device="gpu:0", precision="single").with_config( + resolution=(32, 32), prnu=0.01 + ) + rate = cupy.full(camera.resolution, 1_000.0, dtype=cupy.float32) + + first = camera.expose(rate, 0.01, seed=17) + second = camera.expose(rate, 0.01, seed=17) + series = list(camera.dark_series(0.01, 2, seed=3)) + + assert camera.device == "gpu" + assert camera.device_id == 0 + assert camera.with_config(resolution=(16, 16)).device_id == 0 + assert "device='gpu:0'" in repr(camera) + assert camera._fixed_patterns.prnu_multiplier.device.id == 0 + assert first.data.device.id == 0 + assert first.truth.mean_electrons.dtype == cupy.float32 + assert cupy.array_equal(first.data, second.data) + assert all(frame.data.device.id == 0 for frame in series) + assert first.binned(2).data.device.id == 0 + + +def test_gpu_rng_is_built_on_the_selected_device(monkeypatch: pytest.MonkeyPatch) -> None: + cupy = _cupy() + devices: list[int] = [] + original = cupy.random.RandomState + + def recording_state(*args: object, **kwargs: object) -> object: + devices.append(int(cupy.cuda.runtime.getDevice())) + return original(*args, **kwargs) + + monkeypatch.setattr(cupy.random, "RandomState", recording_state) + gf.Camera.from_preset("generic_cmos", device="gpu:0").with_config(resolution=(8, 8)) + + assert devices and set(devices) == {0} + + +def test_auto_selects_the_gpu_when_one_is_usable() -> None: + cupy = _cupy() + camera = gf.Camera.from_preset("generic_cmos", device="auto", precision="single").with_config( + resolution=(16, 16) + ) + frame = camera.expose(cupy.full(camera.resolution, 100.0, dtype=cupy.float32), 0.01, seed=1) + + assert camera.device == "gpu" + assert camera.device_id == int(cupy.cuda.runtime.getDevice()) + assert isinstance(frame.data, cupy.ndarray) + + +def test_gpu_number_beyond_the_device_count_is_rejected() -> None: + cupy = _cupy() + count = cupy.cuda.runtime.getDeviceCount() + with pytest.raises(ValueError, match=f"CuPy sees {count} device"): + gf.Camera.from_preset("generic_cmos", device=f"gpu:{count}")