Backends & support
This matrix describes the implemented pre-alpha runtime, not a promise of production support. All operations are synchronous; tensors are contiguous and only device index 0 is exposed.
| Capability | CPU | Metal | CUDA |
|---|---|---|---|
| Float32/int32 tensor copy | Yes | Yes | Yes |
| Float32 fill, add, multiply | Yes | Yes | Yes |
| Int32 fill, add, multiply | Yes | Yes | No |
| Float32 subtract, divide, negate and scalar arithmetic | Yes | Yes | Yes |
| Int32 subtract, negate and scalar add/subtract/multiply | Yes | Yes | No |
| Contiguous reshape (float32/int32) | Yes | Yes | Yes |
| Transpose to contiguous copy (float32/int32) | Yes | Yes | Yes |
| Squeeze / expand dims views (float32/int32) | Yes | Yes | Yes |
| Basic indexing / slicing copies (float32/int32) | Yes | Yes | Yes |
| Concat / stack / split copies (float32/int32) | Yes | Yes | Yes |
| Explicit float32/int32 conversion | Yes | Yes | Yes |
| Binary arithmetic broadcasting | Yes | Yes | Yes, float32 |
| Float32 2D matmul | Yes | Custom + optional optimized path | Custom path |
| Float32 sum, max, mean | Yes | Yes | Yes |
| Int32 sum, max | Yes | Yes | No |
Reduction keepdims |
Yes | Yes | Yes, float32 |
| All-axis / multi-axis sum, max, mean | Yes | Yes | Yes, float32 |
| Float32 exp, GELU, SiLU | Yes | Yes | Yes |
| Float32 softmax, RMSNorm, LayerNorm | Yes | Yes | Yes |
| Generated kernels | Optional MLIR subset | Experimental MSL subset | Optional MLIR subset |
The mandatory reference backend. It runs without a GPU and is the baseline for device comparison tests. Start here for the simplest installation and debugging workflow. CPU-only CI covers Python 3.11, 3.12, and 3.13.
Apple Silicon is the primary Metal target. Building requires Xcode’s metal
and metallib tools. Static kernels implement tensor operations; matmul can use
an optional Apple optimized path or the custom correctness-first kernel.
cx.matmul(a, b, backend="custom") selects the custom implementation.
backend="optimized" requires the enabled Apple optimized path. macOS CI has
configurations with that path enabled and disabled. A job that skips device
tests does not establish hardware execution.
CUDA is an opt-in backend. Its recorded real-device validation is limited to a
GTX 980 Ti (sm_52) on Linux x86_64 with CUDA 12.4 and GCC 13. Tests compare
outputs with CPU references and include native contracts and GPU sanitizer runs.
Additional NVIDIA architectures and toolchain combinations still need their own validation. The hosted CUDA build job tests compilation and no-device behavior; it does not run on an NVIDIA GPU. A local push gate exercises the validated GPU. Repository-wide GPU acceptance still needs an isolated runner.
The CUDA backend copies and explicitly casts int32 tensors but rejects int32 factories and
arithmetic. Float32 matmul supports auto and custom, not optimized.
Compiler boundary
Section titled “Compiler boundary”MLIR is an optional compiler path, not a fourth hardware backend. Its current CPU/CUDA runtime integration handles a guarded float32 elementwise subset. See experimental kernels for requirements and non-goals.
ROCm, Vulkan/SPIR-V, float16/bfloat16, asynchronous streams, and multi-device execution are not implemented.
Engineering evidence
Section titled “Engineering evidence”Detailed, versioned records stay with the source: