Correctness has a reference.
CPU execution is a required part of the runtime. Accelerator operations are checked against it with explicit numerical tolerances.
INDEPENDENT · APACHE 2.0 · PRE-ALPHA
A compact tensor runtime for people who want to understand what runs underneath. Python up front. C++20 at the core. CPU, Metal, and CUDA behind one interface.
Experimental by design. Explicit about its limits.
import tensorcx as cx
device = cx.best_device()
x = cx.ones((1_000_000,), dtype=cx.float32, device=device)
y = cx.ones((1_000_000,), dtype=cx.float32, device=device)
z = x + y
print(z.cpu().numpy()[:5])
# [2. 2. 2. 2. 2.][2. 2. 2. 2. 2.]One million elements. One familiar API.01 / THE APPROACH
Explore tensor execution without losing sight of the runtime, the memory, or the hardware.
CPU execution is a required part of the runtime. Accelerator operations are checked against it with explicit numerical tolerances.
Device discovery, buffers, and operation dispatch share a contract. Platform-specific details stay inside each backend.
Custom kernels live under cx.experimental. Metal code generation and optional MLIR compilation have defined, narrow scopes.
02 / HARDWARE, HONESTLY
The required baseline. Tensor operations, matmul, reductions, and normalization.
MetalAPPLE SILICONStatic Metal kernels, an optional optimized matmul path, and experimental generated kernels.
CUDAOPT-INFloat32 primitives validated on sm_52. Broader GPU and toolchain coverage remains open.
Current scope: contiguous tensors, broadcasting, explicit devices and dtype conversion. Autograd and distributed training are outside the current runtime.
03 / OPEN THE WORKBENCH