INDEPENDENT · APACHE 2.0 · PRE-ALPHA

Tensor compute.
From Python
to silicon.

A compact tensor runtime for people who want to understand what runs underneath. Python up front. C++20 at the core. CPU, Metal, and CUDA behind one interface.

Experimental by design. Explicit about its limits.

first_tensor.pyPYTHON
import tensorcx as cx

device = cx.best_device()
x = cx.ones((1_000_000,), dtype=cx.float32, device=device)
y = cx.ones((1_000_000,), dtype=cx.float32, device=device)

z = x + y
print(z.cpu().numpy()[:5])
# [2. 2. 2. 2. 2.]
OUTPUT[2. 2. 2. 2. 2.]One million elements. One familiar API.
Python APIC++20 core backend-neutral
CPUMetalCUDA
INTERFACEPython 3.11+FOUNDATIONC++20 / nanobindEXECUTIONSynchronous & explicitLICENSEApache 2.0

01 / THE APPROACH

Small enough to follow.
Built to be understood.

Explore tensor execution without losing sight of the runtime, the memory, or the hardware.

[01]

Correctness has a reference.

CPU execution is a required part of the runtime. Accelerator operations are checked against it with explicit numerical tolerances.

[02]

The core stays neutral.

Device discovery, buffers, and operation dispatch share a contract. Platform-specific details stay inside each backend.

[03]

Experiments stay explicit.

Custom kernels live under cx.experimental. Metal code generation and optional MLIR compilation have defined, narrow scopes.

02 / HARDWARE, HONESTLY

Know what runs. And where.

Full support matrix

Current scope: contiguous tensors, broadcasting, explicit devices and dtype conversion. Autograd and distributed training are outside the current runtime.

03 / OPEN THE WORKBENCH

Start with a tensor.
Follow it all the way down.