Tensor operations
Start with import tensorcx as cx. The following example works on CPU:
import tensorcx as cx
a = cx.tensor([[1.0, 2.0], [3.0, 4.0]], device="cpu")b = cx.ones((2, 2), dtype=cx.float32, device="cpu")
product = cx.matmul(a, b)print(product.numpy())# [[3. 3.]# [7. 7.]]
print(cx.sum(a, axis=1).numpy())# [3. 7.]
print(cx.softmax(a, axis=-1).shape)# (2, 2)
print(((a - 1) / 2).reshape((4,)).numpy())# [0. 0.5 1. 1.5]
print(a.sum(axis=1, keepdims=True).shape)# (2, 1)
values = cx.tensor([[1, 2, 3], [4, 5, 6]])bias = cx.tensor([0.25, 0.5, 0.75])result = values.astype(cx.float32) + biasprint((result - result.mean(axis=-1, keepdims=True)).numpy())# [[-1.25 0. 1.25]# [-1.25 0. 1.25]]
print(values.T.numpy())# [[1 4]# [2 5]# [3 6]]
expanded = values.expand_dims((0, -1))print(expanded.shape) # (1, 2, 3, 1)print(expanded.squeeze().shape) # (2, 3)
print(a.sum().numpy()) # 10.0print(a.mean(axis=(0, 1), keepdims=True).numpy()) # [[2.5]]
print(values[:, ::-1].numpy())# [[3 2 1]# [6 5 4]]print(cx.concat([values, values], axis=0).shape) # (4, 3)print(cx.stack([values, values], axis=1).shape) # (2, 2, 3)parts = values.split([1], axis=1)print([part.shape for part in parts]) # [(2, 1), (2, 2)]
mask = values > 3print(mask.dtype) # boolprint(values[mask].numpy()) # [4 5 6]print(cx.where(mask, values, 0).numpy())# [[0 0 0]# [4 5 6]]print(mask.any(axis=1).numpy()) # [False True]
# Leading batch dimensions broadcast; the last two axes form each matrix.batch = cx.ones((2, 1, 3, 4))weights = cx.ones((1, 5, 4, 6))print((batch @ weights).shape) # (2, 5, 3, 6)print((cx.tensor([1., 2., 3.]) @ cx.tensor([4., 5., 6.])).numpy()) # 32.0API overview
Section titled “API overview”| Task | API | Contract |
|---|---|---|
| Construct | tensor, empty, zeros, ones, randn |
Explicit shape, dtype, and device; empty contents are unspecified |
| Inspect | Tensor.shape, .dtype, .device |
Metadata for a contiguous row-major tensor |
| Transfer | .to(device), .cpu(), .numpy() |
Explicit device transfers; NumPy export copies to host |
| Elementwise | a + b, a - b, a * b, a / b, -a |
Broadcast-compatible shapes, same dtype/device; real scalars work on either side; division requires float32 |
| Convert dtype | x.astype(dtype), cx.astype(x, dtype) |
Explicit float32/int32/bool conversion on the same device; copies by default |
| Reshape | x.reshape(shape), cx.reshape(x, shape) |
Shares storage; same element count; one inferred -1 allowed |
| Transpose | x.transpose(axes=None), cx.transpose(x, axes=None), x.T |
Reverses axes by default; always copies into a contiguous buffer on the same device |
| Singleton axes | x.squeeze(axis=None), x.expand_dims(axis); also cx.squeeze, cx.expand_dims |
Remove/insert size-one axes as views sharing storage |
| Index / slice | x[key] |
Integers, slices, negative steps, ellipsis and new axes; independent contiguous copies |
| Join / split | cx.concat, cx.stack, cx.split, x.split |
Same dtype/device; split uses equal sections or NumPy-style cut indices; all outputs copy |
| Compare | ==, !=, <, <=, >, >=; cx.equal, etc. |
Same dtype/device; broadcasting and matching scalars; bool output |
| Boolean logic | &, |, ^, ~; cx.logical_and, etc. |
Bool tensors/scalars; broadcasting |
| Conditional selection | cx.where(condition, x, y) |
Bool condition; same-dtype branches; three-way broadcasting |
| Boolean reduction | cx.any, cx.all; x.any, x.all |
Bool inputs; axes and keepdims like numeric reductions |
| Mask selection | x[mask], cx.masked_select(x, mask) |
Bool mask matches leading dimensions; stable contiguous copy |
| Matrix multiply | cx.matmul(a, b), a @ b |
Float32 vectors/matrices/batches; matching contracting dimensions; batch broadcasting |
| Reduce | cx.sum(x, axis=None, keepdims=False), max, mean |
All axes by default; integer or axis sequence; keepdims=True retains selected axes with size 1 |
| Activate | cx.exp(x), gelu, silu |
Float32 |
| Normalize | softmax, rmsnorm, layernorm |
Float32, explicit axis; preserves shape |
Support varies by backend, particularly for int32. See the
support matrix.
Numerical behavior
Section titled “Numerical behavior”- Float inputs are narrowed to
float32; precision loss and overflow are possible. - Broadcasting aligns axes from the right; dimensions must match or be 1.
(2, 3) + (3,)returns(2, 3). A dimension of 0 paired with 1 stays 0. No dtype promotion or device transfer occurs; the result is contiguous. astypetruncates float-to-int toward zero and rejects NaN, infinity and out-of-range values. Int-to-float can lose precision.copy=Falsereturns the input only when the dtype is unchanged; conversion still allocates.- CPU and Metal
int32add/subtract/multiply/negate wrap on overflow. CUDA int32 arithmetic is unsupported. - Scalars preserve the tensor dtype. Int32 requires an integer scalar within int32 bounds; float32 narrows real scalars before arithmetic. Booleans and complex scalars are rejected. Float division follows IEEE infinity/NaN rules.
- Reshape keeps the underlying buffer alive. Experimental Metal kernels writing
through a view also change other views sharing that storage. MLIR CPU/CUDA
launches preserve caller tensors by copying their output. NumPy export still copies.
Empty shape inference such as
(0, -1)is ambiguous and rejected. - Reductions support negative and multiple axes, such as
(0, -1); duplicates are errors.axis=Nonereduces all axes, whileaxis=()returns an independent same-shape copy under the same dtype restrictions. Multi-axis mean sums the selected group then divides once; no intermediate means are taken. Reducing a zero-size axis produces zeros for sum, NaNs for mean, and an error for max. NaNs propagate and float32 accumulation can overflow or lose precision. - Transpose accepts a full axis permutation, with negative indices; even an
identity or scalar transpose owns a new buffer. Squeeze removes all size-one
axes by default, or selected axes. Expand dims accepts one or more positions
in the final output shape:
(2, 3)at(0, -1)becomes(1, 2, 3, 1). Repeated/out-of-range axes and squeezing a non-singleton axis are errors. Both singleton operations share storage using the reshape ownership rules. - Softmax subtracts the maximum before exponentiation. NaNs propagate; slices containing positive infinity or only negative infinities produce NaNs.
- RMSNorm and LayerNorm have no affine weights. Their
epsmust be finite, non-negative, and representable as float32. - Floating-point accumulation order can differ across backends. Compare with numerical tolerances, not bitwise equality.
Basic slicing supports negative indices/steps and clips slice bounds. Integer selection returns a tensor even for a scalar. Concat joins an existing axis; stack inserts a new one. Split returns a tuple of independent copies. Advanced integer indexing, mixed mask/index tuples, indexed assignment, autograd and general strided views are not implemented. Normalizations require one explicit axis. The experimental kernel subset has its own arithmetic rules and does not extend the ordinary tensor operators.
Boolean masks
Section titled “Boolean masks”Boolean input infers cx.bool; mixed boolean/numeric input needs an explicit
dtype. Numeric casts to bool use nonzero truth, including NaN and infinity;
bool casts back to numbers as 0/1. Arithmetic on bool tensors is rejected.
Comparisons use IEEE rules: NaN compares unequal to every value, including itself.
where selects already-computed branches; it is not lazy control flow. All tensor
arguments must share a device. Mask indexing does not broadcast: a (2,) mask
on (2, 3) selects rows, while a (2, 3) mask produces a flat vector. A scalar
bool Tensor mask adds a leading axis of size 0 or 1. Selection copies preserve
order and stored bits. GPU selection reads back only the selected count to size
its output; it does not export the tensor to NumPy.
any/all accept bool inputs, all/single/multiple axes and keepdims; empty
reductions return false/true respectively. Python bool(x) requires exactly one
element and reads that value to the host; use x.any() or x.all() for a mask.
Batched matrix multiplication
Section titled “Batched matrix multiplication”matmul broadcasts leading batch dimensions and multiplies the last two axes.
For example, (2, 1, 3, 4) @ (1, 5, 4, 6) produces (2, 5, 3, 6).
Rank-one inputs are promoted temporarily: (k,) @ (k,) returns a scalar Tensor,
(m, k) @ (k,) returns (m,), and (k,) @ (..., k, n) returns (..., n).
Scalar operands are rejected. Float32 dtype and device must match.
Empty batch/row/column dimensions yield empty outputs; a zero contracting dimension produces zeros. Outputs own independent contiguous storage. Neither broadcasting nor matmul exports input tensors to the CPU.
CPU accepts auto, cpu, reference; CUDA accepts auto, custom. Metal
accepts custom and, when enabled, optimized (MPSGraph); auto selects the
available path. MPSGraph drops common singleton batch axes, then requires at
most 16 graph dimensions including the two matrix axes. Metal auto selects
the custom kernel for input ranks above 16; the custom path has no rank cap.
Empty products and zero contracting dimensions also work with optimized.