Skip to content

Tensor operations

Start with import tensorcx as cx. The following example works on CPU:

operations.py
import tensorcx as cx
a = cx.tensor([[1.0, 2.0], [3.0, 4.0]], device="cpu")
b = cx.ones((2, 2), dtype=cx.float32, device="cpu")
product = cx.matmul(a, b)
print(product.numpy())
# [[3. 3.]
# [7. 7.]]
print(cx.sum(a, axis=1).numpy())
# [3. 7.]
print(cx.softmax(a, axis=-1).shape)
# (2, 2)
print(((a - 1) / 2).reshape((4,)).numpy())
# [0. 0.5 1. 1.5]
print(a.sum(axis=1, keepdims=True).shape)
# (2, 1)
values = cx.tensor([[1, 2, 3], [4, 5, 6]])
bias = cx.tensor([0.25, 0.5, 0.75])
result = values.astype(cx.float32) + bias
print((result - result.mean(axis=-1, keepdims=True)).numpy())
# [[-1.25 0. 1.25]
# [-1.25 0. 1.25]]
print(values.T.numpy())
# [[1 4]
# [2 5]
# [3 6]]
expanded = values.expand_dims((0, -1))
print(expanded.shape) # (1, 2, 3, 1)
print(expanded.squeeze().shape) # (2, 3)
print(a.sum().numpy()) # 10.0
print(a.mean(axis=(0, 1), keepdims=True).numpy()) # [[2.5]]
print(values[:, ::-1].numpy())
# [[3 2 1]
# [6 5 4]]
print(cx.concat([values, values], axis=0).shape) # (4, 3)
print(cx.stack([values, values], axis=1).shape) # (2, 2, 3)
parts = values.split([1], axis=1)
print([part.shape for part in parts]) # [(2, 1), (2, 2)]
mask = values > 3
print(mask.dtype) # bool
print(values[mask].numpy()) # [4 5 6]
print(cx.where(mask, values, 0).numpy())
# [[0 0 0]
# [4 5 6]]
print(mask.any(axis=1).numpy()) # [False True]
# Leading batch dimensions broadcast; the last two axes form each matrix.
batch = cx.ones((2, 1, 3, 4))
weights = cx.ones((1, 5, 4, 6))
print((batch @ weights).shape) # (2, 5, 3, 6)
print((cx.tensor([1., 2., 3.]) @ cx.tensor([4., 5., 6.])).numpy()) # 32.0
Task API Contract
Construct tensor, empty, zeros, ones, randn Explicit shape, dtype, and device; empty contents are unspecified
Inspect Tensor.shape, .dtype, .device Metadata for a contiguous row-major tensor
Transfer .to(device), .cpu(), .numpy() Explicit device transfers; NumPy export copies to host
Elementwise a + b, a - b, a * b, a / b, -a Broadcast-compatible shapes, same dtype/device; real scalars work on either side; division requires float32
Convert dtype x.astype(dtype), cx.astype(x, dtype) Explicit float32/int32/bool conversion on the same device; copies by default
Reshape x.reshape(shape), cx.reshape(x, shape) Shares storage; same element count; one inferred -1 allowed
Transpose x.transpose(axes=None), cx.transpose(x, axes=None), x.T Reverses axes by default; always copies into a contiguous buffer on the same device
Singleton axes x.squeeze(axis=None), x.expand_dims(axis); also cx.squeeze, cx.expand_dims Remove/insert size-one axes as views sharing storage
Index / slice x[key] Integers, slices, negative steps, ellipsis and new axes; independent contiguous copies
Join / split cx.concat, cx.stack, cx.split, x.split Same dtype/device; split uses equal sections or NumPy-style cut indices; all outputs copy
Compare ==, !=, <, <=, >, >=; cx.equal, etc. Same dtype/device; broadcasting and matching scalars; bool output
Boolean logic &, &#124;, ^, ~; cx.logical_and, etc. Bool tensors/scalars; broadcasting
Conditional selection cx.where(condition, x, y) Bool condition; same-dtype branches; three-way broadcasting
Boolean reduction cx.any, cx.all; x.any, x.all Bool inputs; axes and keepdims like numeric reductions
Mask selection x[mask], cx.masked_select(x, mask) Bool mask matches leading dimensions; stable contiguous copy
Matrix multiply cx.matmul(a, b), a @ b Float32 vectors/matrices/batches; matching contracting dimensions; batch broadcasting
Reduce cx.sum(x, axis=None, keepdims=False), max, mean All axes by default; integer or axis sequence; keepdims=True retains selected axes with size 1
Activate cx.exp(x), gelu, silu Float32
Normalize softmax, rmsnorm, layernorm Float32, explicit axis; preserves shape

Support varies by backend, particularly for int32. See the support matrix.

  • Float inputs are narrowed to float32; precision loss and overflow are possible.
  • Broadcasting aligns axes from the right; dimensions must match or be 1. (2, 3) + (3,) returns (2, 3). A dimension of 0 paired with 1 stays 0. No dtype promotion or device transfer occurs; the result is contiguous.
  • astype truncates float-to-int toward zero and rejects NaN, infinity and out-of-range values. Int-to-float can lose precision. copy=False returns the input only when the dtype is unchanged; conversion still allocates.
  • CPU and Metal int32 add/subtract/multiply/negate wrap on overflow. CUDA int32 arithmetic is unsupported.
  • Scalars preserve the tensor dtype. Int32 requires an integer scalar within int32 bounds; float32 narrows real scalars before arithmetic. Booleans and complex scalars are rejected. Float division follows IEEE infinity/NaN rules.
  • Reshape keeps the underlying buffer alive. Experimental Metal kernels writing through a view also change other views sharing that storage. MLIR CPU/CUDA launches preserve caller tensors by copying their output. NumPy export still copies. Empty shape inference such as (0, -1) is ambiguous and rejected.
  • Reductions support negative and multiple axes, such as (0, -1); duplicates are errors. axis=None reduces all axes, while axis=() returns an independent same-shape copy under the same dtype restrictions. Multi-axis mean sums the selected group then divides once; no intermediate means are taken. Reducing a zero-size axis produces zeros for sum, NaNs for mean, and an error for max. NaNs propagate and float32 accumulation can overflow or lose precision.
  • Transpose accepts a full axis permutation, with negative indices; even an identity or scalar transpose owns a new buffer. Squeeze removes all size-one axes by default, or selected axes. Expand dims accepts one or more positions in the final output shape: (2, 3) at (0, -1) becomes (1, 2, 3, 1). Repeated/out-of-range axes and squeezing a non-singleton axis are errors. Both singleton operations share storage using the reshape ownership rules.
  • Softmax subtracts the maximum before exponentiation. NaNs propagate; slices containing positive infinity or only negative infinities produce NaNs.
  • RMSNorm and LayerNorm have no affine weights. Their eps must be finite, non-negative, and representable as float32.
  • Floating-point accumulation order can differ across backends. Compare with numerical tolerances, not bitwise equality.

Basic slicing supports negative indices/steps and clips slice bounds. Integer selection returns a tensor even for a scalar. Concat joins an existing axis; stack inserts a new one. Split returns a tuple of independent copies. Advanced integer indexing, mixed mask/index tuples, indexed assignment, autograd and general strided views are not implemented. Normalizations require one explicit axis. The experimental kernel subset has its own arithmetic rules and does not extend the ordinary tensor operators.

Boolean input infers cx.bool; mixed boolean/numeric input needs an explicit dtype. Numeric casts to bool use nonzero truth, including NaN and infinity; bool casts back to numbers as 0/1. Arithmetic on bool tensors is rejected. Comparisons use IEEE rules: NaN compares unequal to every value, including itself.

where selects already-computed branches; it is not lazy control flow. All tensor arguments must share a device. Mask indexing does not broadcast: a (2,) mask on (2, 3) selects rows, while a (2, 3) mask produces a flat vector. A scalar bool Tensor mask adds a leading axis of size 0 or 1. Selection copies preserve order and stored bits. GPU selection reads back only the selected count to size its output; it does not export the tensor to NumPy.

any/all accept bool inputs, all/single/multiple axes and keepdims; empty reductions return false/true respectively. Python bool(x) requires exactly one element and reads that value to the host; use x.any() or x.all() for a mask.

matmul broadcasts leading batch dimensions and multiplies the last two axes. For example, (2, 1, 3, 4) @ (1, 5, 4, 6) produces (2, 5, 3, 6). Rank-one inputs are promoted temporarily: (k,) @ (k,) returns a scalar Tensor, (m, k) @ (k,) returns (m,), and (k,) @ (..., k, n) returns (..., n). Scalar operands are rejected. Float32 dtype and device must match.

Empty batch/row/column dimensions yield empty outputs; a zero contracting dimension produces zeros. Outputs own independent contiguous storage. Neither broadcasting nor matmul exports input tensors to the CPU.

CPU accepts auto, cpu, reference; CUDA accepts auto, custom. Metal accepts custom and, when enabled, optimized (MPSGraph); auto selects the available path. MPSGraph drops common singleton batch axes, then requires at most 16 graph dimensions including the two matrix axes. Metal auto selects the custom kernel for input ranks above 16; the custom path has no rank cap. Empty products and zero contracting dimensions also work with optimized.