The approach encodes neural computations using statistics of probabilistic bits and designs sparse operations around the physical coupling structure of Z1. Inference is divided across thermodynamic accelerators and conventional digital processors or FPGAs, assigning different computations to suitable hardware within a heterogeneous pipeline.
Z1T offers researchers a concrete architecture and scaling experiments for investigating hardware-software co-design. Its open sparse-transformer implementation is a research resource; the hardware efficiency discussion concerns the proposed Z1 deployment strategy and should not be read as a universally available inference service.

