Advanced Versatile Inference Silicon

ZilionX™ Transformer Engine

Developing an inference accelerator for Transformer workloads on constrained edge systems. The current RTL integrates attention, feed-forward and normalization paths. The Stage 3M memory fabric is implemented; Stage 3N is integrating it with the Transformer RTL and benchmarking the result.

Current status · September 2026
Integrated RTL and an eight-layer simulation benchmark are available. Stage 3M memory fabric implementation is complete. Stage 3N integration and benchmarking are in progress. Published cycle counts are the pre-integration baseline; FPGA evaluation is planned.

Engineering focus

Transformer datapath

Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.

Data movement

Integrating the Stage 3M memory fabric with the RTL, informed by cycle-level baseline measurements. Memory traffic and energy results are still to be measured.

Evaluation path

Cycle-level simulation today; FPGA evaluation and representative model workloads are future milestones.

See the evidence

Architecture overview explains the blocks and intended memory integration. Benchmarks provides the measured cycle counts and test scope. Progress separates completed work from active and planned work.