Advanced Versatile Inference Silicon

ZilionX™ Transformer Engine

Developing an inference accelerator for Transformer workloads on constrained edge systems. The current RTL integrates attention, feed-forward and normalization paths. Stage 3M work is building a memory fabric to improve how data moves between storage and compute.

Current status · September 2026
Integrated RTL and an eight-layer simulation benchmark are available. Stage 3M memory fabric development has started. FPGA evaluation is planned.

Engineering focus

Transformer datapath

Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.

Data movement

Developing a memory fabric informed by cycle-level baseline measurements. Memory traffic and energy results are still to be measured.

Evaluation path

Cycle-level simulation today; FPGA evaluation and representative model workloads are future milestones.

See the evidence

Architecture overview explains the blocks and intended memory integration. Benchmarks provides the measured cycle counts and test scope. Progress separates completed work from active and planned work.