ZilionX™ Transformer Engine

Architecture

The integrated RTL executes the principal blocks of a Transformer layer. A new memory fabric is being developed as Stage 3M to coordinate data delivery to the compute engines.

Layer-level view

Input and residual dataAttention and projectionNormalization and residualFeed-forward pathLayer output

Conceptual view of the integrated RTL. It is not a claim about physical placement or a production data path.

Compute and control

Attention

Fixed-point arithmetic includes a softmax path, integrated within the Transformer RTL.

Feed-forward

Linear operations and GELU activation form the measured FFN path in the current baseline.

Normalization and residual

Layer normalization and residual transfers are included in the integrated layer benchmark.

Scheduling and resilience

Control coordinates the engines. Fault detection and recovery are design goals requiring dedicated verification before performance claims can be made.

Stage 3M memory fabric · in development

The memory fabric sits alongside and serves the compute engines, under the layer scheduler's control. Its purpose is to improve data availability and reduce avoidable waiting during Transformer execution. Banking, local tiles and overlap between loading and computation are being evaluated; a completed fabric and measured gains are not yet claimed.

External and banked storageStage 3M fabric and local working dataTransformer compute engines

Next measurements include memory transactions, bank conflicts, compute stalls and resource use, compared with the baseline in Benchmarks.

Workload direction

The first target is compact language Transformers. Vision Transformers are a planned extension after the current integration and memory work. No end-to-end Tiny LLM or ViT throughput claim has been established by the present test configuration.