Transformer datapath
Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.
Advanced Versatile Inference Silicon
Developing an inference accelerator for Transformer workloads on constrained edge systems. The current RTL integrates attention, feed-forward and normalization paths. The Stage 3M memory fabric is implemented; Stage 3N is integrating it with the Transformer RTL and benchmarking the result.
Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.
Integrating the Stage 3M memory fabric with the RTL, informed by cycle-level baseline measurements. Memory traffic and energy results are still to be measured.
Cycle-level simulation today; FPGA evaluation and representative model workloads are future milestones.
Architecture overview explains the blocks and intended memory integration. Benchmarks provides the measured cycle counts and test scope. Progress separates completed work from active and planned work.