Transformer datapath
Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.
Advanced Versatile Inference Silicon
Developing an inference accelerator for Transformer workloads on constrained edge systems. The current RTL integrates attention, feed-forward and normalization paths. Stage 3M work is building a memory fabric to improve how data moves between storage and compute.
Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.
Developing a memory fabric informed by cycle-level baseline measurements. Memory traffic and energy results are still to be measured.
Cycle-level simulation today; FPGA evaluation and representative model workloads are future milestones.
Architecture overview explains the blocks and intended memory integration. Benchmarks provides the measured cycle counts and test scope. Progress separates completed work from active and planned work.