Close

Presentation

TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
DescriptionMultimodal stacks mixing ViTs, CNNs, GNNs, and transformer NLP strain embedded platforms due to heterogeneous compute and tight real-time constraints. We present TRINE, a single-bitstream FPGA accelerator and compiler that runs end-to-end multimodal inference without reconfiguration. It unifies layers as DDMM/SDDMM/SpMM on a mode-switchable PE array supporting weight/output-stationary systolic, 1×CS SIMD, and a routable adder tree with in-stream top-k token pruning. Dependency-aware layer offloading overlaps independent kernels across RPUs. On Alveo U50 and ZCU104, TRINE achieves up to 22.57× and 6.86× lower latency than RTX 4090 and Jetson Orin Nano at 20–21 W, with <2.5% accuracy loss and state-of-the-art efficiency.