LLM 推理性能决策基线:TTFT/TPOT、KV Cache、吞吐与解码策略对照
-
Updated
Aug 12, 2026 - Python
LLM 推理性能决策基线:TTFT/TPOT、KV Cache、吞吐与解码策略对照
Analytical simulator for SRMIC — a residency-first LLM inference accelerator architecture using distributed on-package SRAM (HRM) to reduce HBM pressure on the decode critical path.
Add a description, image, and links to the decode-acceleration topic page so that developers can more easily learn about it.
To associate your repository with the decode-acceleration topic, visit your repo's landing page and select "manage topics."