Benchmark of token streaming flush policies for LLM serving, measuring user-perceived interactivity, chunk readability, and infrastructure overhead through TTI, TTRC, SPIS, and SPIS-R metrics.
python performance-engineering benchmark streaming simulation latency inference sse systems throughput user-experience serving event-driven-simulation spis ttft llm token-streaming spisflush-policy
-
Updated
Jul 25, 2026 - Python