FlashInfer Profiling
MiniCPM-SALA layer-level profiling at 64 concurrency — W4A16 GPTQ Marlin with FP8 KV Cache, CUDA graph disabled
3.88M input tokens
410K output tokens
64 requests
1,815.86s benchmark
125 prefill steps
30,055 decode steps