FlashInfer Profiling

MiniCPM-SALA layer-level profiling at 64 concurrency — W4A16 GPTQ Marlin with FP8 KV Cache, CUDA graph disabled
3.88M input tokens 410K output tokens 64 requests 1,815.86s benchmark 125 prefill steps 30,055 decode steps