oldMoney SOAR Notes

Single Layer Transformer with KV Cache
Source Code and Architecture Explanation
SALA Forward Process
Sparse Attention and Linear Attention Forward
Compilation & Submission Pipeline
sgl-kernel build process: .cu → .o → .so → wheel → tar.gz
MiniCPM-SALA Model Architecture
9.48B model panorama: Lightning layers, MiniCPM4 layers, parameter breakdown
Compute Analysis — RTX 5090
Per-operation timing analysis: memory bandwidth, FLOP/Byte, decode bottlenecks
Fine-Grained Structure
Latest structure blue print
0403 Meething
0403 Meething
Flashinfer 64 Concurrency Profiling
Flashinfer 64 Concurrency Profiling