Projects across sparse inference, accelerator deployment, and architecture-aware optimization.
Working across cache architecture, sparse runtime paths, benchmarks, and reproducible systems documentation.
Combined profiling, ONNX graph simplification, TensorRT FP16, persistent memory, and CUDA Graph capture.
Exposed parallelism, removed cache conflicts, and validated the optimized benchmark under the standard stability test.