2024 - Present
Harbin Institute of Technology, Shenzhen
B.Eng. candidate, Computer Science and Technology
Undergraduate student, ranked in the top 20% of the cohort. Shenzhen, China.
古权胜
I work where high-performance computing meets AI infrastructure. Across research and industry, I build systems for efficient inference.I work where high-performance computing meets AI infrastructure. Across research and industry, I build systems for efficient inference.
Current direction
AI infrastructure internships and systems research collaboration
2024 - Present
B.Eng. candidate, Computer Science and Technology
Undergraduate student, ranked in the top 20% of the cohort. Shenzhen, China.
May 2026 - Present
Research intern with the M³AIL Research Group at Harbin Institute of Technology, Shenzhen.
Oct 2025 - Present
Core member of the HITSZ supercomputing team, focusing on efficient inference for LLMs and world models.
2026
Team captain
First Prize
ASC26 Final
2026
Team captain
National Final, Third Prize
2024-2025
Individual
Outstanding Student and Academic Scholarship
Experience across five compute ecosystems, from cluster software to distributed inference and runtime optimization.
| Ecosystem | NVIDIA GPU | AMD GPU | Huawei NPU | Huawei CPU | Hygon DCU |
|---|---|---|---|---|---|
| Architecture | Hopper / Ampere | RDNA 3 | AI Core | Armv9 / Armv8.2-A | DCU |
| Models | H100 / H20 / A100 | Radeon PRO W7900D | Ascend 910C / 910B | Kunpeng 920F / 920 | BW1000 |
| Software stack | TensorRT / CUDA Graphs | ROCm / HIP | CANN / Ascend C / HCCL | HPCKit / BiSheng | DTK / DAS |
Cluster and distributed computing
Linux, Distributed systems, MPI, Slurm, Shell, Git
Distributed and efficient model inference
LLM inference, DeepEP, Mooncake, PyTorch, ONNX
Profiling and bottleneck analysis
Nsight Systems, PyTorch Profiler, perf, msprof
Deployment and execution optimization
TensorRT, CUDA Graphs, CANN

Cycling, running, hiking, travel, photography, and music make up the other side of this site.
Open activity