CV
Education, experience, awards, skills, and patents.
Professional Summary
Machine learning systems researcher focusing on efficient training and inference for large-scale language models. My work studies system-level optimizations including MFU improvement, communication–computation overlap, MoE inference optimization, and hardware-efficient execution for large-scale AI infrastructure.
Currently AI Infrastructure / Training Systems Researcher at AIGCode (蔻町科技). Previously Senior Algorithm Researcher and TL at Huawei Cloud Architecture Innovation Lab.
Publications
-
2026 EcoThink: A Green Adaptive Inference Framework for Sustainable and Accessible Agents
The Web Conference (WWW 2026), Oral
Linxiao Li*, Zhixiang Lu*†
-
2026 Deterministic Component Mining for Multi-framework UI2Code Generation
International Conference on Machine Learning (ICML 2026), Poster
Zixiong Yang*, Linxiao Li*, Jiaye Lin, Binrui Wu, Xiaoyu Kang, Jiechao Gao†
-
2026 Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning
arXiv:2606.14187, Preprint
Kaiwen Chen*, Shuhai Zhang*, Zimo Liu, Linxiao Li, Ying Sun, Yuchen Li, Yifan Zhang, Bo Han, Mingkui Tan†, Qiuwu Chen†
-
2026 Toward Executable Multi-framework Front-end Code Generation with Self-Correction
Advances in Neural Information Processing Systems (NeurIPS 2026), Poster
Linxiao Li*, Haoran Ma*, Chenyue Wang, Jiaye Lin, Haochen Sui, Jiechao Gao†
-
2026 Phase-wise MLLM Tuning for Multi-framework WebUI Code Generation
Advances in Neural Information Processing Systems (NeurIPS 2026), Poster
Haoran Ma*, Linxiao Li*, Chenyue Wang, Haochen Sui, Jiechao Gao†
-
2026 Tree Rotary Positional Encoding for Extreme Length Extrapolation from Scratch
Advances in Neural Information Processing Systems (NeurIPS 2026), Poster
Qiuwu Chen, Ziteng Huang, Shuhai Zhang, Zimo Liu, Linxiao Li, Ying Sun, Yuchen Li, Yifan Zhang, Yaofo Chen, Mingkui Tan†
Public Talks
-
Jun 2026 Communication-Aware Scheduling Optimization for Large-Scale MoE Model Training on Ascend NPUs
Frontiers Series · Guangming Laboratory, Academic Seminar
Academic seminar on communication-aware scheduling optimization for large-scale MoE model training on Ascend NPUs.
-
Apr 2026 DeepSeek-V4 Mechanism Deep Dive
Internal Technical Sharing · AIGCode, Technical Talk
Company technical interpretation and discussion of DeepSeek-V4 mechanism details.
-
Dec 2025 Research Methodology, Internship Planning, and Career Development
Graduate Career Workshop · Guangming Laboratory, Invited Talk
Invited talk for graduate students on research methodology, internship planning, and career development.
Industry Research Experience
-
2025 - Shanghai, China
AI Infrastructure / Training Systems Researcher
AIGCode (蔻町科技)
Leading LLM pre-training optimization and training infrastructure. Currently leading pre-training of a 300B+ parameter LLM from scratch.
- Achieved 5x single-operator speedup (MoE Combine + FFN), improving overall MFU by 10% and reducing training cost by >10%
- Tech stack: Triton, TileLang, PyPTO with Tile paradigm programming
- Profiling tools for end-to-end optimization. Building next-generation AI Infra and model architectures.
-
2023 - 2025 Shenzhen, China
Senior Algorithm Researcher Leader / PL (Huawei Fellow Team)
Huawei Technologies Co., Ltd.
TL at Huawei Cloud Architecture Innovation Lab (Cloud Lab), under CTO Fellow Jiongjiong Gu (Chief Architect). Led a 10+ person joint team with the Guangming Laboratory on intelligent resource scheduling. All annual performance ratings A.
- MoE Memory Offloading: Expert offloading for MoE LLM inference on Ascend NPUs. TTFT <3s, TPOT <100ms. Supports BF16/W8A8 (DeepSeek V3). (HDC 2025, ~50 min)
- NPU User-Space Virtualization: Multi-tenant NPU sharing with <3% overhead, utilization 10%→30%+. (HDC 2025, ~50 min)
- Huawei Cloud Flexus FRCP: iTransformer capacity prediction for X/L instances, deployed at scale.
- QoS MoE Model: MAE 3%, classification accuracy 99.995%. Key support for Agricultural Bank of China HCS project.
-
2021 - 2023 Chengdu, China
Software Development Leader
Huawei Technologies Co., Ltd.
Led 5-person team as Committer. Core developer on Data Workshop distributed messaging system. (incl. internship from Sep 2021)
- Built on Pulsar distributed messaging middleware with self-developed storage engine
- Zero P3 incidents over 6 months; supported major customers at scale
Selected Research Projects
-
Ascend NPU User-Space Virtualization
Designed a user-space virtualization framework for Ascend NPUs to improve hardware utilization in large-scale AI training.
- Achieved <3% performance overhead while increasing NPU utilization from 10% to 30%
-
MoE Inference Memory Offloading
Designed a memory offloading framework for large-scale Mixture-of-Experts inference to reduce GPU memory pressure.
- Achieved <3% overhead while maintaining TTFT <3s and TPOT <100ms in production systems
- Huawei Cloud High-Potential Patent: A MoE Model-Oriented Inference Optimization Method
-
Both projects were deployed in Huawei Cloud and announced at Huawei HDC 2025 (~50 min mark).
Education
-
2018 - 2021 Sydney, Australia
The University of Sydney
Computer Science
- GRE: 334 (Q170, V164, AW 5.5)
- Advisor: Prof. Wei (Wilson) Li
-
2021 - 2021 Shanghai, China
-
2020 - 2021 Shanghai, China
Open Source
-
Codex AutoResearch
Creator · 2,651 Stars — GitHub. A self-directed iterative research system for Codex that continuously cycles through modify, verify, retain/discard, and repeat.
-
PyPTO
Contributor — GitCode. High-performance programming framework for AI accelerators using tile-based programming (PTO paradigm), compiling tensor graphs to hardware instructions via multi-level IR.
-
PyPTO-AutoGrad
Creator — GitCode. Source-code level automatic differentiation tool for PyPTO kernels, generating readable and editable backward kernels from forward definitions.
-
Modern LLM Notebook
Contributor · 215 Stars — GitHub, Website. A hands-on PyTorch curriculum for building modern LLM systems from scratch, covering tokenization, attention, MoE, inference, evaluation, and distillation.
-
Pie
Contributor · 217 Stars — GitHub, Website. A programmable LLM serving system for custom inference logic, stateful agents, and serving-side optimization.
Awards
-
2025 Huawei ICT Software Competition — 14th / 5200+
Huawei
Top 3 in Huawei Cloud product line, 1st in Chengdu Research Institute.
-
2024 Golden Cloud Award (金代码)
Huawei Cloud
Top 2% recognition for outstanding code contributions.
-
2024 HCS Cloud Summit Star (云巅之星)
Huawei Cloud
Top 5% internal technical honor for outstanding contributions.
-
2023 -
2023 Huawei Trustworthy Professional Certification
Huawei
Passed Huawei Trustworthy Professional Level Certification.
-
2022 -
2020 Vice Chancellor's Global Mobility Scholarship
The University of Sydney
Awarded on the basis of academic merit and exchange destination.
Patents
-
2026 Runtime Semantic Authorization for Shortcut Actions in Branch-Structured Large Language Model Programs
First inventor. Granted. CN122571589B.
-
2026 Expert Module Selection for Mixture-of-Experts Models
First inventor. Application No. 202611106026.3.
-
2026 Method for Evaluating Large Language Model Inference Interfaces, Terminal Device, and Storage Medium
First inventor. Application No. 202611242873.2.
-
2024 A MoE Model-Oriented Inference Optimization Method
Huawei Cloud High-Potential Patent
Invention patent for MoE large model inference memory offloading optimization.
-
2024 A Latency-Aware Serverless Request Scheduling Strategy
Invention patent for intelligent serverless request scheduling.
Academic Service
-
Teaching
Teaching Assistant, The University of Sydney — Operating Systems, Algorithm Design
-
Reviewing
Reviewer, NeurIPS 2026 and ICLR 2027