Selected work

Projects

Selected systems from the latest CV, written with concrete problem, architecture, and impact rather than abstract positioning.

P01
Agent InfrastructureProduction / Awarded

Coding Agent and Kernel Agent Infrastructure

Built coding-agent infrastructure that improved IPD efficiency by more than 100%, with a kernel agent achieving first place in the 2026 MLSys Challenge.

Problem
Engineering delivery and kernel optimization are bottlenecks when AI systems move from research prototypes to product-grade infrastructure.
Solution
Built agent infrastructure for coding workflows and high-performance kernel generation.
Architecture
Agent harnesses coordinate code generation, optimization, validation, and workflow integration for IPD and kernel-level tasks.
Impact
Improved IPD efficiency by more than 100%; the kernel agent achieved first place in the 2026 MLSys Challenge.
Coding AgentsKernel GenerationMLsys ChallengeDeveloper Productivity
P02
AI InfrastructureBest Practice

Inference Acceleration and Transformer Hardware-Software Co-design

Built inference infrastructure and co-designed hardware-software paths for Transformer hybrid attention and large MoE models on edge devices.

Problem
Hundred-billion-parameter inference is constrained by prefill/decode latency, kernel efficiency, and hardware memory/compute limits.
Solution
Combined quantization, pruning, kernel optimization, high-performance kernel generation, and PD-disaggregated inference.
Architecture
A co-designed stack spanning model structure, serving architecture, kernels, and edge-device hardware constraints.
Impact
Improved TTFT and TPOT by more than 50% in best practice and supported Transformer hybrid attention and large MoE deployment paths.
PD-disaggregated InferenceQuantizationPruningKernel OptimizationMoE
P03
Foundation ModelsDeployed

Hundred-Billion-Parameter LLM Training and Deployment

Led a 20–30 person team training 100B+ parameter large language models that ranked No. 1 on C-Eval in September 2023 and No. 1 on MMBench in April 2024.

Problem
Public-sector deployments need strong Chinese-language reasoning, multimodal understanding, and reliable task adaptation, not just generic model demos.
Solution
Led post-training, evaluation, and deployment of hundred-billion-parameter models for public-sector customers.
Architecture
Large-model training and post-training pipeline connected to evaluation, customer deployment, and downstream public-service applications.
Impact
Models ranked No. 1 on C-Eval in September 2023 and No. 1 on MMBench in April 2024, then deployed for public-sector customers.
100B+ LLMsSFT / RLC-EvalMMBench