13 Aug 2026
Senior Machine Learning Engineer - LLM Quantization & Deployment
$175K – $296K
Global
Full Time
/
Senior
$175K – $296K
·
Global
/
Senior
XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.
Our mission is to build strong foundation for LLM deployment and quality sign-off for next-gen XPENG Turing AI chip. This includes and is not limited to: LLM model fine tuning, PTQ, QAT, on-vehicle inference and related fields.
Key Responsibilities
- Develop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods, including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.
- Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling.
- Build robust model export, calibration, benchmarking, validation, and deployment pipelines.
- Engage early with the VLA model research team to establish performance estimates and prove model feasibility.
- Curate evaluation datasets and establish a comprehensive metric suite to systematically benchmark VLA performance.
- Analyze numerical errors, accuracy regressions, and performance trade-offs.
- Develop PTQ and QAT orchestration workflows.
- Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off.
- Collaborate with the in-vehicle software team on latency analysis and issue triage.
- Collaborate with the training infrastructure team to develop QAT and model distillation.
Basic Qualifications
- Master in CS/CE/EE, or equivalent, with 1-3 years of industry experience. Open to new graduates.
- Strong understanding of Transformer architectures and LLM inference.
- Hands-on experience quantizing or deploying deep learning models in production.
- Proficiency with PyTorch and at least one inference or compilation stack.
- Strong Python programming and software engineering skills.
- Ability to work effectively across research, systems, infrastructure, and product teams.
- Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.
Preferred Qualifications
- Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization.
- Experience with AWQ, GPTQ, SmoothQuant, or related methods.
- Strong numerical analysis and systems engineering skills.
- Experience with one or more LLM runtimes, such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes.
- Experience deploying LLMs on resource-constrained or heterogeneous hardware.
- Contributions to model optimization, inference, compiler, or serving projects.
- Publications at NeurIPS, ICML, ICLR, ACL, or related conferences.
What We Provide
- A fun, supportive and engaging environment.
- Infrastructures and computational resources to support your work.
- Opportunity to work on cutting edge technologies with the top talents in the field.
- Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.
- Competitive compensation package.
- Snacks, lunches, dinners, and fun activities.
Skills
Benefits
Frequently asked questions
What is the salary for Senior Machine Learning Engineer - LLM Quantization & Deployment?
This position pays $175K – $296K.
Is the Senior Machine Learning Engineer - LLM Quantization & Deployment position remote?
No, this position is on-site.
What experience level is required for Senior Machine Learning Engineer - LLM Quantization & Deployment?
This role is at the Senior level.
What is the application deadline for Senior Machine Learning Engineer - LLM Quantization & Deployment?
Applications close on September 12, 2026.
AI ML Engineer Salary
The average yearly salary for a ML Engineer is $221K per year, with a minimum base salary of $90K and a maximum of $372K.
ML Engineer Jobs in AI
S
$220K – $300K
1D
XM
$215K – $364K
3D
2
3D
O
4D
XM
$175K – $296K
4D
R
$266K – $372K
5D