Senior MLOps & Infrastructure Engineer
- VinFast
- Hanoi, Vietnam
- VND 1,800,000,000 – VND 3,000,000,000
About the Role
We are looking for Senior
MLOps
& Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast’s ADAS and autonomous driving
programme
-
on-premise
GPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.
What the Team Covers
You will contribute across
all of
the following over time. We do not expect one person to master every part on day one, but we do expect you to be willing to work in any of them.
GPU and HPC clusters running model training, large-scale inference and annotation workloads
MLOps
model serving, pipeline orchestration, experiment and model registry
Storage at multi-petabyte scale: tiering, lifecycle, backup and recovery
Ingestion of large volumes of recorded sensor data, with integrity verification and metadata extraction
CI/CD,
GitOps
, infrastructure-as-code, monitoring and observability
Access control, audit logging and data protection for internal and external users
Key Responsibilities
AI Compute & Serving
Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling,
utilisation
and capacity
Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
Manage multi-GPU distributed training and reproducible environment sandboxing
Scale pipeline orchestration on Kubernetes for large data processing jobs
Platform Automation & Delivery
Drive
GitOps
-based deployment and maintain infrastructure-as-code across the platform
Build and maintain CI pipelines, secure container builds and release automation
Build monitoring, logging and alerting so that failures are detected and actionable, never silent
Lead incident response and drive follow-up actions to closure
Data Infrastructure & Security
Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
Implement access control and single sign-on across platform services, with audit logging
Apply data protection measures to sensitive content before it reaches external users
Requirements
4+ years in MLOps
, DevOps, SRE or HPC platform engineering
, with production ownership of AI/ML infrastructure
Kubernetes
administration at production scale: Helm, ingress (
Traefik
or Envoy), CNI, and
GitOps
(Flux or
ArgoCD
)
GPU & HPC
SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
Strong Linux systems skills
and infrastructure-as-code (
SaltStack
or Ansible)
Python and Bash
for automation, including Airflow DAGs and custom operators
Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
Willingness to work across the full stack
- compute, storage, networking, automation and security - rather than within a single specialty
Good communication in English - technical documentation and working with international partners
Benefits
Competitive salary
Premium healthcare package, including PVI insurance & annual health check-ups
13th-month salary & performance bonuses to reward your contributions
Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool...
Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain
Work Location:
Technopark Tower, Gia Lam, Ha Noi
With respect to all your personal data shared to VinFast in the application and the entire recruitment process of VinFast, by clicking “Apply”, submitting your resumé/CV and/or participating in VinFast's recruitment process, you agree that you have read VinFast's Personal Data Protection Policy ("Policy") posted at https://vinfastauto.com/vn_vi/dieu-khoan-phap-ly or https://vinfast.vn/privacy-policy/ , you agree to the Policy and consent for VinFast to process your personal data in accordance with the Policy and the applicable regulations on personal data protection.
Skills
- Kubernetes
- GPU cluster management
- MLOps
- CI/CD
- Infrastructure as Code
- Python
- Observability








