AI Platform Engineer
Some ghost-posting signals
- open for 189 days (90+ without a fill is a strong ghost signal)
- 148 open roles at this company in 30 days (mass-hiring blitz)
- no salary disclosed (correlates with ghost postings)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
Key Responsibilities
Kubernetes for AI/ML: Architect and manage Kubernetes clusters tailored to AI/ML workloads.
GPU Orchestration: Implement Run:ai and operators for GPU resource orchestration and workload scheduling.
Automation & Pipelines: Develop and maintain Python-based automation scripts and ML pipelines; automate infrastructure provisioning with Terraform and configuration management with Ansible.
Notebooks & Collaboration: Create and manage Jupyter Notebooks for experimentation and collaboration.
NVIDIA Integration: Integrate and optimize NVIDIA Enterprise Suite components (CUDA, NeMo Framework, Triton, TensorRT, GPU drivers) for accelerated computing.
MLOps Practices: Establish and maintain MLOps best practices for model lifecycle management, CI/CD, and monitoring (e.g., MLflow, Kubeflow).
Collaboration: Work closely with data scientists and platform engineers to ensure efficient resource utilization and scalability across environments.
Required Skills & Experience
Strong proficiency in Python and experience with ML frameworks (TensorFlow, PyTorch).
Hands-on experience with Kubernetes and container orchestration.
Familiarity with Run:ai or similar GPU scheduling platforms.
Expertise in Terraform and Ansible for infrastructure automation.
Experience with Jupyter Notebooks for ML development.
Knowledge of NVIDIA Enterprise Suite (CUDA, NeMo Framework, Triton, GPU drivers).
Solid understanding of MLOps principles and tools (e.g., MLflow, Kubeflow).
Background in deploying and scaling AI workloads in cloud or hybrid environments.
Qualifications
4+ years in platform architecture or solutions architecture, with 2+ years focused on AI/ML workloads.
Experience with high-performance computing (HPC) environments.
Familiarity with distributed training and model optimization techniques.
Certification in Kubernetes or cloud platforms (AWS, Azure, GCP).
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot