Infrastructure/GPU Cluster/Platform Operations Lead
Some ghost-posting signals
- open for 44 days (30+ days starts to look stale)
- 46 open roles at this company in 30 days (mass-hiring blitz)
- no salary disclosed (correlates with ghost postings)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
REQUIREMENTS
8+ years of Infrastructure Engineering or Platform Operations experience
Experience managing GPU clusters for AI workloads
Strong Kubernetes administration skills
Experience with NVIDIA GPU technologies and CUDA ecosystem
Experience with cloud infrastructure (Azure, AWS or GCP)
Knowledge of storage, networking, and high-performance computing environments
Experience implementing Infrastructure as Code (Terraform or similar)
Strong operational leadership skills
Experience supporting AI platform infrastructure
Upper-Intermediate or higher level of English
RESPONSIBILITIES
Lead GPU infrastructure design and operations
Manage Kubernetes-based AI platform environments
Optimize infrastructure for AI training and inference workloads
Define operational standards and reliability practices
Collaborate with AI engineering teams
Implement monitoring, security, and disaster recovery strategies
Lead infrastructure capacity planning
Support technical roadmap and infrastructure evolution
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot