← All verified jobs

Infrastructure/GPU Cluster/Platform Operations Lead

eleks Remote (Canada)

See all open roles at eleks

Ghost-risk verdict

Some ghost-posting signals

  • open for 44 days (30+ days starts to look stale)
  • 46 open roles at this company in 30 days (mass-hiring blitz)
  • no salary disclosed (correlates with ghost postings)

How we score ghost risk →

See your fit for this role and apply with a truthfully tailored résumé.

See my fit, free

About the role

REQUIREMENTS

8+ years of Infrastructure Engineering or Platform Operations experience

Experience managing GPU clusters for AI workloads

Strong Kubernetes administration skills

Experience with NVIDIA GPU technologies and CUDA ecosystem

Experience with cloud infrastructure (Azure, AWS or GCP)

Knowledge of storage, networking, and high-performance computing environments

Experience implementing Infrastructure as Code (Terraform or similar)

Strong operational leadership skills

Experience supporting AI platform infrastructure

Upper-Intermediate or higher level of English

RESPONSIBILITIES

Lead GPU infrastructure design and operations

Manage Kubernetes-based AI platform environments

Optimize infrastructure for AI training and inference workloads

Define operational standards and reliability practices

Collaborate with AI engineering teams

Implement monitoring, security, and disaster recovery strategies

Lead infrastructure capacity planning

Support technical roadmap and infrastructure evolution

Stop applying to ghosts.

OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.