Tech Lead – NOC
Likely real
- 22 open roles at this company in 30 days (mass-hiring blitz)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
Role Overview
The Tech Lead – NOC / Infrastructure Engineering will drive the reliability, scalability, and
operational excellence of critical systems and infrastructure. This role blends deep technical
expertise with leadership responsibilities, including guiding a team of engineers, defining best
practices, and leading automation and transformation initiatives.
The ideal candidate will act as a technical authority across cloud (AWS), networking, telephony
(VOIP/GSM), monitoring, and incident management domains, while championing Infrastructure as
Code (IaC), AI-driven operations, and continuous improvement.
Key Responsibilities
Technical Leadership & Strategy
Act as the primary technical authority for NOC operations and infrastructure engineering.
Define architecture standards, operational best practices, and automation strategies.
Drive adoption of SRE principles, including reliability engineering, error budgets, and
capacity planning.
Lead technical decision-making and design reviews across infrastructure and operations.
Infrastructure & Automation
Architect, design, and implement infrastructure using Terraform (IaC).
Lead automation initiatives to eliminate manual tasks and improve operational efficiency.
Develop reusable modules, pipelines, and frameworks for scalable infrastructure
deployment.
Promote CI/CD practices for infrastructure and operational workflows
Cloud & Platform Engineering
Own and optimize AWS cloud environments (compute, networking, storage, security).
Ensure cost efficiency, high availability, and scalability of cloud resources.
Drive cloud governance, security best practices, and compliance adherence.
Networking & Telephony Systems
Oversee design and operations of networking infrastructure (routing, switching, VPNs,
firewalls).
Lead management of VOIP systems, SIP trunks, PSTN, and GSM gateways.
Troubleshoot complex telecommunication and network issues across distributed systems
Monitoring, Incident & Reliability Management
Define and implement monitoring and observability strategy (Prometheus, Grafana, Datadog,
etc.).
Lead major incident response (P1/P2), acting as Incident Commander when required.
Drive proactive issue detection, alert tuning, and noise reduction.
Ensure high-quality Root Cause Analysis (RCA) and track corrective/preventive actions.
Change Management & Operational Governance
Ensure adherence to ITIL processes for incident, problem, and change management.
Review and approve high-risk changes and deployments.
Continuously improve operational playbooks and SOPs.
AI & Continuous Improvement
Identify and implement AI/ML-driven solutions for predictive monitoring and automation.
Drive innovation initiatives to improve operational maturity and reduce MTTR.
Evaluate new tools and technologies for operational enhancement.
Reporting & Stakeholder Communication
Provide leadership-level reporting on system health, incidents, SLAs, and KPIs.
Build dashboards and insights using Excel, Power BI, or similar tools.
Collaborate with cross-functional teams including Product, Engineering and Support.
Required Skills & Qualifications
Technical Expertise
Strong hands-on experience with Terraform, AWS, and infrastructure automation.
Deep understanding of networking protocols (TCP/IP, VLANs, VPNs, DNS, etc.).
Expertise in VOIP, SIP, PSTN, and GSM gateway systems.
Experience with monitoring/observability tools (Prometheus, Grafana, Datadog, etc.).
Working knowledge of AI/ML use cases in operations (AIOps).
Familiarity with ticketing systems (ServiceNow, Jira).
Working knowledge of MSSQL, Oracle, PostgreSQL, MySQL, and related database
technologies, including monitoring, performance analysis, troubleshooting, upgrades,
patching, and failover activities.
Experience supporting highly available database environments, including backup/recovery
validation, replication monitoring, upgrade coordination, failover testing, and incident
management in collaboration with DBA teams.
Process & Leadership
Strong expertise in ITIL processes (Incident, Change, Problem Management).
Proven experience in leading technical teams and handling escalations.
Ability to design and enforce operational standards and governance.
Soft Skills
Excellent problem-solving and analytical thinking.
Strong communication and stakeholder management skills.
Ability to lead under pressure during critical incidents.
Experience
6–10 years in NOC, SRE, or infrastructure roles.
Minimum 2–3 years in a technical lead / senior leadership role.
Preferred Qualifications
AWS Certified Solutions Architect / SysOps Administrator
HashiCorp Terraform Associate
ITIL Foundation Certification
Networking certifications (CCNA or equivalent)
VOIP/SIP certifications
Experience with Python/Bash scripting
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot