System Reliability Engineer (Data Centre)
See all open roles at centre for strategic infocomm technologies →
Some ghost-posting signals
- open for 41 days (30+ days starts to look stale)
- 68 open roles at this company in 30 days (mass-hiring blitz)
- no salary disclosed (correlates with ghost postings)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
Responsibilities
Oversee and manage IT operations within the data centre, including day-to-day monitoring, incident management, and problem management
Lead the end-to-end incident management lifecycle that encompass immediate troubleshooting, root cause identification, and resolution implementation to restore services, followed by comprehensive post-incident analysis
Develop and maintain documentation on IT infrastructure, operations, and procedures within the data centre
Perform capacity planning to ensure IT infrastructure is scalable for future demands
Collaborate and coordinate with Data Centre Facilities teams on matters related to power, cooling, and physical infrastructure
Design and implement robust observability platform alongside network monitoring tools for performance monitoring and real-time alerting of IT devices and networks
Implement and manage remote management tools for out-of-band access and control of IT devices and servers
Define, implement, and track SRE metrics, including SLO, SLI, and error budgets to improve data centre IT reliability
Requirements (Minimum Qualifications)
Background in Computer Science, Computer or Electrical Engineering, Information Technology or a related field
Good technical knowledge in IT infrastructure, including servers, storage, networking, and cloud technologies
Proficient in IT management software and tools
2 years of working experience in IT operations is preferred
Fresh graduates are welcomed to apply
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot