Site Reliability Engineer (Global Support)

SA - Durban

20 Aug 2026

SA - Durban

Full Time

1

3-20 years

Job Family: Operations | Function: Information Systems / Computer Operations | Location: Durban | Salary: Rneg up to R45k p/m (depending on experience)


About the Role
We're looking for an Engineer: Global Operations to join our 24/7 production operations team, supporting global iGaming platforms. In this role, you'll ensure system stability, reliability, and availability through incident management, root cause analysis, and technical problem-solving — balancing reactive incident response with proactive improvement.
You'll play a key part in leading post-incident reviews, supporting operational readiness assessments, validating AI-assisted operational decisions, and mentoring junior colleagues to build team capability.


What You'll Do

  • Participate in 24/7 shift rotations, managing incident response and triage — owning alert acknowledgement, prioritisation, and investigation decisions, resolving issues within established procedures, and escalating to specialists only when needed.
  • Conduct root cause analysis (RCA) investigations, determine technical solutions, collaborate with engineering teams on complex problems, and take accountability for resolution outcomes.
  • Coordinate and facilitate severity-tiered post-incident reviews, extracting lessons learned and driving organisational learning.
  • Own and maintain incident knowledge and documentation — knowledge base structure, standards, and accessibility for the team.
  • Mentor and coach junior team members, transferring technical knowledge and developing team proficiency in investigation and response procedures.
  • Identify operational bottlenecks and inefficiencies, propose and implement workflow refinements, reduce alert noise, and optimise incident response procedures.
  • Support AI-assisted operational decision validation — establishing validation processes, escalation criteria, curating incident data for agent training, and optimising human-in-the-loop model performance.
  • Assess operational readiness for new features and system changes, including SLI/SLO implications and operational risk identification.
  • Evaluate and recommend operational tools and platforms, assessing supportability, reliability, and integration options.
  • Drive automation and toil-reduction initiatives — identifying repetitive manual tasks and implementing efficiency improvements.

What We're Looking For
Education

  • Advanced Diploma or Bachelor's Degree in Information Technology, Computer Science, Computer Engineering, Information Systems, Cybersecurity or an equivalent technical field.

Experience

  • 2–4 years of progressive experience in IT operations, technical support, cybersecurity or systems administration, with demonstrated capability in incident investigation, troubleshooting, and process improvement.

Skills

  • Incident Management, Troubleshooting, Root Cause Analysis, System Administration, Technical Documentation, Problem-Solving (Developing–Intermediate)
  • Communication Skills (Intermediate–Advanced)
  • Process Improvement, Knowledge Management, Escalation Management (Developing–Intermediate)
  • Automation, Agile Methodology, Quality Assurance, Change Management, Influencing Skills, ITIL 4 (Awareness–Developing)

Knowledge

  • Broad knowledge across incident management, post-incident learning, process improvement, emerging operational technologies, and operational readiness assessment
  • Advanced understanding of incident response practices — alert triage, investigation methodology, root cause analysis, escalation procedures, and documentation standards
  • Demonstrated technical knowledge of production systems, application architecture, and infrastructure components
  • Developing proficiency with AI-assisted operational tools and automation opportunities
  • Working knowledge of operational readiness assessment, SLOs, and quality standards

Why Join Us

  • Be part of a global operations team supporting platforms that operate across multiple countries and continents.
  • Work at the intersection of traditional operations and AI-assisted automation, helping shape how human-in-the-loop processes evolve.
  • Grow your technical and leadership skills through mentoring, cross-functional collaboration with engineering teams, and exposure to a broad range of operational challenges.
  • Contribute directly to platform reliability and customer trust in a regulated, high-availability environment.

This is a full-time role requiring participation in a 24/7 shift rotation.