Job Family: Operations | Function: Information Systems / Computer Operations | Location: Durban | Salary: Rneg up to R45k p/m (depending on experience)
About the Role
We're looking for an Engineer: Global Operations to join our 24/7 production operations team, supporting global iGaming platforms. In this role, you'll ensure system stability, reliability, and availability through incident management, root cause analysis, and technical problem-solving — balancing reactive incident response with proactive improvement.
You'll play a key part in leading post-incident reviews, supporting operational readiness assessments, validating AI-assisted operational decisions, and mentoring junior colleagues to build team capability.
What You'll Do
- Participate in 24/7 shift rotations, managing incident response and triage — owning alert acknowledgement, prioritisation, and investigation decisions, resolving issues within established procedures, and escalating to specialists only when needed.
- Conduct root cause analysis (RCA) investigations, determine technical solutions, collaborate with engineering teams on complex problems, and take accountability for resolution outcomes.
- Coordinate and facilitate severity-tiered post-incident reviews, extracting lessons learned and driving organisational learning.
- Own and maintain incident knowledge and documentation — knowledge base structure, standards, and accessibility for the team.
- Mentor and coach junior team members, transferring technical knowledge and developing team proficiency in investigation and response procedures.
- Identify operational bottlenecks and inefficiencies, propose and implement workflow refinements, reduce alert noise, and optimise incident response procedures.
- Support AI-assisted operational decision validation — establishing validation processes, escalation criteria, curating incident data for agent training, and optimising human-in-the-loop model performance.
- Assess operational readiness for new features and system changes, including SLI/SLO implications and operational risk identification.
- Evaluate and recommend operational tools and platforms, assessing supportability, reliability, and integration options.
- Drive automation and toil-reduction initiatives — identifying repetitive manual tasks and implementing efficiency improvements.
What We're Looking For
Education
- Advanced Diploma or Bachelor's Degree in Information Technology, Computer Science, Computer Engineering, Information Systems, Cybersecurity or an equivalent technical field.
Experience
- 2–4 years of progressive experience in IT operations, technical support, cybersecurity or systems administration, with demonstrated capability in incident investigation, troubleshooting, and process improvement.
Skills
- Incident Management, Troubleshooting, Root Cause Analysis, System Administration, Technical Documentation, Problem-Solving (Developing–Intermediate)
- Communication Skills (Intermediate–Advanced)
- Process Improvement, Knowledge Management, Escalation Management (Developing–Intermediate)
- Automation, Agile Methodology, Quality Assurance, Change Management, Influencing Skills, ITIL 4 (Awareness–Developing)
Knowledge
- Broad knowledge across incident management, post-incident learning, process improvement, emerging operational technologies, and operational readiness assessment
- Advanced understanding of incident response practices — alert triage, investigation methodology, root cause analysis, escalation procedures, and documentation standards
- Demonstrated technical knowledge of production systems, application architecture, and infrastructure components
- Developing proficiency with AI-assisted operational tools and automation opportunities
- Working knowledge of operational readiness assessment, SLOs, and quality standards
Why Join Us
- Be part of a global operations team supporting platforms that operate across multiple countries and continents.
- Work at the intersection of traditional operations and AI-assisted automation, helping shape how human-in-the-loop processes evolve.
- Grow your technical and leadership skills through mentoring, cross-functional collaboration with engineering teams, and exposure to a broad range of operational challenges.
- Contribute directly to platform reliability and customer trust in a regulated, high-availability environment.
This is a full-time role requiring participation in a 24/7 shift rotation.