Senior Site Reliability Engineer
Description
About the Role We’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment. You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale. If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you. Key Responsibilities Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD) Drive automation and proactively improve system resilience, reducing manual inte --- Source: jobicy - Playson