Senior Site Reliability Engineer (100% remote)
⚲ Warszawa
Do uzgodnienia
Wymagania
- Kubernetes
- AWS
- Monitoring Tools
- Terraform
- Kafka
Opis stanowiska
About Playson
Playson is a product-led company and one of the leading B2B suppliers in the iGaming industry, creating online slot games and technology used by partners across more than 30 regulated markets. Our products are powered by a high-traffic, high-load platform built to operate at scale, giving our teams the opportunity to solve complex technical challenges while delivering reliable experiences to millions of players.
About the Role
We’re looking for a Senior SRE / DevOps Engineer to join our Platform Tribe - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.
You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale. If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.
Key Responsibilities
• Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time
• Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment
• Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence
• Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem
• Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)
• Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues
• Maintain and evolve CI/CD pipelines and infrastructure-as-code practices
• Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment
• Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance
• Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load
Requirements
• Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments
• Experience with GitOps tools such as FluxCD or ArgoCD
• Proven experience in incident response, root cause analysis, and postmortems in production systems
• Solid experience with AWS, Terraform, Docker, and CI/CD pipelines
• Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch
• Strong understanding of networking concepts and protocols
• Proficiency in at least one scripting language (e.g. Python, Go, Node.js)
• Experience working with version control systems (Git)
• Familiarity with incident management tools like PagerDuty, Opsgenie, or similar
• Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability
• Proactive, resilient mindset with a focus on continuous improvement and system stability
What We Offer
• Competitive Salary
• Quarterly Bonuses
• Unlimited Paid Time Off
• Unlimited Paid Sick Leave
• Remote & Flexible Working
• Private Medical Insurance
• Financial Support for Life Events
• Professional Development Budget
• International Exposure
• Regular Company Events
*Benefits may vary depending on location and contractual agreement
Recruitment Process
1. HR Interview (30-45 min)
2. Technical interview (90 min)
4. Final Interview with C-level (60 min)
Playson is a product-led company and one of the leading B2B suppliers in the iGaming industry, creating online slot games and technology used by partners across more than 30 regulated markets. Our products are powered by a high-traffic, high-load platform built to operate at scale, giving our teams the opportunity to solve complex technical challenges while delivering reliable experiences to millions of players.
About the Role
We’re looking for a Senior SRE / DevOps Engineer to join our Platform Tribe - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.
You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale. If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.
Key Responsibilities
• Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time
• Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment
• Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence
• Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem
• Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)
• Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues
• Maintain and evolve CI/CD pipelines and infrastructure-as-code practices
• Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment
• Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance
• Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load
Requirements
• Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments
• Experience with GitOps tools such as FluxCD or ArgoCD
• Proven experience in incident response, root cause analysis, and postmortems in production systems
• Solid experience with AWS, Terraform, Docker, and CI/CD pipelines
• Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch
• Strong understanding of networking concepts and protocols
• Proficiency in at least one scripting language (e.g. Python, Go, Node.js)
• Experience working with version control systems (Git)
• Familiarity with incident management tools like PagerDuty, Opsgenie, or similar
• Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability
• Proactive, resilient mindset with a focus on continuous improvement and system stability
What We Offer
• Competitive Salary
• Quarterly Bonuses
• Unlimited Paid Time Off
• Unlimited Paid Sick Leave
• Remote & Flexible Working
• Private Medical Insurance
• Financial Support for Life Events
• Professional Development Budget
• International Exposure
• Regular Company Events
*Benefits may vary depending on location and contractual agreement
Recruitment Process
1. HR Interview (30-45 min)
2. Technical interview (90 min)
4. Final Interview with C-level (60 min)
🔍 Dekoder Ogłoszenia
🔴
lean & senior team where ownership is high and expectations are even higher
Spodziewaj się pracy w małym zespole z dużą odpowiedzialnością i wysokimi wymaganiami, co może oznaczać presję i konieczność samodzielnego rozwiązywania problemów.
🟡
deeply hands-on role at the core of a high-traffic system
Będziesz musiał aktywnie angażować się w rozwiązywanie problemów technicznych i utrzymanie systemu, a nie tylko nadzorować.
🔴
real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation
Oczekuj częstych sytuacji kryzysowych, konieczności szybkiego reagowania na awarie i pracy w systemie dyżurów, który może obejmować nocne i weekendowe interwencje.
🔴
resilience, strong decision-making under pressure
Praca będzie wymagała odporności psychicznej i umiejętności podejmowania szybkich decyzji w stresujących sytuacjach.
🟡
continuously improve systems operating at scale
Będziesz odpowiedzialny za ciągłe optymalizowanie i rozwijanie istniejących systemów, co może oznaczać nieustanne zmiany i adaptację.