Site Reliability Engineer II (SRE) - Guidewire Cloud Platform (Application)
⚲ Kraków, Rzeszów, Katowice, Kielce, Gliwice
12 500 - 15 000 PLN brutto (UoP)
Wymagania
- Python
- AWS
- Kubernetes
- Terraform
Opis stanowiska
At Guidewire, we deliver the software that Property and Casualty (P&C) insurance companies rely on to protect their customers during crises, natural disasters, accidents, and cyber risks. Our core applications enable insurers to sell and underwrite policies, settle claims, and bill their customers. We also offer a suite of innovative products for data management, digital portals, and predictive analytics.
Hundreds of insurers worldwide use Guidewire's products, running on our cutting-edge Guidewire Cloud Platform, to handle billions of dollars in business. We are dedicated to providing the tools and technology that help insurers protect and support their customers when they need it most.
The Opportunity
We are seeking a Site Reliability Engineer II who is eager to contribute to the transformation of the insurance industry with our leading cloud platform. As a member of the SRE-Application team, you'll play a critical role in ensuring the reliability, performance, and scalability of applications running on our Guidewire Cloud Platform. This position offers a unique opportunity to apply your skills in automation, software engineering, and operational discipline to support our cloud-based solutions.
What You'll Do:
• Assist in troubleshooting and resolving issues in collaboration with development teams, reducing customer impact.
• Develop and maintain automated runbooks to address common issues proactively.
• Apply engineering principles and basic automation to enhance our operating environments.
• Monitor applications and help improve their reliability and performance on the Guidewire Cloud Platform.
• Use your software engineering skills to optimize systems and reduce manual tasks.
• Document incidents and assist in refining processes to prevent future occurrences.
• Stay informed about industry trends, tools, and best practices in site reliability engineering.
• Contribute to a culture of innovation, learning, and continuous improvement.
• Participate in on-call rotations to ensure the availability and reliability of our services.
What You Will Bring:
• Experience as an SRE or similar role, focusing on improving system reliability.
• Strong problem-solving skills and the ability to assist in analyzing complex systems and devising effective solutions.
• Effective collaboration and communication skills to work cross-functionally and document processes.
• Experience with automation, monitoring, and performance optimization tools and techniques.
• Commitment to maximizing uptime, scalability, and delivering an exceptional end-user experience.
• Passion for technology and a desire to continuously learn and grow your skills.
• Alignment with Guidewire's mission to leverage technology to help protect and support others.
Required Skills:
• Basic understanding of SLI's, SLO's, and Error Budgets
• Familiarity with application performance monitoring (APM) and telemetry tools
• Some experience with troubleshooting and debugging distributed systems on cloud infrastructure
• Exposure to CICD pipelines within K8S or legacy ecosystems
• Familiarity with creating monitors, dashboards, and synthetic transactions in monitoring tools like Datadog
• Some experience deploying and managing infrastructure within AWS or Kubernetes ecosystems using Terraform or other cloud-native approaches
• Familiarity with infrastructure configuration management using tools such as GitOps, Puppet, or Ansible
• Basic understanding of AWS cloud networking and security
• Comfortable with Linux system administration and the ability to program/script using Python, Go, Java, shell, or equivalent
Preferred Skills:
• Pursuing or interested in pursuing SRE Certification
• Pursuing or interested in pursuing AWS Certification
• Familiarity with SQL, database administration, data pipelines, performance tuning, and schema design
• Exposure to pipelining tools such as Team City, Bitbucket Pipelines, Jenkins, or GitHub Actions
• Interest in learning about open-source distributed data processing frameworks such as Hadoop, Apache Spark, AWS RedShift, etc.
Hundreds of insurers worldwide use Guidewire's products, running on our cutting-edge Guidewire Cloud Platform, to handle billions of dollars in business. We are dedicated to providing the tools and technology that help insurers protect and support their customers when they need it most.
The Opportunity
We are seeking a Site Reliability Engineer II who is eager to contribute to the transformation of the insurance industry with our leading cloud platform. As a member of the SRE-Application team, you'll play a critical role in ensuring the reliability, performance, and scalability of applications running on our Guidewire Cloud Platform. This position offers a unique opportunity to apply your skills in automation, software engineering, and operational discipline to support our cloud-based solutions.
What You'll Do:
• Assist in troubleshooting and resolving issues in collaboration with development teams, reducing customer impact.
• Develop and maintain automated runbooks to address common issues proactively.
• Apply engineering principles and basic automation to enhance our operating environments.
• Monitor applications and help improve their reliability and performance on the Guidewire Cloud Platform.
• Use your software engineering skills to optimize systems and reduce manual tasks.
• Document incidents and assist in refining processes to prevent future occurrences.
• Stay informed about industry trends, tools, and best practices in site reliability engineering.
• Contribute to a culture of innovation, learning, and continuous improvement.
• Participate in on-call rotations to ensure the availability and reliability of our services.
What You Will Bring:
• Experience as an SRE or similar role, focusing on improving system reliability.
• Strong problem-solving skills and the ability to assist in analyzing complex systems and devising effective solutions.
• Effective collaboration and communication skills to work cross-functionally and document processes.
• Experience with automation, monitoring, and performance optimization tools and techniques.
• Commitment to maximizing uptime, scalability, and delivering an exceptional end-user experience.
• Passion for technology and a desire to continuously learn and grow your skills.
• Alignment with Guidewire's mission to leverage technology to help protect and support others.
Required Skills:
• Basic understanding of SLI's, SLO's, and Error Budgets
• Familiarity with application performance monitoring (APM) and telemetry tools
• Some experience with troubleshooting and debugging distributed systems on cloud infrastructure
• Exposure to CICD pipelines within K8S or legacy ecosystems
• Familiarity with creating monitors, dashboards, and synthetic transactions in monitoring tools like Datadog
• Some experience deploying and managing infrastructure within AWS or Kubernetes ecosystems using Terraform or other cloud-native approaches
• Familiarity with infrastructure configuration management using tools such as GitOps, Puppet, or Ansible
• Basic understanding of AWS cloud networking and security
• Comfortable with Linux system administration and the ability to program/script using Python, Go, Java, shell, or equivalent
Preferred Skills:
• Pursuing or interested in pursuing SRE Certification
• Pursuing or interested in pursuing AWS Certification
• Familiarity with SQL, database administration, data pipelines, performance tuning, and schema design
• Exposure to pipelining tools such as Team City, Bitbucket Pipelines, Jenkins, or GitHub Actions
• Interest in learning about open-source distributed data processing frameworks such as Hadoop, Apache Spark, AWS RedShift, etc.
🔍 Dekoder Ogłoszenia
🔴
Assist in troubleshooting and resolving issues in collaboration with development teams, reducing customer impact.
Twoja rola może polegać głównie na reagowaniu na problemy zgłaszane przez innych, a nie na ich proaktywnym zapobieganiu.
🔴
Apply engineering principles and basic automation to enhance our oper
Oczekuje się od Ciebie podstawowej automatyzacji, a nie zaawansowanych rozwiązań, co może oznaczać ograniczone możliwości rozwoju w tym obszarze.
🔴
eager to contribute to the transformation of the insurance industry with our leading cloud platform.
Może to sugerować, że firma jest w trakcie dużych zmian, co może wiązać się z niepewnością i presją.
🟡
play a critical role in ensuring the reliability, performance, and scalability of applications running on our Guidewire Cloud Platform.
Choć brzmi to odpowiedzialnie, może oznaczać, że będziesz głównie utrzymywać istniejący system, a nie budować go od podstaw.
🟡
This position offers a unique opportunity to apply your skills in automation, software engineering, and operational discipline to support our cloud-based solutions.
Podkreślenie 'unikalnej' możliwości może być standardowym zabiegiem marketingowym, a zakres faktycznych zadań może być bardziej ograniczony.