IT - Site Reliability Engineer
⚲ Pune
Do uzgodnienia
Wymagania
- Incident management
- Configuration management
- Configuration Management (ITIL)
- Operations
- Python
- Cloud
- Powershell
- Security
- Microsoft Azure
- CI/CD
Opis stanowiska
Summary: The main function of the Site Reliability Engineer (SRE) role is to ensure the reliability, availability, and performance of critical systems and services in a 24/7 operational environment. The SRE will leverage expertise in cloud platforms (primarily Microsoft Azure, with additional knowledge of AWS) to maintain seamless operations and respond to incidents promptly.Responsibilities:
• Monitor production systems and services using observability tools (logs, metrics, traces, dashboards).
• Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.
• Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.
• Design, implement, and maintain observability solutions (e.g., Prometheus, Grafana, ELK and similar tools).
• Develop and refine dashboards, alerts, and automated health checks for critical infrastructure and applications.
• Analyze system performance and reliability data to identify trends and prevent future incidents.
• Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.
• Automate operational tasks and incident response processes using scripting and configuration management tools.
• Document procedures, runbooks, and incident reports for knowledge sharing and continuous improvement.
• Conduct post-incident reviews and root cause analysis to drive improvements in reliability and response.
• Propose and implement enhancements to monitoring, alerting, and operational processes.
Key Requirements:
• Bachelor's degree in Information Technology, Computer Science, Business Administration, or related field.
• 2-5 years of experience in cloud engineering and operations engineering.
• Experience with Azure services; AWS and GCP knowledge is a plus.
• Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform.
• Strong scripting skills in Python, Bash, or PowerShell.
• Familiarity with GitLab CI/CD tools.
• Proficiency in monitoring and logging tools (e.g., native cloud tools, OpenMetrics, OpenTelemetry).
Nice to Have:
• Master's degree or relevant certifications.
• Experience in developing health checks for critical infrastructure.
Other Details:
• Location: Pune
• Team Structure: 24/7 shift rotation
• Monitor production systems and services using observability tools (logs, metrics, traces, dashboards).
• Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.
• Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.
• Design, implement, and maintain observability solutions (e.g., Prometheus, Grafana, ELK and similar tools).
• Develop and refine dashboards, alerts, and automated health checks for critical infrastructure and applications.
• Analyze system performance and reliability data to identify trends and prevent future incidents.
• Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.
• Automate operational tasks and incident response processes using scripting and configuration management tools.
• Document procedures, runbooks, and incident reports for knowledge sharing and continuous improvement.
• Conduct post-incident reviews and root cause analysis to drive improvements in reliability and response.
• Propose and implement enhancements to monitoring, alerting, and operational processes.
Key Requirements:
• Bachelor's degree in Information Technology, Computer Science, Business Administration, or related field.
• 2-5 years of experience in cloud engineering and operations engineering.
• Experience with Azure services; AWS and GCP knowledge is a plus.
• Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform.
• Strong scripting skills in Python, Bash, or PowerShell.
• Familiarity with GitLab CI/CD tools.
• Proficiency in monitoring and logging tools (e.g., native cloud tools, OpenMetrics, OpenTelemetry).
Nice to Have:
• Master's degree or relevant certifications.
• Experience in developing health checks for critical infrastructure.
Other Details:
• Location: Pune
• Team Structure: 24/7 shift rotation
🔍 Dekoder Ogłoszenia
🔴
ensure the reliability, availability, and performance of critical systems and services in a 24/7 operational environment
Oczekuje się ciągłej gotowości do pracy i reagowania na problemy przez całą dobę, 7 dni w tygodniu.
🔴
Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.
Może oznaczać konieczność natychmiastowego reagowania na problemy, często poza standardowymi godzinami pracy.
🟡
Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.
Praca w systemie zmianowym, obejmującym dyżury w nocy, weekendy i święta, co jest standardem dla SRE.
🟢
Automate operational tasks and incident response processes using scripting and configuration management tools.
Oczekuje się proaktywnego podejścia do usprawniania procesów poprzez automatyzację, co może wymagać nauki nowych narzędzi.
🟡
Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.
Wymaga umiejętności pracy w zespole i komunikacji z różnymi działami, aby osiągnąć wspólne cele.