JustJoin.IT Hybrydowo Junior

IT - Site Reliability Engineer

emagine Polska

⚲ Pune

Do uzgodnienia

Wymagania

  • Incident management
  • Configuration management
  • Configuration Management (ITIL)
  • Operations
  • Python
  • Cloud
  • Powershell
  • Security
  • Microsoft Azure
  • CI/CD

Opis stanowiska

Summary: The main function of the Site Reliability Engineer (SRE) role is to ensure the reliability, availability, and performance of critical systems and services in a 24/7 operational environment. The SRE will leverage expertise in cloud platforms (primarily Microsoft Azure, with additional knowledge of AWS) to maintain seamless operations and respond to incidents promptly.Responsibilities:
• Monitor production systems and services using observability tools (logs, metrics, traces, dashboards).
• Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.
• Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.
• Design, implement, and maintain observability solutions (e.g., Prometheus, Grafana, ELK and similar tools).
• Develop and refine dashboards, alerts, and automated health checks for critical infrastructure and applications.
• Analyze system performance and reliability data to identify trends and prevent future incidents.
• Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.
• Automate operational tasks and incident response processes using scripting and configuration management tools.
• Document procedures, runbooks, and incident reports for knowledge sharing and continuous improvement.
• Conduct post-incident reviews and root cause analysis to drive improvements in reliability and response.
• Propose and implement enhancements to monitoring, alerting, and operational processes.

Key Requirements:
• Bachelor's degree in Information Technology, Computer Science, Business Administration, or related field.
• 2-5 years of experience in cloud engineering and operations engineering.
• Experience with Azure services; AWS and GCP knowledge is a plus.
• Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform.
• Strong scripting skills in Python, Bash, or PowerShell.
• Familiarity with GitLab CI/CD tools.
• Proficiency in monitoring and logging tools (e.g., native cloud tools, OpenMetrics, OpenTelemetry).

Nice to Have:
• Master's degree or relevant certifications.
• Experience in developing health checks for critical infrastructure.

Other Details:
• Location: Pune
• Team Structure: 24/7 shift rotation

🔍 Dekoder Ogłoszenia

🔴
ensure the reliability, availability, and performance of critical systems and services in a 24/7 operational environment
Oczekuje się ciągłej gotowości do pracy i reagowania na problemy przez całą dobę, 7 dni w tygodniu.
🔴
Respond to incidents, alerts, and outages in real-time, ensuring rapid resolution and minimal impact.
Może oznaczać konieczność natychmiastowego reagowania na problemy, często poza standardowymi godzinami pracy.
🟡
Participate in a rotating on-call schedule, providing support during nights, weekends, and holidays.
Praca w systemie zmianowym, obejmującym dyżury w nocy, weekendy i święta, co jest standardem dla SRE.
🟢
Automate operational tasks and incident response processes using scripting and configuration management tools.
Oczekuje się proaktywnego podejścia do usprawniania procesów poprzez automatyzację, co może wymagać nauki nowych narzędzi.
🟡
Collaborate with development, infrastructure, application and security teams to ensure system reliability and scalability.
Wymaga umiejętności pracy w zespole i komunikacji z różnymi działami, aby osiągnąć wspólne cele.