JustJoin.IT Hybrydowo Mid

Service Reliability Engineer

Toyota Digital HUB & SSC (TME NV/SA)

⚲ Wrocław

12 000 - 18 000 PLN brutto (UoP)

Wymagania

  • AWS
  • Java
  • Python
  • Datadog
  • Docker
  • Kubernets
  • SQL
  • CI/CD
  • MongoDB
  • Site Reliability Engineering

Opis stanowiska

Team description: 
We Lead the end-to-end development, deployment, and continuous enhancement of Toyota’s connected services ecosystem across European markets. Our work oversees the integration of vehicle connectivity platforms, mobile applications, and cloud-based services, ensuring scalability, security, and seamless user experience. We Collaborate with engineering, product, IT, legal, and external partners to accelerate service rollout and maintain operational excellence. 

 Role Summary:  
We would like to invite you to join the adventure of expanding TME PL branch in the Toyota Digital HUB in Wroclaw. We are seeking a hands-on Service Reliability Engineer (SRE) to improve service reliability, performance, and operational excellence across our connected services platform. You will work closely with backend engineering teams (primarily Java/Spring Boot) to build automation, ensure observability, operate AWS cloud workloads, drive incident management, and maintain high service availability for millions of Toyota customers across Europe. 
 

Key Responsibilities:  • Partner with product teams to define and enforce SLIs, SLOs, for backend microservices. 
• Improve reliability of a diverse microservices ecosystem (Java/Spring Boot, Python, Node.js). 
• Optimize system performance: latency, throughput, memory/GC tuning, JVM diagnostics. 
• Drive resilience patterns (circuit breakers, retries, graceful degradation). 
• Build deep observability using Datadog (primary): metrics, logs, traces, synthetics, RUM, APM. 
• Triage and restore service during incidents; ensure proactive communication. 
• Conduct post-incident reviews, root cause analysis, and track long-term remediations. 
• Operate backend services across EKS and ECS  
• Manage AWS components: API Gateway, Lambda, ECS, EKS, RDS, IAM, S3, VPC, CloudWatch, ALB/NLB. 
• Work with CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI) 
• Build tooling and automation using Python and/or shell. 
• Operate and troubleshoot MSSQL/SQL and MongoDB (replica sets, indexes, PITR). 
• Ensure backup/restore reliability, failover, performance tuning. 
 

Qualifications & Skills:  
WHAT WE REQUIRE: 
• 3+ years as a Service Reliability Engineer,  
• Strong hands-on experience with Java + Spring Boot (core backend for MyToyota). 
• Good proficiency in Python or another backend programming language. 
• Knowledge of AWS services, including EKS, ECS, IAM, VPC, RDS, S3, and API Gateway. 
• Hands-on experience with Datadog (APM, Logs, Dashboards, SLOs, RUM, and Synthetic Monitoring). 
• Docker & Kubernetes 
• SQL & NoSQL:  
• MSSQL or other SQL dialects 
• MongoDB (indexes, replica sets, perf tuning) 
• CI/CD pipelines: GitHub Actions, Jenkins, GitLab CI. 
• Knowledge about queues & messaging (SQS/SNS, Kafka). 
• Basic Linux and networking fundamentals (HTTP, TLS, DNS, load balancing). 
 
ADDITIONAL REQUIERMENTS (considered an advantage):  
• Previous Experience in the automotive industry 
 
You have a  TOYOTA DNA, this means: 
• Courage: you are ready to let go of the easy path to reach challenging targets 
• Creativity: your passion drives you to explore innovative ideas and challenge the impossible 
• Coaching: you share knowledge and feedback with your colleagues and celebrate each other's success 
• Curiosity: you are open for new experience and able to combine imagination with fact-based observation 
• Collaboration: you are a team player, respectful and inclusive in your style and you take a customer-oriented approach 
 

Formal Role Details: 
• Job Type: Employment contract 
• Starting date: To be agreed. Position available since October 2026 
• Location: Wrocław, Silver Tower Office Centre

🔍 Dekoder Ogłoszenia

🔴
We Lead the end-to-end development, deployment, and continuous enhancement of Toyota’s connected services ecosystem across European markets.
Oznacza to, że zespół jest odpowiedzialny za cały cykl życia usług, od pomysłu po utrzymanie, co może wiązać się z szerokim zakresem obowiązków i potencjalnie dużą odpowiedzialnością.
🔴
You will work closely with backend engineering teams (primarily Java/Spring Boot) to build automation, ensure observability, operate AWS cloud workloads, drive incident management, and maintain high service availability for millions of Toyota customers across Europe.
Choć brzmi to jak standardowe zadania SRE, 'operate AWS cloud workloads' może oznaczać zarówno zarządzanie istniejącą infrastrukturą, jak i jej aktywne budowanie i optymalizację, co wymaga szerokich kompetencji chmurowych.
🔴
Partner with product teams to define and enforce SLIs, SLOs, for backend microservices.
Współpraca z zespołami produktowymi w celu definiowania i egzekwowania SLI/SLO może oznaczać konieczność negocjacji, przekonywania i radzenia sobie z różnymi priorytetami biznesowymi.
🔴
Improve reliability of a diverse microservices ecosystem (Java/Spring Boot, Python, Node.js).
Praca nad 'diverse microservices ecosystem' może oznaczać konieczność radzenia sobie z wieloma różnymi technologiami i architekturami, co wymaga elastyczności i szybkiego uczenia się.
🔴
Build deep observability using Datadog (primary): metrics, l
Choć Datadog jest potężnym narzędziem, 'build deep observability' sugeruje, że obecny stan może być niewystarczający i wymagać znaczącego nakładu pracy w celu stworzenia kompleksowego systemu monitorowania.