JustJoin.IT Praca zdalna Senior

Senior DevOps

ALOKAI SPÓŁKA Z OGRANICZONĄ ODPOWIEDZIALNOŚCIĄ

⚲ Warszawa

25 000 - 32 000 PLN netto (B2B)

Wymagania

  • Kubernetes
  • GCP
  • Terraform
  • Python
  • Istio
  • Elasticsearch
  • Node.js
  • SRE
  • Datadog

Opis stanowiska

Position Overview
We are seeking an experienced Senior Infrastructure/SRE Engineer with strong Python skills to join our team. This is primarily an infrastructure-focused role — you'll contribute to a broader rearchitecture of our infrastructure alongside the rest of the team, while also contributing to our Python codebase. The split is roughly 70% infrastructure / 30% coding in Python. A deep understanding of GCP, Kubernetes, Istio, and SRE best practices is critical. You will drive the reliability, scalability, and cost-efficiency of our infrastructure while enabling rapid and stable application delivery.

Tech Stack
GCP, Kubernetes, Istio, Helm, Terraform, Docker, Spacelift, GitHub Actions, Argo Ecosystem (CD, Rollouts), Fastly + VCL, DataDog, Elasticsearch, Python, Django, Node.js.

Key Responsibilities
• Manage and optimize Kubernetes clusters (including Istio service mesh) to ensure high availability and performance.
• Continuously monitor and refine cloud resource allocation.
• Troubleshoot issues and implement fixes to maintain system stability.
• Develop and enhance monitoring and alerting systems built on top of DataDog and Elasticsearch.
• Ensure proactive identification and resolution of potential issues.
• Evaluate and improve Content Delivery Network (CDN) configurations for optimal performance and cost-efficiency.
• Work with development teams and customer support to create custom dashboards and alerts for specific applications.
• Build Cloud-based features that are available to our customers.
• Develop and maintain APIs to support application configuration management.
• Collaborate with application teams to refine API requirements and ensure seamless integration.

On-Call Rotation
• Participate in an on-call rotation to ensure 24/7 system reliability.
• Respond promptly to incidents and implement corrective actions.

Qualifications

Experience
• 5+ years with GCP (or other cloud providers), Kubernetes, and modern infrastructure tooling.
• Hands-on experience designing, rearchitecting, or rebuilding infrastructure at scale (not just maintaining what's already there).
• 5+ years with high-traffic infrastructures and monitoring tools.
• Hands-on involvement in designing and managing SLO-driven systems.
• Experience with Istio or other service mesh technologies is a strong plus.
• Experience with handling multi-tenant environments is a plus.
• 3+ years in Python with DevOps/SRE exposure.

Technical Skills
• Strong knowledge of containerization and orchestration (Docker, Kubernetes, Istio).
• Hands-on experience with monitoring tools like DataDog, Elasticsearch, Prometheus, or similar.
• Familiarity with CI/CD pipelines (Argo Ecosystem: CD, Rollouts) and infrastructure-as-code tools like Terraform and Helm.
• Strong grasp of CDN technologies, caching mechanisms, and their impact on application performance.
• Good proficiency in Python (Django) — contributing to API/configuration-management code, even as infra work takes priority.
• SRE mindset — proactive reliability work, not just reactive firefighting.

Soft Skills
• Analytical mindset with a focus on proactive problem-solving.
• Excellent communication skills to collaborate with cross-functional teams.
• Ability to work in a fast-paced, dynamic environment.

🔍 Dekoder Ogłoszenia

🔴
This is primarily an infrastructure-focused role — you'll contribute to a broader rearchitecture of our infrastructure alongside the rest of the team, while also contributing to our Python codebase. The split is roughly 70% infrastructure / 30% coding in Python.
Chociaż ogłoszenie podkreśla rolę infrastrukturalną, 30% czasu poświęconego na kodowanie w Pythonie może oznaczać znaczące zaangażowanie w rozwój oprogramowania, a nie tylko skrypty.
🔴
A deep understanding of GCP, Kubernetes, Istio, and SRE best practices is critical.
Określenie 'deep understanding' może sugerować, że oczekiwana jest wiedza ekspercka, a nie tylko podstawowa znajomość wymienionych technologii.
🔴
You will drive the reliability, scalability, and cost-efficiency of our infrastructure while enabling rapid and stable application delivery.
Fraza 'drive the reliability, scalability, and cost-efficiency' może oznaczać odpowiedzialność za rozwiązywanie problemów i optymalizację, które mogą być czasochłonne i wymagające.