Senior Site Reliability Engineer
⚲ Kraków
28 000 - 32 000 PLN (B2B)
Wymagania
- Docker
- Datadog
- Observability
- Cosmos DB
- MySQL
- CI/CD
- GitHub Actions
- ELK Stack (nice to have)
- Grafana (nice to have)
- Prometheus (nice to have)
- .NET (nice to have)
Opis stanowiska
O projekcie:
OneRail seeks a Senior Site Reliability Engineer to help ensure the availability, performance, scalability, observability, and resilience of our SaaS platform. In this role, you’ll be responsible for designing and optimizing platform architecture, improving system reliability, driving automation initiatives, and enabling engineering teams to build and operate highly scalable services.
This is a hands-on role for someone with deep expertise in cloud-native platforms, distributed systems, observability, and performance optimization. You’ll work closely with Engineering, Product, and Operations teams to improve platform reliability, scalability, and operational efficiency while reducing manual overhead through automation. Over time, you’ll have opportunities to influence platform architecture, lead reliability initiatives, and establish Site Reliability Engineering best practices across the organization.
Wymagania:
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.- 5+ years of experience in Site Reliability Engineering, Platform Engineering or a related role.- Strong experience with cloud-native architectures and distributed systems.- Hands-on experience implementing observability solutions, including metrics, logging, tracing, and application performance monitoring.- Experience designing scalable backend architecture using Node.js, TypeScript, .NET- Strong knowledge of database administration, performance tuning, and optimization, including Azure Cosmos DB, MySQL, or similar platforms.- Experience building automation frameworks, scripting solutions, and operational tooling.- Familiarity with CI/CD pipelines and continuous integration practices using GitHub Actions or similar platforms.- Experience with containerization technologies such as Docker.- Strong understanding of event-driven architectures, messaging systems, and real-time data processing platforms.- Experience with monitoring and observability tools such as Datadog, OpenTelemetry, ELK Stack, App Insights, Grafana, or Prometheus.- Excellent troubleshooting, analytical, and problem-solving skills.- Strong written and verbal communication skills with the ability to collaborate across technical and business teams.- Advanced proficiency in English & Polish, both written and spoken (B2+).
Codzienne zadania:
- Serve as the platform subject matter expert, mentoring engineering teams on reliability, scalability, security, and operational best practices.
- Design, implement, and maintain observability solutions covering logs, metrics, traces, APM, and alerting across all platform services.
- Benchmark, analyze, and optimize application performance, cloud services, integrations, databases, and distributed systems.
- Build and maintain performance, load, chaos, and resilience testing frameworks to proactively identify system weaknesses.
- Perform advanced analysis, tuning, and optimization of structured and unstructured databases to improve performance and scalability.
- Develop automation frameworks and operational tooling that eliminate manual processes and reduce human error.
- Lead incident response efforts, conduct root cause analysis, and drive continuous improvement through postmortem reviews and reliability initiatives.
- Optimize the performance of databases, caches, streaming platforms, message brokers, and backend services.
- Collaborate closely with Engineering, Product, and Operations teams to deliver highly available, production-grade platform solutions.
- Establish and maintain monitoring, alerting, and reliability standards across the platform.
- Create and maintain technical documentation, runbooks, architectural diagrams, and operational procedures.
- Evaluate emerging technologies and recommend solutions that improve platform reliability, scalability, observability, and operational efficiency.
OneRail seeks a Senior Site Reliability Engineer to help ensure the availability, performance, scalability, observability, and resilience of our SaaS platform. In this role, you’ll be responsible for designing and optimizing platform architecture, improving system reliability, driving automation initiatives, and enabling engineering teams to build and operate highly scalable services.
This is a hands-on role for someone with deep expertise in cloud-native platforms, distributed systems, observability, and performance optimization. You’ll work closely with Engineering, Product, and Operations teams to improve platform reliability, scalability, and operational efficiency while reducing manual overhead through automation. Over time, you’ll have opportunities to influence platform architecture, lead reliability initiatives, and establish Site Reliability Engineering best practices across the organization.
Wymagania:
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.- 5+ years of experience in Site Reliability Engineering, Platform Engineering or a related role.- Strong experience with cloud-native architectures and distributed systems.- Hands-on experience implementing observability solutions, including metrics, logging, tracing, and application performance monitoring.- Experience designing scalable backend architecture using Node.js, TypeScript, .NET- Strong knowledge of database administration, performance tuning, and optimization, including Azure Cosmos DB, MySQL, or similar platforms.- Experience building automation frameworks, scripting solutions, and operational tooling.- Familiarity with CI/CD pipelines and continuous integration practices using GitHub Actions or similar platforms.- Experience with containerization technologies such as Docker.- Strong understanding of event-driven architectures, messaging systems, and real-time data processing platforms.- Experience with monitoring and observability tools such as Datadog, OpenTelemetry, ELK Stack, App Insights, Grafana, or Prometheus.- Excellent troubleshooting, analytical, and problem-solving skills.- Strong written and verbal communication skills with the ability to collaborate across technical and business teams.- Advanced proficiency in English & Polish, both written and spoken (B2+).
Codzienne zadania:
- Serve as the platform subject matter expert, mentoring engineering teams on reliability, scalability, security, and operational best practices.
- Design, implement, and maintain observability solutions covering logs, metrics, traces, APM, and alerting across all platform services.
- Benchmark, analyze, and optimize application performance, cloud services, integrations, databases, and distributed systems.
- Build and maintain performance, load, chaos, and resilience testing frameworks to proactively identify system weaknesses.
- Perform advanced analysis, tuning, and optimization of structured and unstructured databases to improve performance and scalability.
- Develop automation frameworks and operational tooling that eliminate manual processes and reduce human error.
- Lead incident response efforts, conduct root cause analysis, and drive continuous improvement through postmortem reviews and reliability initiatives.
- Optimize the performance of databases, caches, streaming platforms, message brokers, and backend services.
- Collaborate closely with Engineering, Product, and Operations teams to deliver highly available, production-grade platform solutions.
- Establish and maintain monitoring, alerting, and reliability standards across the platform.
- Create and maintain technical documentation, runbooks, architectural diagrams, and operational procedures.
- Evaluate emerging technologies and recommend solutions that improve platform reliability, scalability, observability, and operational efficiency.
🔍 Dekoder Ogłoszenia
🔴
hands-on role for someone with deep expertise
Oczekuje się, że będziesz aktywnie wdrażać rozwiązania i rozwiązywać problemy techniczne, a nie tylko delegować zadania.
🔴
opportunities to influence platform architecture, lead reliability initiatives, and establish Site Reliability Engineering best practices across the organization
Początkowo możesz mieć ograniczony wpływ, a te możliwości mogą pojawić się dopiero po dłuższym czasie i udowodnieniu swojej wartości.
🔴
reducing manual overhead through automation
Może to oznaczać, że obecne procesy są w dużej mierze manualne i wymagają znaczącej pracy w celu ich zautomatyzowania.
🟡
5+ years of experience
Chociaż podano '5+', często firmy szukają kandydatów z doświadczeniem bliżej górnej granicy lub nawet wyższym.