Site Reliability Engineer (SRE)
Do uzgodnienia
Wymagania
- Node.js
- TypeScript
- AWS
- Cloudflare
- Sentry
Opis stanowiska
About the Client
Founded in 2014 as a dating app, Feeld gathered millions of users in one place to create a safer and more inclusive space online for everyone open to experiencing people and relationships in a new way. Their mission is to elevate the human experience of sexuality and relationships and create a world where everyone is more intimately connected to each other and themselves.
About the Project
You’ll join a consumer mobile product with an engineering and product organization of around 50 people distributed across Europe and the US. The team works in small, autonomous product squads, each responsible for a specific area of the product and critical user journeys. From a technical perspective, the team is focused on building reliable, observable systems that allow engineers to identify issues early, understand their impact and respond quickly when incidents occur.
About the Role
We are looking for an experienced Site Reliability Engineer with a strong backend engineering background in Node.js and TypeScript. You will be embedded within a product squad and take ownership of the reliability and observability of critical user journeys. An important part of the role is being able to interpret production signals, identify when something is going wrong and begin mitigating incidents independently while bringing in the wider engineering team when needed.
Responsibilities:
- Own observability for critical product and user journeys within your squad.
- Define, build and maintain meaningful metrics, dashboards and alerts.
- Define and maintain SLIs/SLOs for key services and product-level metrics.
- Improve monitoring, logging, tracing and alerting across the squad’s systems.
- Act as the first responder for critical P0/P1 production incidents, including out-of-hours incidents.
- Investigate production signals, identify potential root causes and begin mitigating issues independently.
- Coordinate with other engineers when broader support or escalation is required.
- Participate in incident triage, mitigation and postmortems.
- Identify recurring reliability issues and drive improvements to infrastructure, tooling and incident-response processes.
- Work closely with backend and product engineers in a distributed, autonomous squad.
Founded in 2014 as a dating app, Feeld gathered millions of users in one place to create a safer and more inclusive space online for everyone open to experiencing people and relationships in a new way. Their mission is to elevate the human experience of sexuality and relationships and create a world where everyone is more intimately connected to each other and themselves.
About the Project
You’ll join a consumer mobile product with an engineering and product organization of around 50 people distributed across Europe and the US. The team works in small, autonomous product squads, each responsible for a specific area of the product and critical user journeys. From a technical perspective, the team is focused on building reliable, observable systems that allow engineers to identify issues early, understand their impact and respond quickly when incidents occur.
About the Role
We are looking for an experienced Site Reliability Engineer with a strong backend engineering background in Node.js and TypeScript. You will be embedded within a product squad and take ownership of the reliability and observability of critical user journeys. An important part of the role is being able to interpret production signals, identify when something is going wrong and begin mitigating incidents independently while bringing in the wider engineering team when needed.
Responsibilities:
- Own observability for critical product and user journeys within your squad.
- Define, build and maintain meaningful metrics, dashboards and alerts.
- Define and maintain SLIs/SLOs for key services and product-level metrics.
- Improve monitoring, logging, tracing and alerting across the squad’s systems.
- Act as the first responder for critical P0/P1 production incidents, including out-of-hours incidents.
- Investigate production signals, identify potential root causes and begin mitigating issues independently.
- Coordinate with other engineers when broader support or escalation is required.
- Participate in incident triage, mitigation and postmortems.
- Identify recurring reliability issues and drive improvements to infrastructure, tooling and incident-response processes.
- Work closely with backend and product engineers in a distributed, autonomous squad.
🔍 Dekoder Ogłoszenia
🟡
small, autonomous product squads
Małe zespoły = mniej wsparcia, więcej samodzielności i odpowiedzialności za całość