Data Engineer
⚲ Poland, Hungary
Do uzgodnienia
Wymagania
- Apache Beam
- Spark
- Apache Airflow
- SQL
- Python
Opis stanowiska
We are looking for a Data Engineer to build and operate scalable, reliable data pipelines that power analytics, reporting, and downstream data products. This role focuses on distributed data processing, orchestration, and strong engineering fundamentals.
Your responsibilities:
• Design, build, and operate scalable batch and streaming data pipelines.
• Implement distributed data processing using Spark, Apache Beam, or similar frameworks.
• Orchestrate workflows using Apache Airflow.
• Develop and maintain transformations using SQL and dbt.
• Ensure data quality, observability, and performance across pipelines.
• Collaborate closely with Analytics Engineers and Platform Engineers.
• Build reliable, well‑modelled datasets in the data warehouse.
We are looking for you, if you have:
• Strong experience with Google Cloud Platform (GCP) data services, or 5+ years on another major cloud provider.
• Strong Python skills for data pipeline development.
• Strong SQL skills and hands‑on experience with dbt.
• Hands‑on experience with distributed data processing frameworks (Spark, Beam, Flink, etc.).
• Strong experience with Apache Airflow.
• Some experience with streaming systems (Pub/Sub, Kafka, etc.).
• Mandatory experience using AI‑assisted coding tools in day‑to‑day development.
• Experience with AI‑driven data workflows (e.g. ML pipelines, feature generation, LLM‑adjacent pipelines).
• Experience supporting analytics or BI‑driven use cases.
• Exposure to data quality or governance tooling.
We offer:
• Participation in interesting and demanding projects.
• Flexible working hours.
• A great, non-corporate atmosphere.
• Possibility to work remote or hybrid (2 days per week from the office).
• Opportunities for development and promotion.
• Attractive package of benefits.
We reserve the right to contact the selected candidates.
Your responsibilities:
• Design, build, and operate scalable batch and streaming data pipelines.
• Implement distributed data processing using Spark, Apache Beam, or similar frameworks.
• Orchestrate workflows using Apache Airflow.
• Develop and maintain transformations using SQL and dbt.
• Ensure data quality, observability, and performance across pipelines.
• Collaborate closely with Analytics Engineers and Platform Engineers.
• Build reliable, well‑modelled datasets in the data warehouse.
We are looking for you, if you have:
• Strong experience with Google Cloud Platform (GCP) data services, or 5+ years on another major cloud provider.
• Strong Python skills for data pipeline development.
• Strong SQL skills and hands‑on experience with dbt.
• Hands‑on experience with distributed data processing frameworks (Spark, Beam, Flink, etc.).
• Strong experience with Apache Airflow.
• Some experience with streaming systems (Pub/Sub, Kafka, etc.).
• Mandatory experience using AI‑assisted coding tools in day‑to‑day development.
• Experience with AI‑driven data workflows (e.g. ML pipelines, feature generation, LLM‑adjacent pipelines).
• Experience supporting analytics or BI‑driven use cases.
• Exposure to data quality or governance tooling.
We offer:
• Participation in interesting and demanding projects.
• Flexible working hours.
• A great, non-corporate atmosphere.
• Possibility to work remote or hybrid (2 days per week from the office).
• Opportunities for development and promotion.
• Attractive package of benefits.
We reserve the right to contact the selected candidates.
🔍 Dekoder Ogłoszenia
🔴
build and operate scalable, reliable data pipelines that power analytics, reporting, and downstream data products
Oczekuje się, że będziesz nie tylko tworzyć, ale także utrzymywać i rozwiązywać problemy w istniejących potokach danych.
🟡
strong engineering fundamentals
Może oznaczać potrzebę solidnej wiedzy teoretycznej i praktycznej z zakresu inżynierii oprogramowania, nie tylko specyficznych narzędzi.
🔴
Ensure data quality, observability, and performance across pipelines
Oprócz budowania potoków, będziesz odpowiedzialny za ich monitorowanie, diagnozowanie problemów i zapewnienie ich niezawodności.
🔴
5+ years on another major cloud provider
Jeśli nie masz doświadczenia z GCP, musisz mieć bardzo bogate doświadczenie z AWS lub Azure, aby zostać rozważonym.
🔴
Mandatory experience using AI‑assisted coding tools in day‑to‑day development
Oczekuje się aktywnego wykorzystania narzędzi takich jak Copilot, a nie tylko posiadania wiedzy o ich istnieniu.