Data Engineer - Automation & Innovation Department
⚲ Warszawa
Do uzgodnienia
Wymagania
- GCP
- SQL
- Python
- ETL
- data processing
Opis stanowiska
Location: Warszawa, ul. Marynarska 12
Type of contract: Employment contract
Type of cooperation: Hybrid ( 2-3 days a week in the office)
Recruitment online!
You will join a diverse, international team of data professionals working across multiple countries. In this role, you will collaborate with both local experts and international colleagues to build a unified, high-performance data infrastructure. We are looking for a Data Engineer who enjoys solving complex integration challenges and wants to have a real impact on how our organization uses data globally.
What tasks await you?
• Build and maintain data ingestion processes from various sources into the Data Lake.
• Design, develop, and optimize complex data pipelines for reliable data flow.
• Build, develop, and maintain frameworks that facilitate the construction of data pipelines.
• Implement end-to-end testing frameworks for data pipelines.
• Collaborate with data analysts and scientists to ensure the delivery of quality data.
• Ensure robust data governance, security, and compliance practices.
• Explore and implement emerging technologies to improve data pipeline performance.
• Utilize and integrate data from various source system types, including Kafka, MQ, SFTP, databases, APIs, and file shares
What skills will be appreciated?
• Minimum 3 years of experience in a Data Engineering role.
• Practical experience with Cloud (preferably GCP) services, including BigQuery, Dataflow, Google Cloud Storage (GCS), Data Catalog, and dbt.
•
Advanced SQL proficiency, including writing complex queries, query optimization, Common Table Expressions (CTEs), window functions, and related advanced SQL techniques.
• Python coding skills (experience with Pandas and NumPy is a plus)
• Experience working with Apache Airflow for pipeline orchestration, workflow automation, monitoring, data quality validation, alerting, and troubleshooting.
• Experience with Apache Spark, Hadoop, and Hive for large-scale data processing.
• Experience designing, developing, testing, maintaining, and optimizing ETL/ELT pipelines for batch and streaming data processing, including Apache Kafka, API-based data ingestion.
• Experience designing and implementing data integration solutions, data warehouse architectures, data models, and data layers to support analytics and reporting.
• Experience implementing and maintaining CI/CD pipelines and using GitLab for version control and technical documentation.
• Experience working in Agile/Scrum environments.
• Good command of English.
Nice to have:
• Bachelor’s or Master’s degree in Data Science, Statistics, Computer Science, Economics, or a related field.
• Basic understanding of Machine Learning and MLOps principles
• Strong understanding of data governance principles, including metadata management and data quality frameworks.
• Experience in international or multi-country data projects is preferred.
• Strong communication skills to translate technical findings into business-oriented insights.
• Self-starter with a continuous learning mindset and strong attention to detail.
Type of contract: Employment contract
Type of cooperation: Hybrid ( 2-3 days a week in the office)
Recruitment online!
You will join a diverse, international team of data professionals working across multiple countries. In this role, you will collaborate with both local experts and international colleagues to build a unified, high-performance data infrastructure. We are looking for a Data Engineer who enjoys solving complex integration challenges and wants to have a real impact on how our organization uses data globally.
What tasks await you?
• Build and maintain data ingestion processes from various sources into the Data Lake.
• Design, develop, and optimize complex data pipelines for reliable data flow.
• Build, develop, and maintain frameworks that facilitate the construction of data pipelines.
• Implement end-to-end testing frameworks for data pipelines.
• Collaborate with data analysts and scientists to ensure the delivery of quality data.
• Ensure robust data governance, security, and compliance practices.
• Explore and implement emerging technologies to improve data pipeline performance.
• Utilize and integrate data from various source system types, including Kafka, MQ, SFTP, databases, APIs, and file shares
What skills will be appreciated?
• Minimum 3 years of experience in a Data Engineering role.
• Practical experience with Cloud (preferably GCP) services, including BigQuery, Dataflow, Google Cloud Storage (GCS), Data Catalog, and dbt.
•
Advanced SQL proficiency, including writing complex queries, query optimization, Common Table Expressions (CTEs), window functions, and related advanced SQL techniques.
• Python coding skills (experience with Pandas and NumPy is a plus)
• Experience working with Apache Airflow for pipeline orchestration, workflow automation, monitoring, data quality validation, alerting, and troubleshooting.
• Experience with Apache Spark, Hadoop, and Hive for large-scale data processing.
• Experience designing, developing, testing, maintaining, and optimizing ETL/ELT pipelines for batch and streaming data processing, including Apache Kafka, API-based data ingestion.
• Experience designing and implementing data integration solutions, data warehouse architectures, data models, and data layers to support analytics and reporting.
• Experience implementing and maintaining CI/CD pipelines and using GitLab for version control and technical documentation.
• Experience working in Agile/Scrum environments.
• Good command of English.
Nice to have:
• Bachelor’s or Master’s degree in Data Science, Statistics, Computer Science, Economics, or a related field.
• Basic understanding of Machine Learning and MLOps principles
• Strong understanding of data governance principles, including metadata management and data quality frameworks.
• Experience in international or multi-country data projects is preferred.
• Strong communication skills to translate technical findings into business-oriented insights.
• Self-starter with a continuous learning mindset and strong attention to detail.
🔍 Dekoder Ogłoszenia
🟡
wants to have a real impact on how our organization uses data globally
Oznacza, że Twoja praca będzie miała znaczenie, ale niekoniecznie oznacza to dużą autonomię czy możliwość wprowadzania rewolucyjnych zmian od razu.
🟡
Explore and implement emerging technologies to improve data pipeline performance
Może oznaczać, że będziesz miał okazję pracować z nowymi technologiami, ale też że obecne rozwiązania mogą być przestarzałe i wymagać ciągłych poprawek.
🟡
Build, develop, and maintain frameworks that facilitate the construction of data pipelines
Sugestia, że będziesz nie tylko budował pipeline'y, ale także tworzył narzędzia i procesy, które ułatwią ich tworzenie innym, co może oznaczać większą odpowiedzialność i pracę nad infrastrukturą.
🟡
Collaborate with data analysts and scientists to ensure the delivery of quality data
Wskazuje na potrzebę dobrej komunikacji i współpracy, ale może też oznaczać, że będziesz musiał często negocjować wymagania i rozwiązywać problemy wynikające z różnic w potrzebach między zespołami.
🟡
Ensure robust data governance, security, and compliance practices
Podkreśla wagę tych aspektów, ale może też oznaczać, że będziesz musiał poświęcić znaczną część czasu na spełnianie restrykcyjnych wymagań, a nie tylko na rozwój nowych funkcjonalności.