Remote Senior Data Engineer
⚲ Wrocław, Gdańsk, Kraków, Poznań, Warszawa
37 362 - 40 350 PLN (B2B)
Wymagania
- Spark
- AWS
- Linux
- NoSQL
- SQL
- Kafka
- Scala
- Neo4j
- Databricks
- PySpark
- LLM
- GCP
- Python
- Kinesis (nice to have)
- Airflow (nice to have)
- Jenkins (nice to have)
- Python (nice to have)
- TensorFlow (nice to have)
- Parquet (nice to have)
- Delta Lake (nice to have)
Opis stanowiska
O projekcie:
We are looking for Data Engineers to work remotely for an Adtech company that leverages machine learning and data science to build an identity graph that can scale to reach millions of users via brands with programmatically selected households. The work includes scaling our Big Data asset that combines billions of transaction data points including intent, conversions, first party data into an identity graph.
We value technical excellence and you will have both resources and time to deliver world-class code.
This is a 100% remote position, You will be working with team members in NYC.
If you like solving hard and technically challenging problems, join us to use those skills here to create real-time, concurrent, globally distributed systems applications and services.
Wymagania:
• 8+ years of professional software engineering experience, with a focus on data engineering in big data environments.• Expert-level proficiency in Python and at least one other high-level programming language (e.g., Java, Scala, C++, C#).• Proven experience building and optimizing large-scale data pipelines using Apache Spark.• Hands-on experience developing and deploying data solutions in a major cloud platform (AWS, GCP, or Azure).• Experience working with AI, LLMs, Agents, and/or Generative AI technologies, both in product applications and for development productivity.• Relevant technology experience working in the advertising industry.• Experience with Scala• Experience with big data visualization and analytics using OLAP tools.• Familiarity with big data tools and frameworks such as MLFlow, dbt, Kafka, Airflow or Databricks.• Experience with large-scale data management formats like Parquet, Delta Lake, or Iceberg.
Codzienne zadania:
- Design, develop, and maintain highly scalable data pipelines, ETL processes, and data models using Python, Spark, and other big data technologies in a cloud environment (AWS/GCP).
- Lead the integration of AI and LLMs into our data products, working with data scientists to productionize machine learning models and agentic workflows in an efficient and ethical fashion.
- Champion and utilize AI-driven software development tools (e.g., GitHub Copilot) to boost development productivity, improve code quality, and accelerate delivery.
- Work on creating and maintaining reliable and scalable distributed data processing systems
- Become a core maintainer of the data lake
- Maintain our data lake by building searchable data sets for broader business uses
- Scale, troubleshoot and fix existing applications and services
- Own a complex set of services and applications
- Focus ensuring that our data pipelines run 24/7
- Lead technical discussions leading to improvements in tools, processes or projects
- Work on scaling our identity graph to deliver impactful advertising campaigns
- Work on data sets exceeding billions of records
- Scale our MLOps platform by using both traditional ML as well as LLM/Generative AI based applications
We are looking for Data Engineers to work remotely for an Adtech company that leverages machine learning and data science to build an identity graph that can scale to reach millions of users via brands with programmatically selected households. The work includes scaling our Big Data asset that combines billions of transaction data points including intent, conversions, first party data into an identity graph.
We value technical excellence and you will have both resources and time to deliver world-class code.
This is a 100% remote position, You will be working with team members in NYC.
If you like solving hard and technically challenging problems, join us to use those skills here to create real-time, concurrent, globally distributed systems applications and services.
Wymagania:
• 8+ years of professional software engineering experience, with a focus on data engineering in big data environments.• Expert-level proficiency in Python and at least one other high-level programming language (e.g., Java, Scala, C++, C#).• Proven experience building and optimizing large-scale data pipelines using Apache Spark.• Hands-on experience developing and deploying data solutions in a major cloud platform (AWS, GCP, or Azure).• Experience working with AI, LLMs, Agents, and/or Generative AI technologies, both in product applications and for development productivity.• Relevant technology experience working in the advertising industry.• Experience with Scala• Experience with big data visualization and analytics using OLAP tools.• Familiarity with big data tools and frameworks such as MLFlow, dbt, Kafka, Airflow or Databricks.• Experience with large-scale data management formats like Parquet, Delta Lake, or Iceberg.
Codzienne zadania:
- Design, develop, and maintain highly scalable data pipelines, ETL processes, and data models using Python, Spark, and other big data technologies in a cloud environment (AWS/GCP).
- Lead the integration of AI and LLMs into our data products, working with data scientists to productionize machine learning models and agentic workflows in an efficient and ethical fashion.
- Champion and utilize AI-driven software development tools (e.g., GitHub Copilot) to boost development productivity, improve code quality, and accelerate delivery.
- Work on creating and maintaining reliable and scalable distributed data processing systems
- Become a core maintainer of the data lake
- Maintain our data lake by building searchable data sets for broader business uses
- Scale, troubleshoot and fix existing applications and services
- Own a complex set of services and applications
- Focus ensuring that our data pipelines run 24/7
- Lead technical discussions leading to improvements in tools, processes or projects
- Work on scaling our identity graph to deliver impactful advertising campaigns
- Work on data sets exceeding billions of records
- Scale our MLOps platform by using both traditional ML as well as LLM/Generative AI based applications
🔍 Dekoder Ogłoszenia
🔴
leverages machine learning and data science to build an identity graph that can scale to reach millions of users via brands with programmatically selected households
Projekt może być bardziej skoncentrowany na analizie danych i budowaniu modeli niż na tradycyjnym inżynieringu danych, a 'identity graph' może być złożonym i nie do końca zdefiniowanym artefaktem.
🔴
scaling our Big Data asset that combines billions of transaction data points including intent, conversions, first party data into an identity graph
Praca z 'billions of transaction data points' sugeruje bardzo dużą skalę i potencjalnie problemy z wydajnością oraz zarządzaniem danymi.
🟡
you will have both resources and time to deliver world-class code
Obietnica zasobów i czasu może być standardowym frazesem, który nie zawsze przekłada się na rzeczywistość projektu.
🔴
If you like solving hard and technically challenging problems, join us to use those skills here to create real-time, concurrent, globally distributed systems applications and services.
Może to oznaczać, że systemy są faktycznie trudne i wymagają ciągłego rozwiązywania problemów, a nie tylko budowania nowych funkcjonalności.
🟡
Experience working with AI, LLMs, Agents, and/or Generative AI technologies, both in product applications and for development productivity.
Wymaganie doświadczenia z AI/LLM może oznaczać, że firma dopiero zaczyna eksplorować te technologie, a nie ma ugruntowanych, dojrzałych zastosowań.