JustJoin.IT Hybrydowo Senior

Senior Data Engineer / MLOps Specialist

emagine Polska

⚲ Lisbon

Do uzgodnienia

Wymagania

  • GitHub
  • Data storage
  • Microsoft Platform
  • Machine Learning (ML)
  • Data Warehouse (DW)
  • Operations
  • Docker
  • Testing
  • Cloud
  • CI/CD

Opis stanowiska

1. Context and Mission of the Position 
We are expanding our data engineering structure and are seeking a professional with a technical profile, focused on advanced data engineering, solid foundations in DevOps, and knowledge of MLOps. This individual will be responsible for designing, optimizing, and maintaining the infrastructure and pipelines that support local data projects, data science projects, and the analytical operations of the organization, serving as a critical bridge between model development and large-scale production.

2. Main Responsibilities and Activities 
● Complex Pipeline Engineering: Develop, monitor, and automate robust data ingestion pipelines (Batch and Streaming), handling complex scenarios such as advanced integration of third-party APIs and web scraping of application APIs.
● Data Warehouse Architecture and Conventions: Design and organize data storage in Google BigQuery using best modeling practices. Ensure the correct transition and logical segregation of environments, from raw data reception (Data Swamp / Data Lake) to optimized availability layers (Publish Layers).
● Machine Learning Operationalization (MLOps): Actively support the Data Science stack, creating the technical bases necessary for autonomy, local testing, product packaging, deployment, and continuous monitoring of predictive models in a production environment.
● Automation and Infrastructure (CI/CD): Ensure the code lifecycle and that implementations are automated.
● Software Engineering Culture: Contribute to highly scalable, secure, and well-documented solutions. Promote and structure modern repositories in a Monorepo environment, facilitating collaboration and independence of Data Science teams.

3. Technical Profile and Candidate Requirements 

Mandatory Requirements (Hard Skills) 
● Consolidated experience in Data Engineering or Software Engineering focused on large-scale data systems.
● Practical experience in the cloud ecosystem GCP (Google Cloud Platform), with a specialized focus on BigQuery.
● Practical experience with containerization and orchestration tools: Docker and Kubernetes (K8s).
● Proven proficiency in designing automated CI/CD pipelines using preferably GitHub Actions (or equivalent market tools).
● Advanced practical knowledge in organizing and managing code in Monorepo architectures.
● Strong understanding of modern data architectures (Lakehouse, Data Lake) and their respective data layer conventions.
● Proficiency in English: Minimum intermediate-advanced level (e.g., B2 or C1) for communication in technical and international environments.

Valued Requirements (Differentials) 
● Direct practical experience in MLOps concepts and tools (model lifecycle management, model serving APIs, feature stores).
● Experience in deploying solutions in orchestrated clusters (Kubernetes), depending on architectural needs.
● Experience with Observability Using Logging tools like Datadog or Grafana 

4. Summary of Required Tech Stack 

Key Technologies / Concepts 

Cloud & Storage 
Google Cloud Platform (GCP), Google BigQuery 

Infrastructure & DevOps 
Kubernetes (K8s), Docker, Virtual Machines (VMs) 

CI/CD & Repositories 
GitHub Actions, Monorepo 

ML Methodologies 
MLOps (Operationalization and Deployment of Data Science Models) 

Data Architecture 
Data Swamp, Data Lake, Publish Layers, DW Conventions 

Languages 
English (B2 or C1)

🔍 Dekoder Ogłoszenia

🔴
serving as a critical bridge between model development and large-scale production
Może oznaczać, że będziesz odpowiedzialny za przenoszenie modeli z fazy eksperymentalnej do produkcyjnej, co często wiąże się z problemami skalowalności i stabilności.
🔴
handling complex scenarios such as advanced integration of third-party APIs and web scraping of application APIs
Może oznaczać konieczność radzenia sobie z niestabilnymi lub słabo udokumentowanymi API, co wymaga dużej elastyczności i umiejętności rozwiązywania problemów.
🔴
Ensure the correct transition and logical segregation of environments, from raw data reception (Data Swamp / Data Lake) to optimized availability layers (Publish Layers)
Określenie 'Data Swamp' może sugerować, że obecna infrastruktura danych jest chaotyczna i wymaga uporządkowania, co może być czasochłonne.
🔴
creating the technical bases necessary for autonomy, local testing, product packaging, deployment, and continuous monitoring of predictive models
Choć brzmi to jak wsparcie dla Data Science, może oznaczać, że będziesz musiał samodzielnie budować narzędzia i procesy, których brakuje, zamiast korzystać z gotowych rozwiązań.
🔴
Ensure the code lifecycle and that implementations are automated
Może oznaczać, że będziesz musiał zaimplementować i utrzymać całą infrastrukturę CI/CD od podstaw, a nie tylko korzystać z istniejących narzędzi.