JustJoin.IT Praca zdalna Senior

👉Senior MLOps Engineer (AI/ML Platform)

Xebia sp. z o.o.

⚲ Wrocław

Do uzgodnienia

Wymagania

  • GCP
  • VertexAI
  • Terraform
  • Python
  • TensorFlow
  • PyTorch
  • MLflow

Opis stanowiska

🟣 You will be:
• designing, building, and managing AI/ML platform infrastructure on Google Cloud Platform (GCP), leveraging Vertex AI services,
• building and operating ML pipelines, covering model training, evaluation, deployment, and lifecycle management,
• provisioning and automating cloud infrastructure using Terraform and integrating it into CI/CD pipelines,
• developing and maintaining CI/CD workflows with GitHub Actions and automating ML training, deployment, and retraining processes,
• developing production-grade Python solutions and designing secure, scalable REST APIs,
• implementing monitoring and observability for ML models, including performance tracking, data drift detection, and system health,
• collaborating with hybrid cloud and HPC environments to support GPU-based training and large-scale ML workloads.

🟣 Your profile:
• proven experience in MLOps / ML Platform Engineering in production environments,
• hands-on expertise with: GCP (Vertex AI) and cloud-native ML workflows and Terraform (GCP provider) for infrastructure as code,
• strong experience with: Python (3.10+), ML frameworks (TensorFlow, PyTorch, scikit-learn) and MLOps tools (MLflow, KFP SDK),
• experience building secure, production-grade APIs using FastAPI or Flask,
• solid understanding of: CI/CD pipelines (GitHub Actions), identity & access management (OAuth 2.0, WIF),
• experience with Slurm and HPC environments, including GPU scheduling,
• strong knowledge of: ML lifecycle best practices and distributed training and model deployment patterns.

🟣 Recruitment Process:

CV review – HR call – Interview – Client Interview – Decision

🎁 Benefits 🎁

✍ Development:
• development budgets of up to 6,800 PLN,
• we fund certifications e.g.: AWS, Azure,
• access to Udemy, O'Reilly (formerly Safari Books Online) and more,
• events and technology conferences,
• technology Guilds,
• internal training,
• Xebia Upskill.

🩺 We take care of your health:
• private medical healthcare,
• multiSport card - we subsidise a MultiSport card,
• mental Health Support.

🤸‍♂️ We are flexible:
• B2B or employment contract,
• contract for an indefinite period.

🔍 Dekoder Ogłoszenia

🔴
designing, building, and managing AI/ML platform infrastructure on Google Cloud Platform (GCP), leveraging Vertex AI services
Oczekuje się, że będziesz samodzielnie projektować i wdrażać całą infrastrukturę platformy ML, a nie tylko korzystać z gotowych rozwiązań.
🔴
building and operating ML pipelines, covering model training, evaluation, deployment, and lifecycle management
Będziesz odpowiedzialny za cały cykl życia modelu ML, od stworzenia po utrzymanie w produkcji, co może być bardzo czasochłonne.
🔴
provisioning and automating cloud infrastructure using Terraform and integrating it into CI/CD pipelines
Oczekuje się głębokiej znajomości Infrastructure as Code (IaC) i umiejętności tworzenia złożonych, zautomatyzowanych procesów.
🔴
developing production-grade Python solutions and designing secure, scalable REST APIs
Oczekuje się nie tylko skryptowania, ale tworzenia solidnych aplikacji w Pythonie z naciskiem na bezpieczeństwo i skalowalność.
🔴
collaborating with hybrid cloud and HPC environments to support GPU-based training and large-scale ML workloads
Może oznaczać pracę z różnymi, potencjalnie skomplikowanymi i niejednorodnymi środowiskami, wymagającymi elastyczności i rozwiązywania problemów.