Software Golang Engineer (Slurm)
G-CORE INNOVATIONS SOCIETE A RESPONSABILITE LIMITEE
⚲ Kraków
20 000 – 30 000 zł / mies.
Wymagania
- Slurm
- Go
- Kubernetes
- PyTorch
Opis stanowiska
Nasze wymagania:
Hands-on experience using Slurm in production from a user’s perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo
Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops
Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes
Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon
Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges
Mile widziane:
Experience operating large-scale HPC or GPU clusters for external customers
Experience with PyTorch distributed training and other large-scale AI/ML frameworks
Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure
Experience building unified job-submission workflows across Kubernetes and Slurm
Experience in GPU-cloud or HPC product engineering environments
Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects
Zakres obowiązków:
Design and build a managed Slurm service on Kubernetes
Write clean, reliable, and maintainable Go code
Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads
Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, DCGM
Oferujemy:
Competitive compensation
Flexible working hours and hybrid or remote options, depending on your role
Work from anywhere in the world for up to 45 days per year
Private medical insurance for you and your family*
Extra paid vacation and sick leave days*
Support for life’s important moments and celebrations
Language courses to help you connect and grow
Modern, welcoming offices with snacks, drinks, and entertainment*
Team sports and social activities*
Hands-on experience using Slurm in production from a user’s perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo
Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops
Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes
Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon
Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges
Mile widziane:
Experience operating large-scale HPC or GPU clusters for external customers
Experience with PyTorch distributed training and other large-scale AI/ML frameworks
Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure
Experience building unified job-submission workflows across Kubernetes and Slurm
Experience in GPU-cloud or HPC product engineering environments
Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects
Zakres obowiązków:
Design and build a managed Slurm service on Kubernetes
Write clean, reliable, and maintainable Go code
Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads
Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, DCGM
Oferujemy:
Competitive compensation
Flexible working hours and hybrid or remote options, depending on your role
Work from anywhere in the world for up to 45 days per year
Private medical insurance for you and your family*
Extra paid vacation and sick leave days*
Support for life’s important moments and celebrations
Language courses to help you connect and grow
Modern, welcoming offices with snacks, drinks, and entertainment*
Team sports and social activities*
🔍 Dekoder Ogłoszenia
🔴
end-to-end ownership of complex distributed systems
Pełna odpowiedzialność za całość — w tym dyżury i awarie poza standardowymi godzinami pracy
🟡
product mindset and strong customer empathy
Może oznaczać bezpośredni kontakt z klientami i presję na szybkie rozwiązywanie ich problemów