Bulldogjob Praca zdalna Mid

GPU Software Engineer (HPC / Deep Learning Optimization)

Luxoft DXC

Do uzgodnienia

Wymagania

  • C++
  • OpenCL
  • HPC
  • Deep Learning

Opis stanowiska

We are looking for a Software Engineer focused on GPU computing, HPC workloads, and Deep Learning inference optimization on Windows platform. The project is aimed at improving performance and efficiency of GPU-based workloads, including compute kernels and inference pipelines. The role is not limited to graphics APIs and is suitable for candidates with strong experience in CUDA, OpenCL, or similar technologies, as well as shader-based optimization.

Develop, optimize, and maintain GPU compute kernels using C++ and a GPU programming framework (CUDA, HIP, OpenCL, SYCL, DirectCompute / HLSL compute shaders, Metal compute, or equivalent).

Profile GPU workloads and tune memory, compute, and latency to improve performance and efficiency.

Analyze performance bottlenecks and apply targeted optimizations.

Debug and resolve performance and stability issues.

Apply kernel optimization to HPC or Deep Learning inference pipelines where needed.

Collaborate with engineers, QA, and stakeholders.

Follow coding standards and contribute to technical documentation.

🔍 Dekoder Ogłoszenia

🔴
The project is aimed at improving performance and efficiency of GPU-based workloads, including compute kernels and inference pipelines.
Może to oznaczać, że obecne rozwiązania są dalekie od optymalnych i wymagają znaczących usprawnień, a nie tylko drobnych poprawek.
🔴
The role is not limited to graphics APIs and is suitable for candidates with strong experience in CUDA, OpenCL, or similar technologies, as well as shader-based optimization.
Chociaż wymieniono wiele technologii, faktyczne zapotrzebowanie może być skoncentrowane na jednej lub dwóch, a reszta to tylko "nice to have".
🔴
Profile GPU workloads and tune memory, compute, and latency to improve performance and efficiency.
Może to oznaczać, że będziesz musiał spędzić dużo czasu na analizie i mikrooptymalizacji, co może być żmudne.
🔴
Apply kernel optimization to HPC or Deep Learning inference pipelines where needed.
Sformułowanie "where needed" może sugerować, że optymalizacja będzie reaktywna, a nie proaktywna, i może być stosowana tylko w przypadku zgłaszanych problemów.