Job
- Level
- Senior
- Job Field
- Software, Data
- Employment Type
- Full Time
- Contract Type
- Permanent employment
- Location
- Berlin
- Working Model
- Onsite
Job Summary
In this role, you will optimize AI workloads on NVIDIA platforms, analyze performance metrics, resolve cluster issues, and assist customers in scaling high-performance solutions efficiently.
Job Technologies
Your role in the team
- Collaborating with NVIDIA's training framework developers and product teams to stay ahead of the latest features and help partners to adopt them effectively.
- Assisting with deployment, debugging, and improving the efficiency of AI workloads on extensive NVIDIA platforms.
- Benchmarking new framework features, analyzing performance, and sharing actionable insights with both customers and internal teams.
- Working directly with external customers to solve cluster performance and stability issues, identify bottlenecks, and implement effective solutions.
- Build expertise and guide customers in scaling workloads efficiently and reliably on the latest generation of NVIDIA GPUs.
- Contributing to Europe's Sovereign AI initiative by helping customers implement advanced resiliency features within AI training pipelines.
This text has been machine translated. Show original
Our expectations of you
Qualifications
- Strong programming skills in at least one of the following languages: C, C++, or Python.
- Solid understanding of CPU and GPU architectures, CUDA, parallel filesystems, and high-speed interconnects.
- Proficient knowledge of training pipelines and frameworks, encompassing their internal operations and performance attributes.
Experience
- BS, MS, PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering field—or equivalent practical experience.
- 8+ years of experience in accelerated computing technologies at cluster scale, ideally including work with NVIDIA platforms.
- Practical experience in identifying and resolving bottlenecks in large-scale training workloads or parallel applications.
- Hands-on experienced in profiling and debugging large parallel applications.
- Erfahren im Umgang mit großen Rechenclustern mit Verständnis für deren interne Scheduling- und Ressourcenmanagement-Mechanismen (z. B. SLURM oder Cloud-basierte Cluster).
This text has been machine translated. Show original
Topics that you deal with on the job
Job Locations
This is your employer
Nvidia
For the past two decades, NVIDIA has been constantly reinventing itself. The invention of the GPU by NVIDIA in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing.
Description
- Founding year
- 1999
- Language
- English
- Company Type
- Established Company
- Working Model
- Full Remote, Hybrid, Onsite
- Industry
- Trade