Job
- Level
- Experienced
- Job Field
- IT, System, Security
- Employment Type
- Full Time
- Contract Type
- Temporary employment
- Location
- Tübingen
- Working Model
- Hybrid, Onsite
Job Summary
In this role, you will design and operate high-performance computing clusters, develop security architectures, and automate provisioning within a challenging HPC environment for machine learning.
Job Technologies
Your role in the team
- Design and operate their HPC clusters across four data centers, including scheduler (SLURM), parallel filesystems, networks and accelerators, ensuring high availability and throughput for research workloads.
- Conceive and establish the security architecture of the Machine Learning Science Cloud and harden the HPC environment.
- Evolve the automated provisioning and configuration of heterogeneous compute, storage, and network nodes (e.g., image-based provisioning, node lifecycle).
- Run patch and vulnerability management - risk assessment across heterogeneous systems.
- Build and operate logging, monitoring and intrusion detection, and integrate HPC telemetry into both operational dashboards and incident-response workflows.
- Lead incident response for the clusters: detection, containment, forensic support and post-incident review.
- Automate operations and security policy as code (Ansible/IaC).
- Advise researchers on efficient, secure cluster usage (job scheduling, data handling, access workflows) and derive requirements for our further roadmap from their scientific workloads.
This text has been machine translated. Show original
Our expectations of you
Education
- Master's degree in Computer Science or a related field.
Qualifications
- In-depth IT security knowledge: system hardening, network security, applied cryptography and IAM - with the ability to derive architectural decisions from a threat model, not only to apply given baselines.
- Hands-on HPC background: Slurm, parallel file systems (Weka, Lustre, Ceph), GPU workloads and high-speed networks (InfiniBand, 400G Ethernet).
- A plus: security frameworks (ISO 27001, BSI Grundschutz), container security (Apptainer/Singularity/Docker) or offensive-security fundamentals.
- Independent, structured working style and good communication in English; German is a plus.
- A collaborative, user-facing mindset - comfortable supporting and advising researchers and translating their needs into platform design.
Experience
- Solid Linux experience in production environments (RHEL/Almalinux/Ubuntu).
- Experience with virtualization for management-plane and infrastructure services (Proxmox).
- Strong scripting and automation skills (Bash, Python, Ansible) and experience with configuration management / Infrastructure-as-Code.
This text has been machine translated. Show original
What we offer
- Technically deep, architecturally open work that directly enables cutting-edge machine-learning research.
- Flexible working hours and the option to work partially from home.
- Working in an English-speaking, international team of HPC experts.
- A small, senior team with a flat hierarchy where responsibilities are divided by domain.
- Ownership of a technical domain in a production environment of real scale.
- Professional development, conference attendance and real influence on our roadmap.
This text has been machine translated. Show original
Topics You Will Work On
Job Locations
About Your Employer
Cyber Valley GmbH
Cyber Valley GmbH serves as the central management and coordination unit of the Cyber-Valley Innovation Campus in the Stuttgart/Tübingen region. It provides a platform for networking, events, and programs related to artificial intelligence and robotics, supporting the development and visibility of the ecosystem.
Description
- Company Type
- Established Company
- Working Model
- Hybrid, Onsite
- Industry
- Internet, IT, Telecommunication