Logo gridscale GmbH

Senior Site Reliability Engineer - Ceph / On-Prem Cloud Storage

New

Job

  • Level
    Senior
  • Location
    Cologne
  • Working Model
    Onsite
  • Job Field
    IT, System, DevOps
  • Employment Type
    Full Time
  • Contract Type
    Permanent employment

Job Summary

In this role, you will develop the storage architecture for an on-premise cloud platform, operate Ceph clusters, and optimize automation strategies using Kubernetes and OpenStack, ensuring high availability and performance.

Job Technologies

Your role in the team

  • You will help us build, operate and industrialize the storage foundation of our on-premise cloud platform.
  • As part of a small, experienced team, you will own the Ceph clusters behind our block, object, and file storage services - essentially the storage on which every customer workload ultimately runs.
  • The platform is still actively being built, which means you will have real influence on the storage architecture, hardware lifecycle, automation strategy and technical direction.
  • The role goes beyond Ceph itself: you will regularly work across bare metal, networking, OpenStack, Kubernetes and GitOps.
  • As a Senior Engineer, you take ownership of your area and help make storage predictable, durable and highly automated as we scale into the petabytes.

This text has been machine translated. Show original

Our expectations of you

Qualifications

  • You know the Ceph fundamentals from real-world operations: CRUSH, pools and placement groups, replication vs. erasure coding, BlueStore, scrubbing, recovery and backfill under load.
  • You have managed storage infrastructure end to end, from the physical layer with drives, firmware, BIOS, controllers and hardware diagnostics through rebalancing, host drains, capacity expansion and hardware generation migrations.
  • You are comfortable using Kubernetes as a control plane, not only as a workload runtime, and understand operators/controllers, custom resources, reconciliation and GitOps.
  • You can also move confidently around OpenStack and debug issues involving Ironic, Neutron, Nova or Glance when needed.
  • You have dealt with situations where durability really matters - such as degraded or near-full clusters, recovery under load or split-brain conditions - and you are used to making changes in a staged, evidence-based and reversible way.
  • You have strong Linux and bare-metal skills, understand the block layer, filesystems and I/O behaviour, and are comfortable working with Ansible, Terraform and GitOps.
  • AI-assisted engineering is already part of your day-to-day work.
  • You work autonomously, bring a strong sense of ownership, and are comfortable debugging problems where the actual root cause may sit somewhere between storage, networking, Kubernetes and OpenStack.
  • You would rather verify what is happening from evidence than assume how the system should behave.
  • We work in an international environment in English, so you should feel comfortable discussing technical topics, documenting decisions and working with the team in English.

Experience

  • Several years of hands-on experience as an SRE, Storage Engineer or Platform Engineer running production storage infrastructure, with solid experience deploying, operating and debugging Ceph in production.
  • You have practical experience with LLMs and agentic tools and know where they can meaningfully support development, testing, reviews or operations.
  • Ideally, you also bring experience with RGW / S3, RBD mirroring, CephFS or ceph-csi, deeper OpenStack storage integrations and Go and/or Python.
  • Experience with VLAN/BGP, storage performance tuning, NVMe/BlueStore, observability, data protection, encryption, auto-remediation, security baselines, or multi-site storage is also relevant for the role.

This text has been machine translated. Show original

What we offer

  • Exceptional team spirit across all departments and national borders; we live #OneTeam.
  • Exciting work in a highly innovative and international environment with cutting-edge technologies.
  • 32 vacation days, increasing with length of service.
  • Flexible working hours, home office options, and a secure permanent position with market- and performance-based compensation.
  • Employer-funded pension plan and an attractive insurance package.
  • OVHcloud covers 50% of public transportation costs.
  • Up to €400 annual financial contribution from OVHcloud towards sports activities (gym membership, sports classes, etc.).
  • Through Corporate Benefits, you receive attractive discounts at numerous shops and companies.
  • We contribute to the leasing of your cargo bike.
  • Regular company events and free cold and hot beverages.

This text has been machine translated. Show original

Benefits

Work-Life-Integration

Topics You Will Work On

Job Locations

  • Location Cologne

    Nordrhein-Westfalen

    Germany

About Your Employer

gridscale GmbH

gridscale GmbH

gridscale ist ein europäischer IaaS- und PaaS-Anbieter und schafft mit seiner innovativen Technologie die Basis für anspruchsvolle Cloud-Lösungen. Das Unternehmen mit Hauptsitz in Köln bietet zukunftsorientierten und digital ausgerichteten Nutzern mit seiner hochautomatisierten Architektur eine Lösung, in der sie flexibel zwischen einer Vielzahl an Infrastructure-as-a-Service-Komponenten sowie komplementären Platform-as-Service-Elementen wählen können.

Description

  • Founding Year
    2014
  • Company Type
    Established Company
  • Working Model
    Hybrid, Onsite
  • Industry
    Internet, IT, Telecommunication
Logo gridscale GmbH

Senior Site Reliability Engineer - Ceph / On-Prem Cloud Storage

Location
Cologne
Working Model
Onsite
Diversity
Open for all genders
English Only
English only required

More Jobs