Job
- Level
- Senior
- Job Field
- IT, DevOps, Back End
- Employment Type
- Full Time
- Contract Type
- Freelancer
- Location
- Stuttgart
- Working Model
- Hybrid, Onsite
Job Summary
In this role, you will develop an automated Kubernetes platform on bare metal to support mission-critical applications, managing its operations, high availability, as well as integrated observability and security measures.
Job Technologies
Your role in the team
- Technical specification of platform requirements and translation into implementable solutions in coordination with the Platform Domain Architect and the involved specialist teams.
- Development and further enhancement of a largely automated and reproducible Kubernetes-based platform on bare-metal infrastructure for mission-critical applications up to operational readiness and go-live.
- Implementation of reusable platform functions for the provisioning, configuration, and further development of Kubernetes clusters, servers, and other infrastructure components.
- Development, implementation, automated testing, and integration of Kubernetes controllers and operators for cluster and application deployment, applications without native GitOps mechanisms, as well as server and network configuration.
- Further development of automated Bare-Metal and Kubernetes deployment, including hardware and operating system lifecycle, cluster provisioning, and integration of BMC interfaces.
- Further development of GitOps- and Infrastructure-as-Code-based automation solutions, including the necessary testing and integration procedures to reduce manual interventions.
- Integration of external systems and infrastructure components via APIs, declarative interfaces, and controller-based mechanisms.
- Design and implementation of observability for platform and applications with metrics, logs, traces, and alerting, as well as cluster-internal and cross-cluster views and role-based access control.
- Design and implementation of automated backup, restore, and recovery procedures for platform components, persistent volumes, and databases, including external backup targets, verifiable recovery processes, and specific failure scenarios.
- Design and implementation of high availability measures and defined system behavior during failures to improve stability, recoverability, traceability, and security across cluster and protection zone boundaries.
- Design and technical implementation of security policies through Admission Control, policy-based validation and mutation of resources, as well as additional hardening and threat detection measures.
- Planning and implementation of technological advancements of existing platform components to meet new platform requirements.
- Analysis of complex error patterns across multiple technical layers, as well as the execution and support of restart and functionality tests.
- Support for stable and secure platform operation, including incident and fault management, restoring normal operations, operational changes and adjustments, backup, restore, and recovery procedures, go-live and hypercare phases, as well as technical usability during ongoing operations.
- Execution of upgrades, releases, patches, and other operational lifecycle activities as well as maintenance and adjustment of existing configurations.
- Creation and handover of technical and operational documentation as well as conducting knowledge transfer and operational support.
This text has been machine translated. Show original
Our expectations of you
Education
- Completed degree in Computer Science, Engineering, or a comparable qualification.
Qualifications
- Native-level German and business-fluent English skills, both written and spoken.
- Willingness to undergo security clearance according to SÜG and the corresponding pre-screening.
- At least one personal reference within the DACH region and at least one personal reference related to Platform Engineering for a productive Kubernetes-based infrastructure.
- Availability from October or November 2026.
- Willingness to use Atlassian Confluence and Jira as prescribed collaboration and documentation tools.
Experience
- At least 6 years of relevant operational experience in a senior platform engineering role for Kubernetes-based platforms.
- Experience in designing, technically implementing, integrating, and significantly developing Kubernetes-based platforms.
- Practical experience with automated and reproducible deployment as well as lifecycle management of bare-metal systems, servers, and Kubernetes clusters.
- Experience with Bare-Metal and Kubernetes automation as well as with GitOps, Infrastructure as Code, and declarative or controller-based automation mechanisms.
- Extensive hands-on experience in the development or significant enhancement of Kubernetes controllers and/or operators that have been deployed in production.
- At least 3 years of experience with Kubernetes Day-2 operations, lifecycle management, upgrades, as well as the analysis and resolution of complex technical issues.
- Experience in designing and implementing observability for platforms and applications, including metrics, logging, tracing, alerting, as well as cluster-internal and cross-cluster views.
- Experience in designing, automating, and practically testing high availability, backup, restore, and recovery procedures, including defined failure scenarios and recovery tests.
- Experience with the design and technical implementation of Kubernetes security, policy enforcement, admission control, policy-based validation and mutation of resources, as well as other hardening and threat detection measures.
- Experience with integrating external systems and technical interfaces via APIs, declarative interfaces, and similar mechanisms.
- Extensive technical experience as well as the ability to independently handle complex and not fully specified tasks across multiple technical layers up to a deployable solution and assume responsibility for the results.
- Experience in the energy industry, critical infrastructure, or a comparable regulated or safety-critical environment.
This text has been machine translated. Show original
What we offer
- The role is approximately 90% remote; in-person days take place in the Stuttgart area.
This text has been machine translated. Show original
Benefits
Work-Life-Integration
Topics You Will Work On
Job Locations
About Your Employer
Wavestone
Wien
Wavestone is a renowned consulting firm focused on the strategic transformation of organizations. It offers comprehensive consulting services that encompass both industry-specific and transversal expertise.
Description
- Company Type
- Established Company
- Working Model
- Hybrid, Onsite
- Industry
- Consulting