Job
- Level
- Experienced
- Job Field
- IT, DevOps
- Employment Type
- Full Time
- Contract Type
- Permanent employment
- Location
- Karlsruhe
- Working Model
- Hybrid, Onsite
Job Summary
In this role, you will develop the infrastructure for container-based web services and ensure continuous optimization, monitoring, and troubleshooting in a highly available environment on Kubernetes.
Job Technologies
Your role in the team
- As a Site Reliability Engineer (SRE) in our Application Hosting Team, you form the technical backbone of our product platform for Managed Nextcloud, Nextcloud Workspace, IONOS GPT, and other web services that we operate on our Kubernetes platform.
- Together with experienced colleagues, you design new services and products that remain performant and fault-tolerant even under maximum load.
- Your main area of responsibility is the further development of the infrastructure/platform of our products, as well as the integration of new products/web services into our Kubernetes and cloud infrastructure.
- You are responsible for the stable and secure operation of our product platform.
- Your expertise is in demand when it comes to in-depth analyses and the optimization of our primarily containerized and Kubernetes-based application infrastructure.
- You live automation.
- With tools like Terraform, GitLab CI/CD, and ArgoCD, you provision and manage our entire infrastructure declaratively and reproducibly.
- You analyze and resolve complex issues in a distributed system landscape and work on the continuous improvement of our platform.
- You develop and maintain our monitoring, logging, and alerting solution (e.g., with Prometheus, Grafana, ELK Stack) to proactively identify bottlenecks and sources of errors.
This text has been machine translated. Show original
Our expectations of you
Qualifications
- You have a proactive, solution-oriented, independent working style and the ability to systematically analyze and sustainably resolve complex technical problems.
- Good German and English skills are required.
Experience
- You have several years of experience as a Site Reliability Engineer or in a related role (Linux System Administrator, Platform Engineer, DevOps Engineer, Full Stack Developer) in a Linux and Kubernetes environment.
- Excellent knowledge and several years of experience in using the Linux operating system, container technologies, and especially Kubernetes.
- You have experience with Infrastructure as Code (preferably Terraform), CI/CD pipelines (e.g., GitLab CI/CD or GitHub Actions), and in the use and application of Helm Charts.
- You are proficient in at least one programming or scripting language (e.g., Go, Python, Bash) to solve automation and monitoring tasks and may already have some experience in building operators.
- Experience with operating and troubleshooting highly available and distributed production environments, including monitoring, alerting, and log analysis of distributed applications (e.g., Prometheus, Grafana, FluentD, ELK, VictoriaMetrics, Icinga).
This text has been machine translated. Show original
What we offer
- Hybrid work model.
- Flexible working hours through trust-based working time.
- At some locations, a subsidized canteen and various free beverages.
- Modern office spaces with excellent transportation links.
- Various employee discounts for activities and products.
- Employee events such as summer and winter parties, as well as workshops.
- Numerous opportunities for further training and development.
- Various health offerings, such as sports and wellness courses.
This text has been machine translated. Show original
Benefits
Health, Fitness & Fun
Work-Life-Integration
Food & Drink
Topics You Will Work On
Job Locations
About Your Employer
1&1 Internet AG
With our strong brands 1&1, GMX, WEB.DE and mail.com, we are the leading provider of consumer applications in Germany with over 30 million active users. We make communication even safer - with up to 500 million incoming e-mails per day! With our advanced security facilities, we ensure that your data is always protected. So you can relax and concentrate on the really important things.
Description
- Company Size
- 250+ Employees
- Company Type
- Established Company
- Working Model
- Full Remote, Hybrid, Onsite
- Industry
- Internet, IT, Telecommunication
Employer reviews
by devworkplaces.com
Total
(1 Review)Career Growth
3.4Workingconditions
4.4Engineering
2.7Culture
3.5