Job
- Level
- Senior
- Location
- Berlin
- Working Model
- Hybrid, Onsite
- Job Field
- Back End
- Employment Type
- Full Time
- Contract Type
- Permanent employment
Job Summary
In this role, you will develop solutions for our Ceph-based object storage platform, collaborate closely with the operations team, and optimize the availability, performance, and security of our systems.
Job Technologies
Your role in the team
- Staff Storage Reliability Engineer for our global public Object-Storage platform based on Ceph.
- You develop solutions that scale in production and collaborate with our operations team to build, deploy, and maintain the platforms.
- We are currently in the double-digit petabyte range, distributed across multiple locations.
- Our Ceph platform is growing rapidly and is a critical component of our internal and public infrastructure.
- You will actively participate in the further development, improve and maintain our production environments, and ensure that availability, performance, and security are maintained during scaling.
- Provision, configure, and operate Ceph clusters - including the management of storage pools, placement groups, and other core components.
- Tuning Ceph for optimal performance, eliminating bottlenecks, and ensuring efficient resource utilization.
- Develop and implement automation strategies for Ceph deployments, upgrades, and maintenance tasks.
- Diagnose and resolve complex technical issues related to Ceph storage, often in collaboration with other teams.
- Collaborate closely with development teams, system administrators, and other stakeholders to integrate Ceph into various systems and applications.
- Follow the latest Ceph developments, new features, and best practices.
- Actively participate in the Ceph community and share knowledge.
This text has been machine translated. Show original
Our expectations of you
Qualifications
- In-depth knowledge of Ceph architecture and administration.
- Experience with automation tools (e.g., Ansible) as well as monitoring and observability solutions.
- Familiarity with cloud platforms and container technologies (e.g., Docker).
- Excellent troubleshooting and problem-solving skills, strong communication and collaboration abilities.
Experience
- 5+ years of experience as a Senior Linux Engineer or Site Reliability Engineer; deep and broad understanding of Linux systems and networks.
- Experience with cloud storage technologies (File, Object, Block).
This text has been machine translated. Show original
What we offer
- Hybrid work model.
- Flexible working hours through trust-based working time.
- At some locations, a subsidized canteen and various free beverages.
- Modern office spaces with excellent transportation links.
- Various employee discounts for activities and products.
- Employee events such as summer and winter parties, as well as workshops.
- Numerous opportunities for further training and development.
- Various health offerings, such as sports and wellness courses.
This text has been machine translated. Show original
Benefits
Health, Fitness & Fun
Work-Life-Integration
Food & Drink
Topics You Will Work On
Job Locations
About Your Employer
1&1 Internet AG
With our strong brands 1&1, GMX, WEB.DE and mail.com, we are the leading provider of consumer applications in Germany with over 30 million active users. We make communication even safer - with up to 500 million incoming e-mails per day! With our advanced security facilities, we ensure that your data is always protected. So you can relax and concentrate on the really important things.
Description
- Company Size
- 250+ Employees
- Company Type
- Established Company
- Working Model
- Full Remote, Hybrid, Onsite
- Industry
- Internet, IT, Telecommunication
Employer reviews
by devworkplaces.com
Total
(1 Review)Career Growth
3.4Workingconditions
4.4Engineering
2.7Culture
3.5