Job
- Level
- Senior
- Location
- Hamburg, Stuttgart
- Working Model
- Hybrid, Onsite
- Job Field
- IT, DevOps, Security
- Employment Type
- Full Time
- Contract Type
- Permanent employment
Job Summary
In this role, you own the reliability of systems, automate manual processes, define SLOs, manage the Kubernetes cluster, and implement security measures to enhance system performance.
Job Technologies
Your role in the team
- You join the Platform team and own the reliability of the systems our AI agents run on.
- This is a reliability engineering role: you treat operations as a software problem, so where others run a manual procedure, you write the automation that makes it unnecessary.
- You define what "healthy" means in numbers, measure it, and hold the line on it in production.
- When something breaks, you bring the system back and then make sure it cannot break the same way twice.
- Define and own SLOs and SLIs for the platform and manage error budgets against them.
- Carry on-call, act as incident commander, and run blameless post-incident reviews that produce real follow-up.
- Run production readiness and capacity planning ahead of demand, not after the page fires.
- Run and harden the Kubernetes platform (Helm, GitOps, service mesh) and the cloud underneath it (Terraform, multi-region).
- Own observability: metrics, logs, and distributed tracing, so problems surface before users feel them.
- Eliminate toil through automation, self-healing systems, and automated remediation.
- Drive cost visibility and FinOps practice across cloud and LLM spend.
- Bake security into the platform: least-privilege access, secrets management, policy as code, and vulnerability management.
- Use coding agents to build automation, write infrastructure code, and reason through failure modes faster.
- Automate operational toil and incident workflows with AI in the loop.
- Collaborate with the product team to give real-world feedback on Blockbrain's own tools from an operator's perspective.
- Bleiben Sie neugierig auf aufkommende KI-Fähigkeiten und wenden Sie diese auf Plattform- und Zuverlässigkeitsarbeit an.
This text has been machine translated. Show original
Our expectations of you
Education
- Degree in computer science or a related field, or equivalent hands-on experience.
Qualifications
- Clear and calm under pressure.
- Can coordinate an incident and write a post-mortem others learn from.
- Arbeitet auf Englisch; Deutsch ist ein Plus.
- Behandelt Infrastruktur als Code und Betrieb als eine Software-Disziplin.
- Genuinely enjoys automating manual work away.
- Misst den Erfolg an Vorfällen, die nicht eingetreten sind.
- Behebt die Ursachen, nicht nur die Symptome.
- Plans capacity and reliability work ahead of demand and balances on-call, project work, and toil reduction.
- We care about what you can operate, not the certificate.
- Kubernetes, Terraform / IaC, CI/CD (e.g., GitHub Actions), observability (Prometheus, Grafana, distributed tracing), a major cloud (AWS, Azure, or GCP), scripting (Python, Go, or TypeScript), secrets management, and policy as code.
- Strong ownership, blameless culture, a bias toward automation, calm in incidents, and a security-by-default mindset.
Experience
- 5+ Jahre Erfahrung in DevOps, SRE oder Platform Engineering, idealerweise bei der Betreuung von skalierbaren SaaS-Produkten im Produktionsumfeld.
- Hands-on experience running Kubernetes in production is essential.
This text has been machine translated. Show original
What we offer
- Vollzeitstelle vor Ort in Stuttgart, Hamburg oder München (3 Tage pro Woche) mit flexiblen Arbeitszeiten.
- Germany Ticket, Wellpass fitness membership, access to the latest AI tools, and regular company and team off-sites.
- International team with exceptional talents.
- Steile Lernkurve in einem schnell wachsenden AI-Startup.
- A high level of personal responsibility and the freedom to actively shape processes.
- MacBook, iPhone, headset, and all the tools you need to perform at your best.
This text has been machine translated. Show original
Topics You Will Work On
Job Locations
About Your Employer
Blockbrain
Blockbrain focuses on automating document processes and enhancing knowledge management. The platform increases efficiency and optimizes the use of corporate knowledge.
Description
- Company Type
- Startup
- Working Model
- Hybrid, Onsite
- Industry
- Internet, IT, Telecommunication