Job
- Level
- Senior
- Location
- Munich
- Working Model
- Hybrid, Onsite
- Job Field
- IT, Software, DevOps
- Employment Type
- Full Time
- Contract Type
- Permanent employment
- Salary
- 75.000 to 90.000€ gross/year
Job Summary
In this role, you will develop automated solutions for monitoring and enhancing the production pipeline, create reliable alerting systems, and build AI agents that proactively diagnose and resolve issues.
Job Technologies
Your role in the team
- Our crawlers collect regulatory updates from 80+ sources and push them through an automated pipeline: extraction, embedding, search, classification, consolidation and translation. As we add regions, keeping it healthy has become a job of its own.
- We want an engineer who makes production tell us what's wrong before customers do, fixes what can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach a developer, it should arrive with a root cause and a suggested fix.
- Not a ticket-driven ops role: you'll write production Python from week one and own platform reliability alongside a core-team developer.
- Reduziere das Rauschen. Trenne vorübergehende Fehler (Netzwerkaussetzer, Timeouts, Fehler, die beim erneuten Versuch verschwinden) von echten Fehlern und entwickle eine Alarmierung, auf die sich das Team verlassen kann.
- Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right developer.
- Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. 'every source checked on time', 'every document produced text and embeddings'.
- Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
- Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and measuring how often they're right.
- Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
- Catch memory, timeout and cost problems across Cloud Run before they become silent OOM kills.
- Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
This text has been machine translated. Show original
Our expectations of you
Qualifications
- A track record of turning an ignored alert channel into one people act on.
- Solid GCP (or similar), containers and CI/CD.
- Starkes Verständnis der Ausfallmodi verteilter Systeme: Wiederholungen, Idempotenz, Teilfehler, Backpressure.
- Pragmatism and clear communication within a small team.
Experience
- 5+ years in software engineering, SRE or platform roles, with real production Python.
- Hands-on experience building with LLMs or agents in production, and judgment about when a plain if-statement is better.
This text has been machine translated. Show original
What we offer
- Hybrid work culture: join us in our Munich office (min. 2 days/week).
- Flexible working hours.
- 26+4 vacation days per year (4 fixed 'company rest days' over Christmas).
- 30 days of 'workation' per year, within the EU and selected countries.
- High autonomy and flat hierarchies.
- EGYM Wellpass for unlimited access to fitness courses and gyms.
- Udemy-Zugang für Bildungsinhalte.
This text has been machine translated. Show original
Benefits
Work-Life-Integration
Topics You Will Work On
Job Locations
About Your Employer
Certivity
Certivity is a RegTech startup based in Munich that develops an AI-powered platform to support companies in regulated industries. This platform transforms regulatory documents into structured information and assists teams in monitoring regulatory changes. Since its founding in 2021, the company has gained over 15,000 users in 10 countries.
Description
- Company Type
- Startup
- Industry
- Internet, IT, Telecommunication