Job
- Level
- Senior
- Job Field
- IT, DevOps
- Employment Type
- Part Time/Full Time
- Contract Type
- Permanent employment
- Location
- Berlin, Dusseldorf, Cologne, Darmstadt
- Working Model
- Full Remote, Hybrid, Onsite
Job Summary
In this role, you manage our Apache Kafka infrastructure, operate change data capture pipelines, and automate access management and infrastructure as code while closely collaborating with an international team of data engineers.
Job Technologies
Your role in the team
- You run and own our self-hosted Apache Kafka backbone and Schema Registry end-to-end, managing ACLs, client quotas, certificates, upgrades, and disaster recovery.
- You operate our Debezium-based change data capture pipelines, overseeing connectors, snapshots, offsets, replica lag, and BigQuery sink loaders.
- You run our workloads on Kubernetes, including Navarch - our self-built workflow orchestrator executing ~7,500 SQL and Python tasks daily - owning resource management, autoscaling, and debugging.
- You automate and own access management and infrastructure as code across our self-hosted GitLab CI/CD pipelines, replacing manual tickets with code-driven policies.
- You own observability and infrastructure costs - Datadog across Kafka and sink loaders, Slack-native alerting, and maintaining an honest view of K8s, Kafka retention, and log volume expenses.
- You carry your share of triage, one week in three - handling ingestion and access requests, schema alerts, and supporting analysts, while helping automate this layer via AI-assisted review.
- You act as a senior sparring partner alongside two senior data engineers (reporting to the Head of Data & Analytics) - driving design reviews and multi-quarter initiatives, with room to grow into BigQuery architecture over time: access frameworks, partitioning and slot strategy, the pipelines and models on top.
- Optional: You participate in our 24/7 on-call rotation.
This text has been machine translated. Show original
Our expectations of you
Qualifications
- You have deep hands-on expertise in Apache Kafka broker operations, including cluster sizing, partition rebalancing, replication, ISR behavior, and version upgrades.
- You have operated cross-system data replication in production - ideally Debezium on Kafka Connect, or database replication via binlog, GTID, WAL, or agent-based pipelines.
- You use production-grade Python and SQL as practical tools, applying a cost-conscious mindset to query performance, scanning volume, and data pruning.
- You thrive in a low-process, Kanban-driven environment where seniors self-organize, rotate weekly triage duties without micro-management, and prefer automating recurring requests over manually repeating them.
- Optional: You bring skills or interest in object-oriented design, data modeling/dbt, BigQuery architecture, orchestration frameworks (Airflow, Dagster), stream processing (Beam, Flink, Spark), or internal automation (n8n, AI triage bots).
- You are fluent in English (C1 level) and enjoy working in an international team environment.
Experience
- You have several years of experience operating production infrastructure you were personally responsible for, including on-call rotation. You have run systems that broke, can explain what went wrong and how you permanently fixed it, and can detail a specific trade-off you navigated backed by concrete numbers.
- You demonstrate strong production experience with Kubernetes, master resource management and autoscaling, and possess sharp debugging reflexes for complex failure modes.
- You manage infrastructure and access as code across multiple environments (ideally Terraform and Ansible; Pulumi, Helm or serious configuration management also counts) following a strict least-privilege IAM approach across multiple environments, ideally with GCP experience or transferable AWS/Azure knowledge.
This text has been machine translated. Show original
What we offer
- Create your own work-life balance: You have the flexibility to choose between working remotely (within Germany) or from one of our locations in Cologne, Darmstadt, Düsseldorf, or in Berlin!
- Möchten Sie nach Deutschland ziehen? Kein Problem - wir bieten Ihnen ein attraktives Umzugspaket, um Ihnen einen reibungslosen Start zu ermöglichen.
- Urban Sports Club and RSG Group Fitness Studios: Get top deals for fitness, swimming, yoga and more.
- Mental Well-Being: We support you on your well-being journey with special offerings such as Instahelp, the digital platform for online psychological counseling.
- Vacation & Sabbatical: Enjoy 30 days of vacation per year and the opportunity to take a sabbatical once you have been part of the team for a certain period!
- Option for Pluxee restaurant vouchers: Buy Pluxee vouchers through us and benefit from tax-advantaged meal allowances!
- 'Germany Ticket': We subsidize your train season ticket for more mobility.
- Employee Discount: You will receive a monthly coupon for Kaufland.de.
- Free choice of operating system: MacOS or Ubuntu Linux, it's up to you.
- Boost your growth: Benefit from our online language learning programs, diverse in-house training, and our automated 360-degree feedback. We cover the costs for relevant conferences, training opportunities, and approved team workshops to strengthen personal interactions.
- This is who we are: Our dynamic culture combines flat hierarchies, a start-up mentality, an international team of over 65 nationalities, and the strength of the Schwarz Group to provide you an agile and secure working environment.
This text has been machine translated. Show original
Benefits
Work-Life-Integration
Higher Take-Home Pay
Health, Fitness & Fun
Topics You Will Work On
Job Locations
About Your Employer
Kaufland Stiftung & Co. KG
Our business's strength lies in our branches. We're present nationwide with 670 locations. Our customers can choose from over 30,000 items. Make every shopping experience an enjoyable one for our customers.
Description
- Company Size
- 50-249 Employees
- Company Type
- Established Company
- Working Model
- Full Remote, Hybrid, Onsite
- Industry
- Trade