Logo Cloudfactory GmbH

Senior Site Reliability Engineer

Job

  • Level
    Senior
  • Job Field
    IT, DevOps, Back End
  • Employment Type
    Full Time
  • Contract Type
    Permanent employment
  • Location
    Berlin
  • Working Model
    Hybrid, Onsite
  • Job Summary

    In this role, you monitor production systems, optimize ML and LLM infrastructure, and develop advanced solutions for automation, security, and reliability of platforms in a global environment.

    Job Technologies

    Your role in the team

    • As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly.
    • You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.
    • The SRE team owns the foundation of AI Platform's Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features.
    • We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.
    • Reliability of platform (includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services.
    • Observability (including ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, integrated into the same metrics and logging stacks we use everywhere else.
    • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team.
    • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster.
    • Wiederverwendbare Komponenten, die gängige Open-Source-Tools (Grafana, Istio, CloudNative-Stack und ML-Tools wie Model-Registries und Feature Stores) verpacken, damit Teams sie in jeder Umgebung bereitstellen können.
    • Secure-by-default infrastructure - integrating security, compliance audits, cost governance, and audit trails into the platform in close collaboration with our lead/backend/staff engineers.

    This text has been machine translated. Show original

    Our expectations of you

    Qualifications

    • Good proficiency in Python or Go or general scripting for automation and tooling (automation with higher language preferred).
    • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with.
    • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off).
    • Mindestens ein Infrastrukturaufbau, den du vollständig eigenständig durchgeführt hast - mit der zugehörigen Ergebniskennzahl (Bereitstellungszeit, MTTR, Kosten, Akzeptanz, Verfügbarkeit).
    • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
    • Ausführung von ML-Workloads auf Kubernetes - GPU-Planung, Kapazitäts- und Kostenmanagement.
    • Model serving and inference at production scale (e.g., KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints (preferably RayServe).
    • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents).
    • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request).
    • Governing ML/LLM workloads as platform capabilities: data residency and PII controls, and audit trails.
    • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
    • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
    • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
    • Availability: Willingness to support processes for 24x7 operational support.

    Experience

    • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes.
    • Production operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or CloudFormation, on at least one major cloud (AWS preferred).

    This text has been machine translated. Show original

    What we offer

    • At CloudFactory, we believe that work should be more than just a job - it should be a platform for growth, impact, and community.
    • Hier wirst du mit Sinnhaftigkeit verdienen, jeden Tag lernen und einer Mission dienen, die wirklich zählt.
    • If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we'd love to have you on this journey!

    This text has been machine translated. Show original

    Benefits

    Work-Life-Integration

    Topics You Will Work On

    Job Locations

    • Location Berlin

      Germany

    About Your Employer

    Cloudfactory GmbH

    Cloudfactory GmbH

    CloudFactory GmbH, a subsidiary of the international company CloudFactory, focuses on IT services, including consulting, project management, and system engineering. With multiple locations in Germany and Switzerland, it provides comprehensive solutions in the IT sector.

    Description

  • Company Type
    Established Company
  • Working Model
    Hybrid, Onsite
  • Industry
    Internet, IT, Telecommunication
  • Logo Cloudfactory GmbH

    Senior Site Reliability Engineer

    Location
    Berlin
    Working Model
    Hybrid, Onsite
    Diversity
    Open for all genders
    English Only
    English only required

    More Jobs