Job
- Level
- Senior
- Job Field
- IT, DevOps, Back End
- Employment Type
- Full Time
- Contract Type
- Permanent employment
- Location
- Berlin
- Working Model
- Full Remote
Job Summary
In this role, you will develop cloud infrastructure using Go and manage Kubernetes clusters. You will optimize the platform, automate workflows, and enhance system observability to create a reliable and scalable solution.
Job Technologies
Your role in the team
- As a Senior Software Engineer - Platform and Reliability, you'll be a key contributor to the design and operation of our cloud-native infrastructure.
- You will build and scale a reliable, cost-efficient platform that enables fast iteration and secure delivery of our vector database solutions to customers around the world.
- This role sits at the intersection of software engineering and infrastructure.
- You will write production-grade Go code, manage Kubernetes at scale, and design internal tooling and systems that enable developer autonomy and operational excellence.
- Design, implement, and operate our cloud-native platform architecture.
- Build and maintain Kubernetes clusters and develop custom Kubernetes operators.
- Write production-grade Go code for platform services and automation tooling.
- Optimize cloud infrastructure (AWS/GCP/Azure) for performance, cost, and reliability.
- Improve observability across the platform through monitoring, logging, and alerting.
- Automate workflows and integrations across systems and tools.
- Collaborate with engineering, operations, and data teams to understand and support their infrastructure needs.
- Contribute to incident response, root cause analysis, and system hardening.
- Continuously improve performance, developer experience, and reliability at scale.
This text has been machine translated. Show original
Our expectations of you
Qualifications
- Strong proficiency in Go and Python, or deep expertise in one with willingness to work with both.
- Solid understanding of distributed systems and microservices architecture.
- Comfort participating in on-call rotations and managing production incidents.
- Proactive, ownership-driven mindset with strong communication skills.
- Vertrautheit mit Observability-Standards wie OpenTelemetry.
- Contributions to open-source projects.
Experience
- 5-7+ years of experience in platform engineering or SRE roles.
- Deep experience developing Kubernetes operators.
- Hands-on experience with cloud providers (AWS, GCP, or Azure).
- Experience with CI/CD, infrastructure-as-code, and automation best practices.
- Experience in a SaaS, database, or systems-level product company.
- Experience with Prometheus, Grafana, and service meshes (e.g., Istio, Linkerd).
This text has been machine translated. Show original
What we offer
- A remote-first, international team working on cutting-edge AI infrastructure.
- A competitive salary with additional perks.
- Flexible working hours and an asynchronous-friendly culture.
- High ownership and real impact.
- Open-source, engineering-driven culture.
- Choose your own laptop equipment.
This text has been machine translated. Show original
Topics that you deal with on the job
Job Locations
This is your employer
Qdrant
Qdrant Solutions GmbH, based in Berlin, is an advanced company developing an open-source vector search engine. This vector database engine serves as the infrastructure for modern AI applications and is utilized by prominent companies such as Canva and Deutsche Telekom.
Description
- Company Type
- Startup
- Working Model
- Full Remote
- Industry
- Internet, IT, Telecommunication