Doctolib
Berlin, DE · On-site · Full-time
<h3><strong>Your Impact</strong></h3> <p>We are looking for a <strong>Staff Site Reliability Engineer</strong> to join our SRE team dedicated to platform reliability within Platform Engineering.</p> <p>Your mission will be to act as a technical leader driving Doctolib's reliability and scalability at a European scale, ensuring our platform remains reliable, debuggable, and resilient across infrastructure, observability, and cross-cutting reliability initiatives. You will play a pivotal role in a team driving reliability standards across 170+ applications, contributing directly to supporting 520,000 health professionals and 90 million patients in their daily healthcare journey.</p> <p>This role sits at the intersection of infrastructure, developer experience, and product engineering. You'll act as a technical leader and strategic partner to SREs, software engineers, and product teams, guiding decisions, mentoring engineers, and driving cross-cutting initiatives that elevate our operational maturity.</p> <h3><strong>What you'll do</strong></h3> <p>Your responsibilities include but are not limited to:</p> <ul> <li>Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management</li> <li>Identify and drive improvements to incident detection, response, and postmortem analysis capabilities</li> <li>Define and evolve SLOs, error budgets, and alerting standards across multiple product teams</li> <li>Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry</li> <li>Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability engineering across the company</li> <li>Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions</li> <li>Partner with software engineering teams to embed reliability practices early in the development lifecycle</li> </ul> <h2><strong>Who you are</strong></h2> <p>Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.</p> <p><strong>You'll be a great fit if you:</strong></p> <ul> <li>Have extensive experience (8+ years) in SRE, platform engineering, or infrastructure roles within a large-scale, multi-team production environment</li> <li>Have proven experience with cloud platforms such as AWS, GCP, or Azure</li> <li>Have strong experience with containerization and orchestration
Doctolib
Posted via Arbeitnow
Apply Now takes you to Rozgoo, where auto-apply can submit your application for this role. Updated today.
Apply Now