Sr. Site Reliability Engineer - SRE
À distance ✓ENWe are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance of our mission-critical services. This is an opportunity to shape our SRE practices, drive automation, reduce operational toil, and significantly impact our product's operational excellence. What You'll Do • Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experiences. • Serve as a subject matter expert for observability, including monitoring, alerting, logging, tracing, dashboards, and synthetic testing. • Develop robust, maintainable software and self-service tooling to automate operational tasks and improve reliability. • Identify and eliminate operational toil through automation, process improvements, and systematic problem solving. • Lead incident response, participate in on-call rotations, and drive blameless post-mortems. • Define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. • Leverage infrastructure as code, GitOps practices, and CI/CD automation using Terraform, Flux, and GitHub Actions. • Provide reliability expertise during system design reviews and influence architectural decisions. • Document processes, build runbooks, and mentor engineers across the organisation. • Leverage AI responsibly to accelerate investigations, improve documentation, reduce toil, and build intelligent operational workflows while maintaining appropriate human oversight, security, and governance. What You'll Bring Core SRE Capabilities • Demonstrated experience operating and improving production systems at scale in an SRE, Production • Engineering, or Platform Engineering role. • Ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains. • Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis. • Experience defining and using SLIs, SLOs, and error budgets to guide reliability decisions. • Excellent written and verbal communication skills. Technical Domains Experience across several of the following areas: • Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio. • Cloud infrastructure and services within AWS. • Identity and access management systems, including Auth0 and AWS IAM.\ • Networking fundamentals, including DNS, load balancing, routing, TLS, and connectivity troubleshooting. • GitOps workflows and infrastructure automation using tools such as Flux and Terraform. • Observability platforms and practices, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring. • CI/CD systems and engineering workflows. • Application logging and distributed system debugging. • Engineering Mindset A strong SRE: • Prioritizes service stability and customer impact during incidents. • Slows down under pressure, gathers facts, and communicates clearly. • Reduces operational complexity through automation and simplification. • Identifies and eliminates toil through self-service tooling and process improvement. • Demonstrates strong scripting and automation instincts. • Brings a systems-thinking approach to problem-solving. • Balances short-term remediation with long-term reliability improvements. Software Engineering for Reliability • Demonstrated ability to build and maintain automation, tooling, and self-service capabilities using one or more programming or scripting languages such as Python, Go, or Bash. • Focuses on applying software engineering practices to improve reliability, reduce toil, and enhance developer productivity. Behavioral Expectations • Calm and effective during high-severity incidents. • Skilled at managing complex situations involving multiple teams and competing priorities. • Able to lead blameless post-mortems and drive meaningful follow-up actions. • Passionate about continuous improvement and fostering a culture of shared ownership. AI-Native Operations • Uses AI assistants effectively to accelerate troubleshooting, root cause analysis, and operational decision-making. • Validates AI-generated recommendations and understands when human judgement is required. • Applies AI to automate repetitive tasks, improve runbooks, and create self-service capabilities. • Identifies opportunities to build AI-assisted workflows for incident response, observability, and platform operations. • Continuously evaluates emerging AI capabilities to improve reliability, developer experience, and operational efficiency. • Uses AI to summarise incidents, analyse logs and metrics, accelerate scripting, support code reviews, and generate operational …
Responsabilités:
• Design, implement, and maintain highly available, scalable, and resilient systems that deliver exceptional customer experiences. • Serve as a subject matter expert for observability, including monitoring, alerting,…