Hybrid job
HybridMgr. SRE
LivePerson
Key points from the posting
- Tech stack
- Google Cloud Platform (GCP)PythonTerraformAnsibleBashKubernetesHelmFluxCDGitLab CI/CD
- Seniority:
- Lead
- Languages:
- English
Read out of the job posting automatically
Our assessment
- 7 more open roles from this employer in our index. 6 of them fully remote.
This section only: calculated automatically by nomado24, from our own job index and our own reading of the posting text. Not stated by the employer.
Job description
LivePerson (NASDAQ:LPSN) is a leading customer engagement company, creating digital experiences powered by Curiously Human AI. Every person is unique, and our technology makes it possible for companies, including leading brands like HSBC, Orange, and GM Financial, to treat their audiences that way at scale. Nearly a billion conversational interactions are powered by our Conversational Cloud each month.
You'll be successful at LivePerson if you are excited to build something from the ground up. You excel by finding daily opportunities to grow at the same pace as the technology we're building, and you build partnerships that improve our business. Likewise, you're someone who sees feedback as a chance to learn and grow and believes decisions powered by data are the norm. You care about the wellbeing of others and yourself.
Overview:
LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences at scale. As the market leader in real-time intelligent customer engagement, we are a B2B SaaS company with 20 years of experience and the heart of a startup. We work day in and day out to help our customers live out our mission of creating lasting, meaningful connections with their customers.
The Cloud DevOps team at LivePerson is looking for an SRE Team Lead to lead a team of Site Reliability Engineers responsible for building, operating, and continuously improving highly reliable, scalable, and secure cloud infrastructure and services.
The ideal candidate is a strong technical leader who combines deep experience in Site Reliability Engineering and cloud infrastructure with a passion for developing people and building high-performing engineering teams. This role requires someone who can balance strategic thinking with hands-on technical leadership and is comfortable making decisions in complex and fast-changing environments.
As an SRE Team Lead, you will be responsible for the team's technical direction, execution, operational excellence, and engineering practices. You will work closely with engineering, product, security, networking, and other infrastructure teams to ensure our platforms and services meet the reliability, scalability, security, and performance expectations of our customers.
You will lead initiatives that improve system reliability, eliminate operational toil, strengthen automation and observability, and establish engineering best practices. You will also mentor engineers, provide technical guidance, support career development, and foster a culture of ownership, collaboration, continuous improvement, and operational excellence.
We will provide you with an environment where you can make a meaningful impact, develop talented engineers, and solve complex technical challenges at scale.
Role and Responsibilities:
- Lead, mentor, and develop a team of SREs, fostering a culture of ownership, collaboration, technical excellence, and continuous improvement.
- Provide technical leadership and direction for the design, implementation, and operation of highly available, scalable, secure, and resilient infrastructure and services.
- Own the team's technical roadmap and ensure alignment with broader engineering and business objectives.
- Plan and prioritize team initiatives, balancing new development, reliability improvements, technical debt, operational work, and business priorities.
- Design and maintain cloud infrastructure and services across cloud and hybrid environments, with a strong focus on Google Cloud Platform (GCP).
- Guide the development of automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and other modern DevOps/SRE technologies.
- Lead the design, deployment, and operation of Kubernetes-based platforms and workloads, including troubleshooting complex production issues.
- Establish and maintain GitOps-based deployment workflows using Kubernetes, Helm, and FluxCD.
- Drive the design, implementation, and continuous improvement of CI/CD pipelines using GitLab CI/CD.
- Establish and improve observability practices using metrics, logs, traces, dashboards, and alerting to provide actionable insights into system health and performance.
- Define and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics for critical services.
- Lead incident response and ensure effective handling of critical production incidents, including root cause analysis and follow-up corrective actions.
- Drive initiatives to reduce operational toil, eliminate recurring incidents, and improve the overall reliability and operational maturity of the platform.
- Partner with software engineering, security, networking, product, and other infrastructure teams to influence architecture and deliver reliable and secure solutions.
- Lead capacity planning, performance analysis, …
This role is provided by an external source. Applications are handled on the source website.
