
Opportunity at
LambdaSenior Incident Manager
Lambda is seeking a Senior Incident Manager to direct the handling of critical service disruptions within its AI data center infrastructure. The company operates an AI…
Career Tools
About This Role
Lambda is seeking a Senior Incident Manager to direct the handling of critical service disruptions within its AI data center infrastructure. The company operates an AI cloud platform that provides computing power to enterprises, researchers, and hyperscalers, with a mission to make compute widely accessible. In this position, the professional will serve as the central command during major outages, coordinate cross-functional technical teams, manage communication with leadership, and drive post-incident…
Job Description
Lambda is seeking a Senior Incident Manager to direct the handling of critical service disruptions within its AI data center infrastructure. The company operates an AI cloud platform that provides computing power to enterprises, researchers, and hyperscalers, with a mission to make compute widely accessible. In this position, the professional will serve as the central command during major outages, coordinate cross-functional technical teams, manage communication with leadership, and drive post-incident reviews to enhance overall system reliability.
This role suits an experienced operations professional with a deep background in distributed infrastructure, large-scale GPU clusters, and high-availability environments. The person in this position will own the complete incident lifecycle, establish response playbooks, track key reliability metrics like MTTR, and participate in an on-call rotation. Candidates will collaborate closely with hardware vendors, networking groups, platform engineers, and data center operations to maintain stability across complex technical layers.
Lambda offers competitive compensation including cash and equity, health insurance for employees and dependents, 401(k) matching for US staff, wellness stipends, and flexible paid time off. This is a full-time, on-site role with a stated annual salary range between USD 125,000 and USD 166,000.
Responsibilities
- Manage high-severity incidents affecting AI infrastructure, storage, networking, and GPU clusters
- Act as Incident Commander during major outages to direct engineering, facilities, networking, and vendor personnel
- Oversee the entire incident response lifecycle from technical triage and escalation to resolution and post-mortem analysis
- Create and maintain operational playbooks, incident documentation, and internal response procedures
- Participate in an on-call rotation to coordinate and lead emergency responses
- Collaborate with security operations, platform reliability, data center operations, and infrastructure engineering teams
- Lead root cause analyses and post-incident reviews to implement corrective actions and track metrics such as MTTR and MTTD
- Deliver executive-level reports, active outage updates, and operational health summaries
Requirements
- Minimum of eight years of professional experience in infrastructure operations, site reliability engineering, or incident management
- Demonstrated background handling incidents within large-scale distributed systems
- Solid understanding of networking, storage, GPU compute clusters, and data center operations
- Familiarity with cloud or hybrid infrastructure platforms
- Experience utilizing incident management frameworks like SRE or ITIL
- Proficiency with monitoring and tracking software such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus, and Grafana
- Strong stakeholder management and communication abilities under high-pressure conditions
Qualifications
- Prior experience running high-performance computing or artificial intelligence infrastructure
- Familiarity with high-density GPU settings including NVIDIA clusters and InfiniBand networking
- Background working in colocation or hyperscale data center facilities
- Knowledge of automation tools and incident command systems
Core Skills
Benefits
- Cash and equity compensation
- Medical, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for eligible roles
- 401(k) retirement plan with a 2 percent company match for USA employees
- Flexible paid time off plan
Frequently Asked Questions
Answers are based only on the employer’s listing; where it doesn’t say, neither do we.
What is the salary range for the Senior Incident Manager position?
The annual salary range is set between USD 125,000 and USD 166,000, though adjustments may occur based on candidate qualifications.
Is this a remote position?
No, the posting indicates that the employment status is on-site.
What level of experience is required for this role?
Applicants must have eight or more years of experience in infrastructure operations, site reliability engineering, or incident management.
What employment type is this?
This is a full-time position.
Sample Interview Questions
AI-generated questions tailored to this specific role — a preview of the full practice set.