Open Role
Lambda

Opportunity at

Lambda

Senior Incident Manager

Lambda is seeking a Senior Incident Manager to direct the handling of critical service disruptions within its AI data center infrastructure. The company operates an AI…

View Company

Role Snapshot

Hiring Now
Work arrangement
Not specified
Location
United States
Salary
USD 125,000 - 166,000/yr
Apply by
Not specified
Employment type
Full-time
Experience
Not specified
Published2h ago
Listing Views0
Applications0
Apply BeforeNo deadline

Career Tools

About This Role

Lambda is seeking a Senior Incident Manager to direct the handling of critical service disruptions within its AI data center infrastructure. The company operates an AI cloud platform that provides computing power to enterprises, researchers, and hyperscalers, with a mission to make compute widely accessible. In this position, the professional will serve as the central command during major outages, coordinate cross-functional technical teams, manage communication with leadership, and drive post-incident…

Job Description

Lambda is seeking a Senior Incident Manager to direct the handling of critical service disruptions within its AI data center infrastructure. The company operates an AI cloud platform that provides computing power to enterprises, researchers, and hyperscalers, with a mission to make compute widely accessible. In this position, the professional will serve as the central command during major outages, coordinate cross-functional technical teams, manage communication with leadership, and drive post-incident reviews to enhance overall system reliability.

This role suits an experienced operations professional with a deep background in distributed infrastructure, large-scale GPU clusters, and high-availability environments. The person in this position will own the complete incident lifecycle, establish response playbooks, track key reliability metrics like MTTR, and participate in an on-call rotation. Candidates will collaborate closely with hardware vendors, networking groups, platform engineers, and data center operations to maintain stability across complex technical layers.

Lambda offers competitive compensation including cash and equity, health insurance for employees and dependents, 401(k) matching for US staff, wellness stipends, and flexible paid time off. This is a full-time, on-site role with a stated annual salary range between USD 125,000 and USD 166,000.

Responsibilities

  • Manage high-severity incidents affecting AI infrastructure, storage, networking, and GPU clusters
  • Act as Incident Commander during major outages to direct engineering, facilities, networking, and vendor personnel
  • Oversee the entire incident response lifecycle from technical triage and escalation to resolution and post-mortem analysis
  • Create and maintain operational playbooks, incident documentation, and internal response procedures
  • Participate in an on-call rotation to coordinate and lead emergency responses
  • Collaborate with security operations, platform reliability, data center operations, and infrastructure engineering teams
  • Lead root cause analyses and post-incident reviews to implement corrective actions and track metrics such as MTTR and MTTD
  • Deliver executive-level reports, active outage updates, and operational health summaries

Requirements

  • Minimum of eight years of professional experience in infrastructure operations, site reliability engineering, or incident management
  • Demonstrated background handling incidents within large-scale distributed systems
  • Solid understanding of networking, storage, GPU compute clusters, and data center operations
  • Familiarity with cloud or hybrid infrastructure platforms
  • Experience utilizing incident management frameworks like SRE or ITIL
  • Proficiency with monitoring and tracking software such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus, and Grafana
  • Strong stakeholder management and communication abilities under high-pressure conditions

Qualifications

  • Prior experience running high-performance computing or artificial intelligence infrastructure
  • Familiarity with high-density GPU settings including NVIDIA clusters and InfiniBand networking
  • Background working in colocation or hyperscale data center facilities
  • Knowledge of automation tools and incident command systems

Core Skills

Benefits

  • Cash and equity compensation
  • Medical, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for eligible roles
  • 401(k) retirement plan with a 2 percent company match for USA employees
  • Flexible paid time off plan

Frequently Asked Questions

Answers are based only on the employer’s listing; where it doesn’t say, neither do we.

What is the salary range for the Senior Incident Manager position?

The annual salary range is set between USD 125,000 and USD 166,000, though adjustments may occur based on candidate qualifications.

Is this a remote position?

No, the posting indicates that the employment status is on-site.

What level of experience is required for this role?

Applicants must have eight or more years of experience in infrastructure operations, site reliability engineering, or incident management.

What employment type is this?

This is a full-time position.

Sample Interview Questions

AI-generated questions tailored to this specific role — a preview of the full practice set.

Related Jobs

Lambda
LambdaPosted 2h ago
Full-timeUSD 271,000 - 361,000/yr

Posted by Lambda

Source: Jobicy — USA

View the original listing ↗

Posted by employer: Not specified

Added to Jobsiz: Oct 10, 2026

2 sources detected for this job

The description was written by Jobsiz from the employer’s listing. Always check the original listing before applying.

Lambda

Lambda

2Open Jobs
—No reviews yet
View Company Profile