Open Role
Reddit

Remote opportunity at

Reddit

Staff Site Reliability Engineer, Ads

Reddit operates a massive online platform connecting hundreds of millions of daily active visitors through over one hundred thousand active communities. As the network continues its…

View Company

Role Snapshot

Hiring Now

Remote from

Remote

Salary

Undisclosed

Department

General

Employment

Full-time

Experience

Not specified

Published16d ago
Listing Views17
Applications0
Apply BeforeNo deadline

Career Tools

About This Role

Reddit operates a massive online platform connecting hundreds of millions of daily active visitors through over one hundred thousand active communities. As the network continues its global expansion, the Site Experience Site Reliability Engineering team focuses on maintaining speed, resilience, and performance across web portals, mobile applications, APIs, media distribution, and real-time features. The organization seeks a Staff Site Reliability Engineer to head up reliability initiatives for high-volume, user-facing systems.…

Job Description

Reddit operates a massive online platform connecting hundreds of millions of daily active visitors through over one hundred thousand active communities. As the network continues its global expansion, the Site Experience Site Reliability Engineering team focuses on maintaining speed, resilience, and performance across web portals, mobile applications, APIs, media distribution, and real-time features.

The organization seeks a Staff Site Reliability Engineer to head up reliability initiatives for high-volume, user-facing systems. This technical leadership position is well-suited for engineers experienced in managing large-scale distributed architectures who want to tackle complex availability hurdles and shape engineering practices company-wide.

In this capacity, you will collaborate closely with product and infrastructure engineering groups to boost availability, throughput, and operational standards. Responsibilities involve steering architectural planning for high global traffic, mitigating operational hazards, building automation solutions, guiding incident responses, and mentoring team members.

Responsibilities

  • Lead reliability engineering initiatives for critical user-facing systems and services
  • Enhance performance and resilience across APIs, content delivery, feeds, search, messaging, and real-time features
  • Partner with product and infrastructure teams on architectural decisions covering failover, redundancy, traffic management, and capacity planning
  • Identify systemic risks and operational bottlenecks to build proactive mitigation strategies
  • Develop automation and tooling to eliminate manual tasks and improve deployment safety
  • Lead complex incident response efforts, drive blameless postmortems, and ensure long-term fixes
  • Establish standards for reliability, SLIs and SLOs, release engineering, and operational maturity
  • Provide technical leadership and mentorship to engineers across SRE and software engineering groups

Requirements

  • At least 8 years of professional background in Site Reliability Engineering, Infrastructure Engineering, or related fields
  • Experience operating large-scale distributed systems and high-traffic user-facing production environments
  • Deep familiarity with distributed systems, networking, Linux systems, or cloud-native architectures
  • Demonstrated skill in designing highly available systems with robust reliability practices
  • Programming proficiency in languages such as Python or Go
  • Understanding of observability tooling including metrics, logs, traces, and alerts
  • Track record of improving reliability through SLOs, automation, and performance tuning
  • Ability to troubleshoot complex problems spanning applications, infrastructure, and networking

Qualifications

  • Experience operating systems handling internet-scale traffic volumes
  • Background using Kubernetes, containers, cloud infrastructure, and modern deployment platforms
  • Familiarity with technologies including Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, or Redis
  • Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure
  • Contributions to open-source software or participation in technical communities
  • Experience leading large-scale incident response and operational transformation initiatives

Core Skills

Benefits

  • Global benefit programs tailored to lifestyle, workspace, professional development, and caregiving support
  • Family planning support and gender-affirming care
  • Mental health and coaching benefits
  • Private medical, dental, and vision coverage
  • Personal retirement savings account with employer matching contributions
  • Cycle to work and tax saver schemes
  • Flexible vacation and paid volunteer time off
  • Generous paid parental leave
  • Eligibility to receive equity in the form of restricted stock units

Frequently Asked Questions

Is this position remote?

Yes, the job is listed with a remote location status.

What is the employment type?

The posting indicates that this is a full-time position.

What is the base salary range for this role?

The base salary range for U.S.-based employees is $217,000 to $303,900 USD, though final offers depend on skills, experience, and location.

What level of experience is required?

Candidates must have eight or more years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.

Sample Interview Questions

AI-generated questions tailored to this specific role — a preview of the full practice set.

Search similar jobs

Related Jobs

Reddit
RedditPosted 1d ago
Full-timeUSD 217,000 - 303,900/yr
Advertisement
320 × 50

Posted by Reddit

Source: Reddit

Reddit

Reddit

136Open Jobs
—No reviews yet
View Company Profile