Senior Site Reliability Engineer, Reddit

$190.8-267.1k

+ Equity in the form of restricted stock units

Kubernetes
Python
Linux
Go
Prometheus
Grafana
loki
Senior and Expert level
New York
Reddit

Online platform for thoughts, experiences and discussions

Job no longer available

Reddit

Online platform for thoughts, experiences and discussions

1001+ employees

B2CB2BPublishingContentSocial MediaCommunity

Job no longer available

$190.8-267.1k

+ Equity in the form of restricted stock units

Kubernetes
Python
Linux
Go
Prometheus
Grafana
loki
Senior and Expert level
New York

1001+ employees

B2CB2BPublishingContentSocial MediaCommunity

Company mission

Reddit's mission is to bring community and belonging to everyone.

Role

Who you are

  • We are looking for someone who thrives at the intersection of infrastructure and software development
  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python
  • Experience with Kubernetes and Cloud systems
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems
  • A demonstrated ability to debug, fix, and optimize code
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems
  • Strong working knowledge of Linux and containers
  • Excellent communication and collaborative skills

What the job involves

  • Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before
  • This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet
  • As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services
  • This team will work very closely with the Compute, Traffic, and Observability infrastructure teams
  • They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale
  • We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more
  • In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems
  • You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience
  • Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence
  • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure
  • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components
  • Identify and engineer away risk across Reddit’s systems
  • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution
  • Automate critical aspects of the event driven development process
  • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem
  • Share on-call responsibilities
  • Observe and improve performance, reduce cost, and improve the experience for millions of users
  • Contribute upstream changes to the open source projects we use

Share this job

View 99 more jobs at Reddit

Insights

Top investors

57% employee growth in 12 months

Company

Company benefits

  • Paid volunteer time off
  • 4+ months paid parental leave
  • Personal and professional development stipend
  • Work from home opportunities
  • Health insurance

Funding (last 2 of 6 rounds)

Aug 2021

$410m

SERIES F

Feb 2021

$250m

SERIES E

Total funding: $1.2bn

Our take

Reddit is a website that facilitates thousands of message board communities, known as subreddits, with an aim to promote authentic human connection. There are more than 100,000 communities on Reddit, covering everything, from food, entertainment, sports, and books, to more niche topics that cater for very specific audiences.

The simple platform is used by more than 52 million people a day and attracts over 50 billion monthly views. While these are impressive numbers, they do pale in comparison with social media giants Facebook and Twitter, however, Reddit distinguishes itself by providing easy-to-find communities for a truly endless range of topics.

Reddit makes money through advertising as well as offering a premium ad-free membership plan. The company has enjoyed continuous user & revenue growth, acquisitions by Conde Nast in 2006 and Advance in 2011, and plans to launch an IPO bid in 2024.

Kirsty headshot

Kirsty

Company Specialist at Welcome to the Jungle