Sr. Site Reliability Engineer, Data Center & Network Infrastructure
Description
The Enterprise Infrastructure SRE team is seeking a seasoned Sr. Site Reliability Engineer to own the reliability, automation, and observability of our data center and network infrastructure. This role bridges network operations and data center facility infrastructure—including bare-metal server reliability, power/cooling systems, and network fabric—with a relentless focus on uptime, scalability, and user experience.
You will minimize manual toil through advanced automation, drive operational excellence through blameless postmortems, and proactively mitigate infrastructure risks. As a senior engineer, you will lead high-impact projects, mentor peers by example, and foster continuous improvement across the global SRE organization.
Responsibilities
- Enhance Telemetry Platforms: Scale observability, logging, and alerting systems using Grafana, Splunk, and Prometheus
- Build Correlated Dashboards: Design visualization tools that connect server health, network telemetry, and facility power/cooling performance across fragmented data sources
- Drive Data Insights: Write complex SQL and SPL queries to analyze infrastructure trends, isolate production bottlenecks, and surface environmental health insights via IPMI interfaces
- Eliminate Manual Toil: Develop robust automation scripts and tooling to handle hardware incident triage, alert noise reduction, and log correlation
- Manage Source of Truth: Maintain and scale Netbox inventory systems, building automated API pipelines to track physical layout, device lifecycles, and cable topologies
- Standardize Operational Playbooks: Create and maintain high-quality runbooks, KB articles, and SOPs to enable bot-assisted incident resolution
- Lead Incident Response: Participate in high-severity infrastructure on-call rotations, directing rapid triage, mitigation, and root cause analysis (RCA) for server, network, and facility anomalies
- Cross-Functional Partnership: Collaborate with architecture, deployment, hardware engineering, and facility operations teams to ensure new implementations are supportable and monitored from day one
- Optimize Performance: Continuous monitoring of network and infrastructure performance to execute sustainability and optimization changes
Requirements
- Bachelor's Degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
- 8+ years of experience in Site Reliability Engineering (SRE), network operations, or data center infrastructure management across highly distributed, large-scale environments
- Strong proficiency in Python, Go, or Shell scripting, combined with configuration management frameworks like Salt or Ansible
- Observability Expertise: Deep production experience with Prometheus, Grafana, Alert Manager, and Splunk for enterprise metrics and logging
- Strong SQL skills (PostgreSQL, MySQL) and a proven track record of consuming and building RESTful APIs to integrate infrastructure tooling
- Advanced Linux system fundamentals paired with hands-on experience provisioning, troubleshooting, and managing bare-metal enterprise server architectures
- Deep understanding of TCP/UDP, IPv4/IPv6, BGP, EVPN, VxLAN, Segment Routing, and load balancing
- Experience managing enterprise hardware vendors like Arista Networks, Juniper Networks, Cisco, Palo Alto Networks firewalls, and F5 load balancers
- Practical knowledge of out-of-band management (IPMI), PDU architecture, and data center physical cooling systems (liquid cooling, HVAC, hot/cold aisle containment)
Compensation and Benefits
Benefits
Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:
- Medical plans > plan options with $0 payroll deduction
- Family-building, fertility, adoption and surrogacy benefits
- Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
- Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
- Healthcare and Dependent Care Flexible Spending Accounts (FSA)
- 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
- Company paid Basic Life, AD&D
- Short-term and long-term disability insurance (90 day waiting period)
- Employee Assistance Program
- Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
- Back-up childcare and parenting support resources
- Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
- Weight Loss and Tobacco Cessation Programs
- Tesla Babies program
- Commuter benefits
- Employee discounts and perks program

