Staff Site Reliability Engineer, Infrastructure Engineering
Description
This team manages multiple functions across Tesla that includes Platform Engineering managing fleet of Kubernetes clusters, Devops, MLOps, Cloud Infrastructure, Sandbox platform and Factory SRE as well. Continued development and automation of deployment, monitoring, self-healing, and alerting processes is imperative to the success of our engineering groups.You will be responsible for the internal platform that lets every team securely spin up sandboxed AI agents, run ephemeral workloads, train models at scale, and serve them in production all as a reliable, self-service experience. This is a high-leverage, high-ownership role where you will design, build, and operate the systems that sit at the intersection of Kubernetes, high-performance networking, GPU infrastructure, and modern ML platforms.
Responsibilities
- Build and own the end-to-end AI/ML platform (training, inference, experimentation) as a self-service product for all internal users
- Write production Kubernetes operators and controllers in Go for GPU workloads, training jobs, model deployments, sandboxes, and cluster lifecycle
- Operate large-scale GPU fleets (A100/H100/B200) scheduling, MIG, topology-aware placement, health monitoring
- Own RoCE/RDMA networking for distributed training
- Build and operate inference infrastructure using KServe, Triton, vLLM, and Ray Serve autoscaling, model versioning, and request batching
- Build and operate training-as-a-service; distributed training (PyTorch , FSDP), MLflow, and checkpoint management
- Design secure, isolated sandbox environments for AI agents and ephemeral execution contexts for untrusted code (gVisor, Kata, Firecracker)
- Architect serverless/Lambda-style ephemeral workloads, scale-to-zero, event-driven compute (Knative, KEDA, Firecracker microVMs)
- Integrate GPU compute, training, inference, and sandboxes as first-class services in the internal cloud platform
- Manage Kubernetes cluster lifecycle at fleet scale using Cluster API (CAPI) provisioning, upgrades, scaling, and decommissioning
Requirements
- Built agent sandbox/ephemeral compute platforms (gVisor, Kata, Firecracker, or similar)
- Built GPU sandboxes ; GPU passthrough (VFIO/SR-IOV), fractional allocation, secure isolated environments
- Configured and troubleshot RoCE v2 / RDMA fabrics for distributed training
- Production experience with KServe, Triton, vLLM, or Ray Serve for model inference
- Built or operated training-as-a-service, distributed training, MLflow, job scheduling
- Integrated platform services into an internal cloud (full stack development) - shared IAM, networking, storage, billing, service catalog
- Root cause analysis and systems thinking - failure modes, back-pressure, resource contention, blast radius
- Deep Kubernetes internals expertise (scheduler, API server, etcd, admission controllers, CRDs, control plane at scale)
- Built production Kubernetes operators in Go (controller-runtime / Kubebuilder)
- Production Cluster API (CAPI) experience; management clusters, custom providers, ClusterClass, fleet-scale lifecycle
Compensation and Benefits
Benefits
Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:
- Medical plans > plan options with $0 payroll deduction
- Family-building, fertility, adoption and surrogacy benefits
- Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
- Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
- Healthcare and Dependent Care Flexible Spending Accounts (FSA)
- 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
- Company paid Basic Life, AD&D
- Short-term and long-term disability insurance (90 day waiting period)
- Employee Assistance Program
- Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
- Back-up childcare and parenting support resources
- Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
- Weight Loss and Tobacco Cessation Programs
- Tesla Babies program
- Commuter benefits
- Employee discounts and perks program

