Senior Site Reliability Engineer - FedRAMP
At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.
Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.
Role Summary
We are seeking an experienced, security-focused Senior Site Reliability Engineer (Senior SRE) to drive the reliability, architectural design, and continuous compliance of our FedRAMP-authorized cloud platform. In this senior role, you will combine hands-on operational leadership with infrastructure architecture, technical governance, and audit readiness across AWS (Commercial and GovCloud) and Kubernetes environments.
As a Senior SRE, you will design and operate highly resilient, multi-cluster Amazon EKS infrastructure, direct observability architectures using Prometheus and Grafana, maintain automated deployment pipelines using GitLab CI/CD, and lead technical response in a 24/7 rotational on-call schedule. You will also serve as a key technical liaison for Third-Party Assessment Organization (3PAO) FedRAMP audits, ensuring strict adherence to NIST SP 800-53 controls, architectural security patterns, and vulnerability management SLAs.
Role Snapshot
Primary Focus: Production reliability, hybrid cloud & edge infrastructure, FedRAMP Compliance, Secure Architectural Design & L3 Escalation Support.
Cloud & Edge Platforms: AWS (GovCloud / Commercial EKS), EKS, S3, KMS, RDS, IAM bare-metal on-premises edge servers.
Core Tooling: Kubernetes, ArgoCD, Terraform, GitLab CI/CD (FIPS-compliant runners).
Observability & Logging: Prometheus, Grafana, Alertmanager, Slack, Elasticsearch, Kibana, SIEM integration.
Programming: Python, Go (Golang), Bash.
On-Call Rotation: L3 escalation support
Key Responsibilities
1. L3 On-Call Escalation & Production Support
• Act as the final technical escalation tier (L3 support) for critical platform and production incidents, troubleshooting complex issues escalated by L1/L2 operations or customer support teams.
• Lead high-priority incident response bridges, coordinating across development, security, and networking teams to drive rapid issue containment and resolution under strict service-level agreements (SLAs).
• Troubleshoot deep, transient infrastructure errors (e.g., routing, network package drops, Kubernetes control-plane failures) that go beyond standard SOPs.
• Establish clear, documented escalation pathways, translating complex L3 resolutions into actionable runbooks and empower L1/L2 teams
2. FedRAMP Compliance, Audit Readiness & Log Auditing
• Implement and enforce security controls aligned with the NIST SP 800-53 framework to achieve and maintain our FedRAMP Authorization to Operate (ATO).
• Act as the technical lead for SRE during annual FedRAMP 3PAO audits, gathering evidence, demonstrating compliance, and proving operational control implementation.
• Build and maintain secure, tamper-proof audit logging pipelines—forwarding application logs, system logs, API call records, and Kubernetes audit trails securely into Elasticsearch (and central SIEMs) with strict, compliant retention and index-lifecycle management policies.
3. Deep Log Troubleshooting & Analysis (ELK Stack)
• Utilize Elasticsearch and Kibana as primary investigative tools to perform deep-dive troubleshooting of complex, distributed system anomalies and application errors across AWS and edge environments.
• Build, customize, and curate high-signal Kibana dashboards, search queries (KQL/Lucene), and visualizations to provide real-time operational visibility and dramatically reduce Mean Time to Resolution (MTTR).
• Troubleshoot log-ingestion pipelines (Vector, Fluentd) to resolve bottlenecks, parsing errors, or missing metadata in high-volume production environments.
4. Secure Architectural Design & Hybrid Infrastructure
• Design and document highly resilient, secure-by-default architectural topologies for hybrid networks spanning AWS GovCloud and on-premises edge servers.
• Architect and scale multi-tenant Kubernetes (EKS) clusters enforcing strict physical or logical boundaries, zero-trust network policies, and identity federation.
• Configure and optimize high-availability L4/L7 load balancer networking, ingress controllers, and FIPS 140-3 validated traffic routing for secure, low-latency edge-to-cloud communication.
• Manage physical bare-metal edge nodes, defining operating system hardening baselines (STIG compliance), secure boot processes, and automated physical host provisioning.
5. Infrastructure as Code & Deployments
• Author clean, modular, and secure Terraform code to provision cloud infrastructure, networking topologies, and security boundaries.
• Build declarative deployment pipelines using GitLab CI/CD and ArgoCD to achieve true GitOps-driven delivery, ensuring all code modifications are traceable, signed, and fully audited.
• Develop custom automation tools, controllers, and CLI utilities in Python or Go (Golang) to eliminate repetitive toil and automate continuous compliance reporting.
6. Observability & Operational Excellence
• Design end-to-end monitoring and dashboarding systems utilizing Prometheus, Grafana, and Alertmanager.
• Build actionable alerting pipelines integrated with Slack, ensuring notifications are high-signal and low-noise.
• Facilitate blameless post-mortem reviews following high-severity incidents, utilizing evidence harvested from Kibana logs to document timelines, determine root causes, and programmatically prevent recurrences.
Required Qualifications
This position will require you to be a US Citizen residing in the United States.
• 5+ years of production experience operating high-availability systems in a Site Reliability Engineering, DevOps, or Systems Architecture role.
• L3 Production Escalation: Proven track record of handling high-pressure L3 production escalations, driving live incident triage, and managing stakeholder communication.
• Logging & Deep Troubleshooting: Advanced hands-on experience using Elasticsearch and Kibana to troubleshoot critical system events, write complex queries (KQL), parse logs, and construct operational dashboards.
• FedRAMP & Audit Expertise: Direct experience implementing, maintaining, and defending systems during FedRAMP audits (Moderate or High), NIST SP 800-53, or SOC 2 Type II audits.
• Architectural Design Skills: Strong capability in designing hybrid cloud network architectures, secure multi-tenant clusters, and formulating system security plans (SSPs).
• Deep Linux Internals: Strong proficiency in Linux administration (RHEL, Rocky Linux, or hardened Ubuntu), networking (TCP/IP, iptables, DNS), and OS hardening.
• Kubernetes & AWS EKS: Expert-level mastery of Kubernetes orchestration, pod autoscaling, ingress management, and cloud infrastructure on AWS/GovCloud.
• Observability Expertise: Hands-on experience configuring and scaling Prometheus, writing PromQL queries, building Grafana dashboards, and managing Alertmanager rules connected to Slack.
• IaC & GitOps: Proven track record implementing Terraform at scale and continuous deployment with ArgoCD and GitLab CI.
• Software Development: Proficiency in writing production-grade automation scripts and tools in Go (Golang) and/or Python.
• Edge & Network Engineering: Practical experience working with on-premises edge servers and configuring hardware/software load balancers (e.g., F5 BIG-IP, HAProxy, NGINX, AWS ALB/NLB) in secure zones.
• Communication & Documentation: Clear, structured written communication skills with a proven habit of writing detailed runbooks and compliance-aligned documentation.
#LI-KA1
The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.
The annual base pay for this position is: $161,900.00 - $242,900.00F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5’s differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change.
You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5’s benefits can be found at the following link: https://www.f5.com/company/careers/benefits. F5 reserves the right to change or terminate any benefit plan without notice.
Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or @myworkday.com).
Equal Employment Opportunity
It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting accommodations@f5.com.

