Site Reliability Engineer, Cloud Platform

QualysPublished 18 hours agoFirst seen 8 hours ago

Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!

Site Reliability Engineer (SRE)

The Opportunity

The Site Reliability Engineer (SRE) plays a critical role in ensuring the reliability, scalability, performance, and operational excellence of Qualys platforms and services. This role operates at the intersection of software engineering and operations, applying automation, observability, troubleshooting, and reliability engineering practices to maintain highly available production systems and improve customer experience.

The SRE will work closely with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to monitor production systems, troubleshoot issues, automate operational processes, and continuously improve system reliability.

Where This Role Sits

The SRE role partners closely with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to ensure production systems remain resilient, scalable, observable, and operationally efficient.

The role provides hands-on support for cloud-native and distributed systems, including Kubernetes, Apache Spark, Kafka, and cloud-based data platforms, while contributing to incident response, automation, monitoring, and reliability improvements.

Key Responsibilities

Reliability and Operations

  • Monitor production applications, infrastructure, and distributed systems to identify issues and maintain service health.
  • Participate in incident response, troubleshooting, escalation, and root-cause analysis.
  • Support highly available and scalable production services and workloads.
  • Monitor system performance, resource utilization, availability, latency, and overall service health.
  • Participate in on-call rotations and follow established incident management processes.
  • Support capacity planning and identify potential reliability and performance risks.
  • Assist engineering teams in resolving production issues and implementing corrective actions.
  • Contribute to post-incident reviews and continuous reliability improvements.

Data Platform and Distributed Systems

  • Support the operation and monitoring of Apache Spark workloads and Kafka pipelines.
  • Monitor distributed data-processing workloads and investigate performance or availability issues.
  • Assist with troubleshooting Spark jobs, Kafka pipelines, application failures, resource constraints, and related production issues.
  • Develop an understanding of distributed systems concepts and their operational requirements.
  • Support data platforms and technologies such as Delta Lake and S3 where applicable.
  • Apply basic SQL skills for troubleshooting, validation, and operational analysis.

Kubernetes and Cloud Operations

  • Assist with Kubernetes deployments, resource monitoring, application troubleshooting, and operational support.
  • Monitor containerized applications and investigate resource, availability, and performance issues.
  • Support cloud environments such as AWS, OCI, or Google Cloud Platform.
  • Work with infrastructure and engineering teams on configuration, deployments, and operational workflows.
  • Develop familiarity with Infrastructure as Code (IaC) and cloud-native operational practices.

Monitoring and Observability

  • Monitor applications, infrastructure, and distributed workloads using tools such as Prometheus and Grafana.
  • Create and maintain dashboards, alerts, and monitoring configurations.
  • Analyze logs, metrics, and system behavior to identify and troubleshoot production issues.
  • Track reliability indicators such as availability, latency, performance, resource utilization, and service health.
  • Contribute to improvements in observability and proactive issue detection.

Automation and Platform Engineering

  • Develop scripts and automation to reduce repetitive manual operational tasks.
  • Create tooling for monitoring, log analysis, troubleshooting, deployments, and operational workflows.
  • Use Python, Bash, Go, or similar programming/scripting languages to automate operational processes.
  • Support CI/CD and automated deployment workflows.
  • Contribute to Infrastructure as Code and self-service capabilities where applicable.
  • Continuously improve operational processes through automation and standardization.

Collaboration and Continuous Improvement

  • Collaborate with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to improve system reliability.
  • Assist development teams in designing and operating production-ready services.
  • Document system architectures, operational procedures, troubleshooting steps, and best practices.
  • Create and maintain runbooks, SOPs, and troubleshooting guides.
  • Proactively identify reliability risks and opportunities for operational improvement.
  • Contribute to SRE best practices, knowledge sharing, and continuous improvement initiatives.

What Good Looks Like

  • Proactively identifies reliability and performance risks before they significantly impact customers.
  • Responds effectively to production incidents and contributes to meaningful root-cause analysis.
  • Uses automation to reduce repetitive manual operational work.
  • Demonstrates strong troubleshooting skills across Linux, Kubernetes, cloud, and distributed systems.
  • Effectively monitors and analyzes system metrics, logs, and application behavior.
  • Collaborates effectively with engineering and infrastructure teams to resolve production issues.
  • Continuously improves system reliability, observability, scalability, and operational efficiency.
  • Builds a strong understanding of production systems and applies SRE principles to day-to-day operations.

Distinguishing Expectations

The SRE is measured by improvements in reliability, automation, operational efficiency, observability, and service availability.

For an early-career SRE, success is demonstrated through:

  • Effective monitoring and troubleshooting of production systems.
  • Rapid learning of complex production environments.
  • Reduction of manual operational effort through automation.
  • Consistent participation in incident response and root-cause analysis.
  • Improvements to monitoring, documentation, and operational processes.
  • Growing ownership of production services and reliability-focused initiatives.

Required Qualifications

  • 0–3 years of experience in SRE, DevOps, Cloud Engineering, Systems Engineering, or a related technical role.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Basic understanding of Linux/Unix systems, processes, memory, networking, and system troubleshooting.
  • Fundamental understanding of distributed systems and cloud-native technologies.
  • Experience or familiarity with Kubernetes and containerized applications.
  • Familiarity with Apache Spark, Kafka, or similar distributed data-processing platforms.
  • Strong programming and scripting fundamentals in Python, Bash, Go, Java, Scala, or similar languages.
  • Ability to write scripts for automation, monitoring, log analysis, and operational tasks.
  • Basic SQL knowledge and strong problem-solving skills.
  • Familiarity with monitoring and observability tools such as Prometheus and Grafana.
  • Basic understanding of cloud platforms such as AWS, OCI, or Google Cloud Platform.
  • Understanding of basic networking, security, and distributed systems concepts.
  • Willingness to learn production systems, debugging techniques, and SRE practices.

Preferred Qualifications

  • Familiarity with Delta Lake, S3, or distributed data platforms.
  • Understanding of CI/CD and automated deployment processes.
  • Familiarity with Terraform, CloudFormation, or other Infrastructure as Code tools.
  • Experience with scripting-based automation or internal tooling projects.
  • Familiarity with tools such as Jenkins, GitHub, Bitbucket, ELK, Splunk, or AppDynamics.
  • Experience with large-scale or distributed systems.
  • Understanding of reliability engineering principles and automation best practices.
  • Exposure to incident management, on-call processes, and postmortem practices.
  • Cloud certifications are a plus.
  • Strong analytical, communication, collaboration, and troubleshooting skills.

Core Competencies

SRE | Linux | Python / Bash / Go | Coding & Scripting | Spark | Kafka | Kubernetes | Cloud Platforms | SQL | Monitoring & Observability | Automation | Distributed Systems | Troubleshooting | Incident Response | CI/CD | Reliability Engineering

Work Environment

  • Full-time role.
  • On-call participation may be required.
  • Hybrid or remote flexibility depending on company policy.
  • Cross-functional collaboration with Engineering, DevOps, Infrastructure, Security, and Data Platform teams.
  • Opportunity to work with cloud-native, distributed, and large-scale production systems.

ABOUT QUALYS

Qualys, Inc. is a pioneer and leading provider of cloud-based IT, security, and compliance solutions, serving more than 10,000 customers across over 130 countries. Qualys delivers innovative security and compliance solutions that help organizations simplify security operations and reduce risk.