Operations Engineer

AppleApplyPublished 1 hours agoFirst seen 1 hours ago
Apply

Summary

The Operations Engineer in Crypto Services team manages key technical infrastructure. An ideal candidate will have experience in Systems Administration. The Operations Engineer will monitor infrastructure and application services and drive incident management. The Ops engineer will work closely with SRE's, PKI Engineers, systems engineers, network engineers, database administrators and information security teams to effectively ensure availability and reliability. For this position, strict application security and high availability requirements must be balanced to achieve optimal solutions.

Description

The successful candidate will participate in troubleshooting issues following established procedures, documenting problems, managing incidents and owning the issue from the initial contact to resolution. Additionally the engineer will be responsible for critical compliance tasks, system patching and upgrades, and building out observability tooling to support the team's operations.

Responsibilities

  • Responsibilities of the Operations Engineer include the following:
  • • Serve as a full time, primary on-call, responding and mobilizing efforts to address outages
  • • Follow change management procedures and deploy code using configuration management (e.g. Puppet, Chef, Ansible, etc)
  • • Perform routine and emergency patching and RHEL version upgrades across the fleet
  • • Set priorities and work efficiently in a fast-paced environment
  • • Measure and optimize system performance
  • • Monitor telemetry and address alerts to ensure smooth operations
  • • Design, build, and maintain dashboards and create alerting rules to improve visibility into system health
  • • Process and manage log pipelines to support troubleshooting, auditing, and compliance needs
  • • Strong communication skills and ability to work effectively across multiple business and technical teams
  • • Demonstrate ability to deliver results on time with high quality

Minimum Qualifications

  • 5+ years of expertise with Linux (any distro, but especially RHEL), including experience with OS patching and upgrade cycles. Standard UNIX utilities and programs
  • Strong understanding of SRE principles and goals, coupled with prior on-call experience, including leading incident command and management efforts.
  • Proficiency with configuration management tools (e.g., Puppet, Chef, Ansible) and scripting languages (e.g., Bash, Python).
  • Experience with monitoring tools (e.g., Icinga/Nagios) and log aggregation platforms (e.g., Splunk, OpenSearch) , and building dashboards/alerts
  • A proven track record of practical problem-solving, coupled with excellent communication and documentation skills, particularly in conveying complex technical information to diverse audiences.

Preferred Qualifications

  • Experience with distributed teams or managing operations across multiple time zones.
  • Working knowledge of cloud platforms (AWS/GCP/AliCloud) and container technologies (Podman, Docker, Kubernetes).
  • A deep understanding of security and compliance best practices, especially within a cryptographic or PKI environment.
  • Experience supporting Java applications and/or managing critical hardware like Hardware Security Modules (HSMs).
  • Flexibility to travel for work