Software Engineer - Observability

AppleApplyPublished 3 days agoFirst seen 3 days ago
Apply

Summary

The Apple Services Engineering (ASE) team is building the next generation of foundational tools that empower software developers at Apple to build products that our customers love. The Observability team within ASE is a fast moving, highly skilled team that is designing and building a suite of platforms and services that help Apple engineers observe and get insights into their systems. If the thought of working with petabytes of data interests you, this is the place to be. Our systems must scale globally, stay highly available, and "just work," while supporting some of the largest services in the world.

We are seeking a versatile engineer who thrives wearing multiple hats. In this role you will primarily design and build the control plane tooling and automation that operates our time series database (TSDB) in production at scale, spanning provisioning, orchestration, scaling, health, and self-healing. Just as importantly, you will not be confined to tooling: you will regularly pick up core developer tasks on the data platform itself, contributing to the distributed systems that ingest, store, and query time series data. If you love the intersection of software engineering and large-scale operations, and you want your work to make massive fleets effortless to run, this is the place to be. If you'd love to join this amazing team, we'd love to hear from you!

Description

We are looking for an engineer who blends strong software engineering fundamentals with a deep operational mindset. You will own the control plane that makes our TSDB reliable, scalable, and easy to operate, and you will fluidly move between building operational tooling and shipping product features. You should be comfortable owning problems end to end, from requirement gathering through design, implementation, rollout, and on-call operability.

Responsibilities

  • Designing and building control plane tooling to provision, orchestrate, scale, upgrade, and self-heal a time series database operating in production at scale.
  • Building automation for deployment, configuration management, capacity planning, and fleet lifecycle management across globally distributed clusters.
  • Wearing multiple hats: in addition to tooling, picking up developer tasks on the core TSDB and distributed data services, contributing to ingestion, storage, and query paths.
  • Requirement gathering across cross functional teams and leading technical design discussions.
  • Driving down operational toil through automation, and improving reliability, performance, and cost efficiency of the fleet.
  • Gaining in-depth understanding of the domain, leading independent research where needed, and mentoring other engineers on the team.
  • You will have the courage and experience to be frank and ambitious but humble enough to listen to others. We want your thoughts on how we can move faster, be more creative, and deliver tools and ideas to empower developers. We expect you to challenge the status quo, to care about the details, the end user, and how it all comes together. You should be someone with ideas and passion for software and systems delivered as a service to maximize reuse, efficiency, and simplicity. Your work will impact millions of Apple users and is necessary to the success of some of the most visible current and future features.

Minimum Qualifications

  • BS or MS in CS or equivalent
  • 5+ years of industry experience
  • Deep understanding of core CS concepts including data structures, algorithms, and concurrent programming
  • Proficiency in one or more programming languages such as Java, Scala, Go, or Python
  • Experience designing, implementing, and operating highly scalable infrastructure services in production
  • Deep understanding and work experience in distributed systems
  • Hands-on experience building automation and tooling for deployment, orchestration, and operations at scale
  • Strong attention to detail, excellent analytical capabilities, and a bias toward reducing operational toil
  • Passion for developing clear, robust, and maintainable software and systems

Preferred Qualifications

  • Experience building control plane or fleet management tooling for stateful/data systems in production
  • Strong DevOps skillset: CI/CD pipelines, infrastructure as code (e.g. Terraform), configuration management, and GitOps workflows
  • Hands-on experience with container orchestration and scheduling (Kubernetes, operators, custom controllers)
  • Familiarity with time series database internals and the operational challenges of running them at scale
  • Experience with Observability solutions using OpenTelemetry, Prometheus, Grafana
  • Experience with columnar storage systems and/or building query engines is a plus
  • Experience building high performant systems in Rust is a plus
  • Comfortable wearing multiple hats, moving between operational tooling and core product/feature development
  • Great communication skills