Application Reliability Engineer

ApplePublished 1 days agoFirst seen 1 days ago

Summary

At Apple, great ideas have a way of becoming great products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish.

Payments and Financial Engineering builds the systems that power Apple's Finance and Accounting capabilities, processing very high-volume micro-transactions across Apple Pay, App Store, iTunes, Retail, Online, and Reseller channels to ensure accurate invoicing, disbursements, receipts, and payments. We are looking for an Application Reliability Engineer (ARE) to own the operational health of this platform - driving problem management, root cause resolution, and automation across a distributed, high-scale system that underpins a majority of Apple's revenue.

Description

In this role, you will be a strong software engineer with production-debugging and distributed-systems expertise, responsible for the reliability of a mission-critical payments platform - not simply triaging and closing incidents, but diagnosing, automating, and permanently eliminating the underlying application issues. You will reproduce failures, read and modify production code to implement the fix, and add detection or prevention mechanisms so the same class of issue doesn't recur.

Responsibilities

  • You will dig into problems such as consumer/thread starvation, blocked threads, connection pool exhaustion, file watcher behavior, queue offset and retry/dead-letter behavior, stale connections, resource leaks, scheduler behavior, application state issues, dependency timeouts, and race conditions - using logs, metrics, and traces to isolate root cause across a distributed system.
  • You will work closely with Application Engineering, Production Support, Infrastructure, and DBA teams, and communicate production trends, RCAs, and remediation plans to stakeholders across Finance, Business Operations, and Engineering.

Minimum Qualifications

  • 5-7 years of experience in software engineering, application support, or SRE roles for critical, high-scale production systems, ideally in a payments, financial, or transaction-processing domain.
  • Strong programming skills in Java or another JVM language, with working proficiency in at least one of Python, Go, or C++, or Shell Scriptng.
  • Ability to read, debug, and modify production code to implement permanent fixes, not just workarounds.
  • SQL and database troubleshooting experience (e.g., Oracle, MongoDB, PostgreSQL), including diagnosing connection pool exhaustion, stale connections, and slow queries.
  • Practical understanding of concurrency, threading, and asynchronous processing, and the failure modes they introduce (thread/consumer starvation, blocked threads, race conditions)
  • Experience with observability stacks - logs, metrics, and traces - to diagnose application-level observability gaps and reproduce failures
  • Strong Linux/Unix fundamentals, including process management, file permissions, and cron
  • Performance profiling and troubleshooting experience, including identifying resource leaks and dependency timeouts
  • Experienced in troubleshooting and driving root cause resolution of production incidents, including identifying performance bottlenecks and proposing fixes
  • Self-starter with strong customer and product focus, and a desire to gain awareness of an ecosystem and how frontend and backend systems collaborate.
  • Ability to communicate thoughtfully, write engineering proposals and support playbooks, leverage problem-solving skills, build a learning mindset and establish long-term relationships.

Preferred Qualifications

  • Experience with containerization and orchestration (Docker, Kubernetes)
  • Familiarity with CI/CD systems and infrastructure fundamentals across prod/non-prod environments
  • Experience with cloud platforms (AWS, Azure, GCP), including managed services and cloud migration
  • Experience with caching systems (Redis, Memcached) and messaging/queueing systems (Kafka, SQS, RabbitMQ), including queue offset management and retry/dead-letter handling.
  • Working knowledge of system design principles applicable to high-throughput, high-availability platforms.
  • Experience with distributed storage systems (S3, GridFS).
  • Track record of building tools or frameworks that detect and prevent recurring classes of production issues.
  • Experience setting up trend analysis or observability processes to proactively identify chronic issues.
  • Solid grounding in REST/API and distributed-systems fundamentals
  • Contributions to system design reviews or architecture discussions focused on reliability and scalability.