Senior Distributed Systems Engineer - Services Special Projects

AppleApplyPublished 2 hours agoFirst seen 2 hours ago
Apply

Summary

At Apple, great ideas turn into phenomenal products, services, and customer experiences at a pace few companies can match. We're looking for an experienced backend engineer to design and build massively scalable, highly available services that power experiences for Apple customers.

Description

As a senior member of the Services Engineering Team, You'll build and operate high-throughput, low-latency backend services that ingest, process, and serve data at scale across a range of mission-critical workloads — from real-time transactions to analytics and content delivery. You'll also help evolve a multi-tenant platform, including AI/ML-powered services, by shipping new capabilities, scaling what exists, and applying distributed-systems best practices from design through production.

Responsibilities

  • Design and build new features and services with a focus on scalability, responsiveness, fault tolerance, and high availability.
  • Design and build large-scale backend services that are resilient to failures, without breaking downstream consumers.
  • Partnering with search/ranking teams, ML Teams and business teams to ship customer-facing features through performant, well-modeled service APIs.
  • Provide technical guidance and mentorship, fostering a culture of learning and growth.

Minimum Qualifications

  • Master's degree in Computer Science or a related field
  • 8+ years of professional software development experience building scalable, distributed systems in production
  • Experience building, authoring, and operating large-scale, multi-tiered distributed systems and customer-facing web services: including API design, authentication, authorization, scaling for high availability, concurrency, and reliability.
  • Strong understanding of concurrency and multi-threaded programming, fundamental data structures, and efficient algorithm design
  • Strong proficiency in Java; working knowledge of a second systems language (Go, C++) is a plus. Solid OO analysis and design skills.
  • Strong proficiency in application frameworks (Spring boot)
  • Hands on experience with JVM performance tuning and profiling - GC selection/tuning, JFR, async-profiler, heap/thread-dump analysis for low-latency services.
  • Hands-on with async and reactive JVM stacks: Netty, Project Reactor, RxJava, Vert.x, or Micronaut/Quarkus.
  • Rigorous testing discipline with JUnit 5, Mockito, AssertJ, Testcontainers, and contract testing.
  • Hands on experience with build and dependency management with Gradle
  • Deep understanding of transactional consistency models - ACID semantics, with deep knowledge of tradeoffs between relational and NoSQL database technologies
  • Expertise with synchronous and asynchronous network I/O and RPC frameworks (gRPC)
  • Experience building and maintaining CI/CD pipelines (e.g., Jenkins, GitHub Actions, GitLab CI, or similar) for automated testing, build, and deployment of production services.
  • Experience with AWS or GCP and cloud-native tooling (Docker, Kubernetes) in the context of deploying scalable production grade services.
  • Experience with event streaming and queueing systems (specifically Kafka) and stream processing frameworks and high-throughput, append-only write paths for durable, queryable historical records.
  • Hands on Experience of leveraging data storage (Iceberg, Cassandra) and caching technologies (Redis) in Production services
  • Hands on experience with Serialization/schema tooling: Jackson, Protobuf, Avro, and Schema Registry
  • Experience identifying, triaging, and remediating security vulnerabilities in production services (dependency management, secure code review, threat modeling).

Preferred Qualifications

  • Self-motivated, with strong collaboration and communication skills, and experience in a fast-paced, agile environment
  • Experience with machine learning systems, ML frameworks, libraries and algorithms
  • Familiarity with deployment and optimization of Large scale Production grade AI Services that require GPUs in the path of the transaction.
  • Hands-on experience deploying, serving, and optimizing LLMs or ML models directly in the transaction/request path
  • Experience with security and cryptography (e.g., TLS, X.509 certificates) identity and access management protocols (OAuth2/OIDC/SAML), and secure token/session lifecycle management.