Core Distributed Systems Database Engineer
Summary
Apple's Services Engineering organization (ASE) is seeking a deeply technical Distributed Systems Database Engineer to help architect the next generation of our massive-scale data platform. ASE Data Services operates Apache Cassandra and related coordination systems across Apple's datacenters worldwide, forming a foundational storage layer for iCloud and hundreds of millions of global users.
As our fleet grows, we are evolving our platform to support true multi-tenancy, dense resource utilization, and highly predictable performance. We are looking for an engineer to design and build an intelligent, auto-balancing cluster control plane that dynamically optimizes our hardware footprint. You will move comfortably between the abstract math of capacity modelling and the concrete realities of running distributed database at scale.
Description
The ASE Data Services team is building an automated, multi-tenant platform layer designed to extract maximum efficiency out of our global footprint without compromising on reliability. This requires an extraordinary degree of engineering rigor. You will own the architecture for resource quota management, node-level capacity estimation, driving performance efficiencies and dynamic cluster rebalancing algorithms across our fleet.
You will figure out how to model CPU, memory, IOPS, and data-streaming requirements per node, and build the software that automatically redistributes data and workloads to optimize the total node count. Over time, the platform will lean on modern workload-scheduling and orchestration primitives to enforce these boundaries, turning raw hardware into a highly efficient, self-healing, self-balancing multi-tenant utility. This work demands an innovative spirit, deep systems intuition, and the communication skills to partner with platform, SRE, and product teams across Apple.
Responsibilities
- Auto-Balancing Control Plane: Design and implement automated cluster balancing software that dynamically redistributes data and workloads across the fleet, optimizing for criteria like node count while ensuring safe thresholds for CPU, memory, and network utilization.
- Streaming Load Optimization: Architect rebalancing mechanisms that safely throttle and schedule heavy data streaming operations (cluster expansions, node migrations, ring balancing) to prevent live production traffic degradation.
- Multi-Tenant Architecture: Design and implement robust multi-tenancy frameworks within our distributed storage systems, ensuring hard isolation, tenant safety, and zero noisy-neighbor impact.
- Resource Modeling & Capacity Engineering: Develop algorithmic models to accurately estimate and forecast CPU, memory, IOPS, and network bandwidth requirements per node that tracks evolving data workloads.
- Quota Management & Rate Limiting: Build and maintain global and local quota management systems, dynamic rate-limiting, and admission control systems to protect cluster health.
- Failure Domain Isolation: Reason about and implement strict boundary controls across diverse infrastructure, factoring in cross-datacenter connectivity, host placement strategies, and cloud/bare-metal failure domains.
- Cross-Team Delivery: Partner closely with SRE, platform, and integration teams to ship, operate, and harden this system in production from design through on-call-grade reliability.
Minimum Qualifications
- BS or MS in Computer Science, Computer Engineering, or equivalent work experience.
- 7+ years of experience in infrastructure engineering, systems programming, or distributed systems development.
- Professional experience developing distributed systems (consensus protocols, replication strategies or consistency models.
- Experience profiling and modeling node-level resource usage (e.g., CPU, memory, I/O)
- A bias toward experimentation, you form hypotheses, prototype quickly, measure honestly, and let data settle the hard architectural questions.
Preferred Qualifications
- Cluster Orchestration & Balancing: Hands-on experience building auto-balancing systems, scheduling frameworks, or multi-dimensional resource allocation algorithms for large-scale distributed systems.
- Database Internals & Streaming: Experience with distributed database internals LSM-tree architectures. Deep understanding of streaming mechanics during node join/leave/move operations is a major plus.
- Workload Scheduling at Scale: Experience running large-scale stateful workloads on modern orchestration platforms (e.g., Kubernetes), with an understanding of container runtime boundaries, resource limits, and scheduling topologies.
- Quota & Traffic Control: Experience implementing admission control, token/leaky-bucket algorithms, or distributed rate limiting in high-throughput environments while considering fairness.
- Kernel & OS-Level Intuition: Familiarity with Linux cgroups (v1/v2), eBPF, I/O schedulers, and network topologies that govern resource isolation or the appetite to dig into them on demand.
- Willingness to ramp on the JVM, since much of our current stack is Java.
- Experience transitioning large single-tenant legacy systems into modern shared-resource, auto-scaling architectures.
At Apple, we're not all the same. And that's our greatest strength. We draw on the differences in who we are, what we've experienced and how we think. Because to create products that serve everyone, we believe in including everyone. Therefore, we are committed to treating all applicants fairly and equally. As a registered Disability Confident employer, we will work with applicants to make any reasonable accommodations. Apple will consider for employment all qualified applicants with criminal backgrounds in a manner consistent with applicable law. Learn more