Senior Data Engineer

AppleApplyPublished 1 days agoFirst seen 1 days ago
Apply

Summary

Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, or Apple Store experience we deliver is the result of us making each other's ideas stronger.

That happens because every one of us shares a belief that we can make something wonderful and share it with the world, changing lives for the better. It's the diversity of our people and their thinking that inspires the innovation that runs through everything we do.
When we bring everybody in, we can do the best work of our lives. Here, you'll do more than join something — you'll add something.

Description

Within Worldwide Sales & Operations Support, the Pacific Operations team is responsible for flawlessly executing Apple's supply chain in the Pacific region to fulfill millions of customer orders from Greater China, North East Asia, South Asia and Australia-New Zealand while making sure that the right Apple products reach our customers at the right place and at the right time.

We are seeking an experienced and pragmatic Senior Data Engineer to join our Business Analytics team. In this role, you will design, build, and maintain the end-to-end data pipelines and operational software applications that our Operations team depends on. The ideal candidate excels at bridging complex data ecosystems and functional software, and is tool-agnostic: you choose the most resilient solution for the problem — whether that is a deterministic SQL model, an asynchronous worker queue, a graph traversal algorithm, or a modern machine learning/LLM component.

In this role, you will work closely with cross-functional teams to solve complex problems, operationalize data into decision-making tools, and contribute to the strategic direction of our business analytics initiatives.

Responsibilities

  • As a Senior Data Engineer in our organization,
  • You will design robust ingestion and transformation pipelines,
  • Build polyglot storage strategies (relational, vector, and graph), and build the downstream backend services, APIs, and user interfaces that turn data into operational tools.
  • You will build and maintain scalable batch and streaming data pipelines across diverse storage paradigms, including relational warehouses, vector stores for semantic search, and graph databases for multi-hop network and entity modeling.
  • Work in tandem with our data science and analytics teams, you will develop production-ready backend services and lightweight functional interfaces that operationalize clean data for decision-making and automation. -Part of this work is serving data to users through agentic solutions — building the retrieval, tools, and interfaces that let AI agents answer questions and take action on our data.
  • A crucial part of your role is to apply systems thinking: you will diagnose complex, multi-variable business problems, map feedback loops and bottlenecks, and design resilient architectures that account for latency, system degradation, and upstream/downstream dependencies.
  • You will evaluate technology trade-offs objectively, implementing deterministic rules engines, constraint optimization, or statistical models first, and integrating applied AI/LLMs strictly where dynamic reasoning provides clear leverage.
  • Additionally, you will ensure operational rigor and observability through data contracts, automated integration tests, CI/CD pipelines, and monitoring that detects silent pipeline failures, index drift, and model latency before they impact end users.

Minimum Qualifications

  • 5+ years of practical experience building production data pipelines, distributed systems, and backend services.
  • Deep proficiency in Python, modern SQL, and data modeling frameworks (e.g., dbt).
  • Hands-on experience with orchestration tooling (e.g., Airflow, Dagster, Prefect) and modern data warehouses (e.g., PostgreSQL, Snowflake, DuckDB).
  • Experience applying vector databases/extensions (e.g., Milvus, pgvector, Qdrant, Pinecone) to build features — embedding, semantic search, hybrid search, and RAG retrieval.
  • Experience applying graph databases (e.g., TigerGraph, Neo4j, Memgraph, AWS Neptune) to build features — data modeling, writing queries (Cypher/Gremlin), and multi-hop traversals.
  • Strong API development experience using modern frameworks (e.g., FastAPI, Go, Node.js) and async task queues/streams (e.g., Celery, Redis, Kafka).
  • Working knowledge of how AI agents operate — tool/function calling, retrieval, context management, and multi-step reasoning — and experience serving data to users through agentic solutions (e.g., MCP servers, agentic frameworks, LLM tool use).
  • Practical exposure to applied AI/ML workflows (embeddings, semantic caching, model serving endpoints, or RAG pipelines) with the discipline to use them judiciously.
  • Strong systems-thinking abilities, with the capacity to identify second-order consequences and systemic constraints before writing code.
  • Ability to operate autonomously across the entire stack to diagnose ambiguous operational problems and deliver working solutions.
  • Bachelor's degree in Computer Science, Data Engineering, or a related field.

Preferred Qualifications

  • Experience delivering end-to-end applications to user-facing interfaces (e.g., React, Streamlit).
  • Working knowledge of Docker for packaging and deploying their own applications (containerize a service, write a Dockerfile, run it).
  • Experience implementing data observability and tracing tools (OpenTelemetry, Great Expectations, Monte Carlo, Prometheus).
  • Background in operations research, graph algorithms, or mathematical optimization.
  • Strong knowledge of supply chain processes, including demand forecasting, inventory management, and replenishment.
  • Demonstrated skill in data visualization and communication tools (e.g., Tableau).
  • Master's degree in Computer Science, Data Engineering, or a related field.