Data Governance Engineer

ApplePublished 2 hours agoFirst seen 2 hours ago

Summary

At Apple, we focus deeply on the customer experience. Apple Ads brings this same approach to advertising, helping people find exactly what they're looking for and helping advertisers grow their businesses. Our technology powers ads and sponsorships across Apple Services, including the App Store, Apple News, Apple Maps, MLS, and F1. Everything we do is designed for trust, connection, and impact: we respect user privacy, integrate advertising thoughtfully into the experience, and deliver value for advertisers of all sizes — from small app developers to big global brands. Because when advertising is done right, it benefits everyone.

The Apple Ads Data Governance team develops privacy-centric advertising solutions that leverage advanced data engineering and machine learning technologies at massive scale. The Data Governance Engineer role involves collaborating with cross-functional teams to implement critical privacy safeguards and data management controls.

Description

At Apple Ads, we are building the next generation of privacy-focused advertising capabilities. In the Data Governance team, we work at the cutting edge of data engineering, machine learning, and privacy at Apple's scale. We are constantly developing data and privacy management products to provide amazing user experiences and to drive value for developers and partners.

The team is seeking a Data Governance Engineer to support our data governance objectives. You will play a crucial role in delivering on Apple's privacy commitments to our customers. You will partner across our engineering, product, privacy and reliability organizations to deliver on data access, use, protection, minimization, retention, and other data governance execution areas.

Responsibilities

  • Design autonomous governance agents for petabyte-scale data pipelines (e.g. Kafka, Spark, Hadoop) ensuring compliance and quality across all processing stages
  • Build MCP integrations enabling automated policy enforcement, PII detection, and data quality validation across diverse storage and compute environments
  • Partner with engineers and PMs to develop intelligent monitoring that detects anomalies, generates compliance reports, and auto-remediates issues
  • Create reusable agent templates and workflows that adapt to changing regulations while maintaining performance and reducing manual compliance
  • Contribute to distributed agent ecosystem emphasizing reliability through self-healing mechanisms, automated failovers, and dynamic resource allocation for scalability
  • Use AI coding agents to accelerate implementation, test coverage, and debugging, validating every result for correctness, privacy, and cost

Minimum Qualifications

  • B.S. in Computer Science or related field with 5+ years of software development experience, including exposure to ML/AI applications
  • Production experience in Python, Java, Rust or similar languages, with familiarity in data processing and API development
  • Understanding of distributed systems and cloud platforms, including CAP theorem tradeoffs and basic ML model deployment
  • Experience with containerization (Docker), orchestration (Kubernetes), and infrastructure as code (e.g. Terraform, CloudFormation)
  • Proficiency in CI/CD pipelines and DevOps practices using Git, GitHub Actions/Jenkins/GitLab CI, with experience in automated testing and deployment workflows
  • Familiarity with observability and monitoring tools (e.g. Prometheus, Grafana, DataDog) and logging frameworks for production systems
  • Basic knowledge of machine learning concepts and MLOps, including data pipelines, model versioning, and experiment tracking tools

Preferred Qualifications

  • Expertise in open source data analytics and governance platforms: architecture, deployment, and performance tuning of Datahub, Apache Spark, Flink, Hive, Hadoop/HDFS, and Iceberg Rest Catalog
  • Experience building multi-agent AI systems: proficiency with LangChain, LangGraph, or AutoGen frameworks; strong prompt engineering and LLM integration skills; ability to design event-driven architectures for autonomous workflows
  • Skills in integration and communication layers: implement MCP servers and APIs using Python, REST/GraphQL, and message queuing (e.g. Kafka, RabbitMQ); experience with modern data platforms including Snowflake, Databricks, and vector databases
  • MLOps and observability capabilities: deploy containerized AI systems with comprehensive monitoring; track experiments using MLflow or Weights & Biases; implement distributed tracing for agent workflows and model performance