Technical Operation Engineer

AppleApplyPublished 2 hours agoFirst seen 2 hours ago
Apply

Summary

Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.

Customer Systems is part of IS&T and drives the technology behind Apple's customer support experience — from contact center operations to the software powering the iconic Genius Bar. The team also builds and operates AppleCare's online support platform, which handles 6 billion visits per year, delivering seamless, high-quality support to Apple customers around the globe.

Description

The Technical Operations Engineer will support and improve business-critical Customer Systems
services used across Apple’s global customer-support organizations. The role combines advanced
production support, automation, reliability engineering and Operational Excellence.
The engineer will be responsible for service health, incident response, root-cause analysis and
continuous improvement of operational processes, tooling and platform reliability.

Responsibilities

  • Operate globally distributed, business-critical customer-support applications and platforms.
  • Monitor service availability, latency, capacity, throughput and error rates.
  • Lead troubleshooting of complex application, infrastructure, database, integration and
  • network issues.
  • Participate in or lead major-incident response, service restoration and stakeholder
  • communication.
  • Perform root-cause analysis and drive permanent corrective actions.
  • Identify recurring incidents, operational gaps and reliability risks, and convert them into
  • improvement initiatives.
  • Build automation, diagnostics and self-service tools to reduce manual operational work.
  • Develop dashboards, health checks, actionable alerts and automated remediation.
  • Drive Operational Excellence across incident, problem, change, release and capacity
  • management.
  • Define and track service metrics, operational KPIs, SLIs and SLOs.
  • Support Java applications, REST APIs, microservices and asynchronous integrations.
  • Coordinate production releases, configuration changes, deployment validation and rollback
  • planning.
  • Lead production-readiness and operational-readiness reviews for new services and major
  • changes.
  • Partner with software engineering teams to improve application operability, resilience and
  • supportability.
  • Maintain runbooks, troubleshooting guides and service documentation.
  • Support capacity planning, resilience testing and disaster-recovery exercises.
  • Participate in on-call rotations and support critical launches and business events.
  • Mentor junior engineers and promote a culture of ownership and continuous improvement.

Minimum Qualifications

  • 7+ years of experience in Technical Operations, production engineering, application support
  • or Site Reliability Engineering.
  • Strong troubleshooting skills across Linux, applications, APIs, databases and networks.
  • Experience deploying, monitoring and supporting Java-based enterprise applications.
  • Experience with scripting or software development using Python, Java, Go or shell.
  • Experience with logging, monitoring and observability tools such as Splunk.
  • Understanding of HTTP, DNS, TCP/IP, load balancing and distributed systems.
  • Experience with incident, problem, change and release-management processes.
  • Demonstrated ability to use operational metrics and incident trends to drive service
  • improvements.
  • Strong communication, documentation and stakeholder-management skills.

Preferred Qualifications

  • Experience with Kubernetes, container platforms and Alicloud services.
  • Knowledge of Oracle, PostgreSQL, MongoDB, Kafka or similar technologies.
  • Experience building operational automation and automated remediation.
  • Experience with high availability, disaster recovery and production-readiness practices.