Lead Infrastructure Engineer
At EA, we believe games are powerful because they bring together multiple ways people engage: play, watch, create, and connect. And increasingly, the biggest entertainment platforms aren't just places to consume content — they're places where communities build.
Creator-made content is already a proven part of EA's history — from community creation tools in Battlefield to The Gallery in The Sims 4. We believe new creative technologies and tools will expand how players engage with and contribute to our experiences, supported by thoughtful product design, safety systems, and global reach. Our focus is on enabling more players to participate in creative expression by making creation easier, safer, and more rewarding.
As Lead Infrastructure Engineer, you will own the GPU fleet our researchers train on including capacity, scheduling, diagnostics, and support. You will set technical direction for GPU operations and infrastructure architecture. You will additionally lead Infrastructure as Code setup, granting permissions, and debugging infrastructure problems.
This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver.
You will report to the Head of Data and Infrastructure.
Responsibilities:
- You will own GPU fleet operations across our AWS estate.
- You will build the scheduling layer from zero.
- You will diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels.
- You will run researcher support as a first-class product including holding office hours, owning the support channel, and driving the recurring causes out of existence with self-service tooling, preflight checks, and documentation
- You will instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.
- You will partner with our external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.
- You will author runbooks, decision records, and onboarding docs.
Qualifications:
- 8+ years of experience operating production infrastructure, with deep, current, hands-on AWS depth — EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3 and FSx
- Experience scheduling, diagnosing, and managing GPUs in AWS specifically
- Experience operating GPU fleets at 1000+ GPU scale
- Expertise in scripting and automation with Python, PowerShell, bash, or equivalent
- Expertise in infrastructure as code (Terraform or equivalent)
- Familiarity with a GPU scheduling or orchestration layer (like Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot)
- Observability practice including Grafana, Prometheus, or equivalent
COMPENSATION AND BENEFITS
The ranges listed below are what EA in good faith expects to pay applicants for this role in these locations at the time of this posting. If you reside in a different location, a recruiter will advise on the applicable range and benefits. Pay offered will be determined based on a number of relevant business and candidate factors (e.g. education, qualifications, certifications, experience, skills, geographic location, or business needs).
PAY RANGES
* British Columbia (depending on location e.g. Vancouver vs. Victoria) *$169,500 - $242,600 CAD* California (depending on location e.g. Los Angeles vs. San Francisco) *$193,100 - $296,500 USDPay is just one part of the overall compensation at EA.
In the US, we offer a package of benefits including paid time off (3 weeks per year to start), 80 hours per year of sick time, 16 paid company holidays per year, 10 weeks paid time off to bond with baby, medical/dental/vision insurance, life insurance, disability insurance, and 401(k) to regular full-time employees. Certain roles may also be eligible for bonus and other incentive programs.
For Canada, we offer a package of benefits including vacation (3 weeks per year to start), 10 days per year of sick time, paid top-up to EI/QPIP benefits up to 100% of base salary when you welcome a new child (12 weeks for maternity, and 4 weeks for parental/adoption leave), extended health/dental/vision coverage, life insurance, disability insurance, retirement plan to regular full-time employees. Certain roles may also be eligible for bonus and other incentive programs.
About Electronic Arts
We’re proud to have an extensive portfolio of games and experiences, locations around the world, and opportunities across EA. We value adaptability, resilience, creativity, and curiosity. From leadership that brings out your potential, to creating space for learning and experimenting, we empower you to do great work and pursue opportunities for growth.
We adopt a holistic approach to our benefits programs, emphasizing physical, emotional, financial, career, and community wellness to support a balanced life. Our packages are tailored to meet local needs and may include healthcare coverage, mental well-being support, retirement savings, paid time off, family leaves, complimentary games, and more. We nurture environments where our teams can always bring their best to what they do.
Electronic Arts is an equal opportunity employer. All employment decisions are made without regard to race, color, national origin, ancestry, sex, gender, gender identity or expression, sexual orientation, age, genetic information, religion, disability, medical condition, pregnancy, marital status, family status, veteran status, or any other characteristic protected by law. We will also consider employment qualified applicants with criminal records in accordance with applicable law. EA also makes workplace accommodations for qualified individuals with disabilities as required by applicable law.

