SDM, ML Data Infra, Amazon Traffic Engineering
We are seeking an experienced Software Development Manager to lead a team of Software Development Engineers (SDEs) and Data Engineers (DEs) building the Core Data Infrastructure that underpins our ML and Science initiatives for Bot Management. You will own the end-to-end data platform—from source ingestion through transformation, feature engineering, and serving—ensuring Science and ML Platform teams have reliable, scalable, and timely access to the data they need for training, evaluation, and inference.
This is a high-impact leadership role. Our Science teams are building increasingly sophisticated models—and each requires different data formats, latencies, and serving patterns. Your team will be the backbone that makes this possible: ingesting billions of events from diverse source systems, building production-grade pipelines that transform raw signals into ML-ready feature groups, and operating the Feature Store that serves these features consistently across all model types. You will partner closely with Science leadership to translate model requirements into data infrastructure investments, and with ML Platform leadership to ensure seamless integration between your data layer and their training/inference systems. This role demands a leader who can navigate ambiguity across organizational boundaries, drive technical alignment between data engineering, software engineering and science teams, and build systems that scale with the rapid pace of model innovation.
Key job responsibilities
Data Infrastructure & Feature Engineering — Own the Feature Store, feature pipelines, and data serving layer. Build versioned feature groups across multiple storage backends (S3 for tabular, OpenSearch for embeddings) and production pipelines that transform disparate datasets into ML-ready features for Science teams.
Streaming & Real-Time Systems — Design and operate Apache Flink applications and live stream data processing for near real-time feature computation. Build event-driven architectures leveraging Kinesis and Kafka to support low-latency bot detection signals.
Data Pipelines & Ingestion — Own batch and near real-time pipelines spanning Trails (raw + aggregated) and Non-Trails sources (AIT, Clickstream, Customer Segmentations, OPS). Evolve pipelines from Cradle/POC to production-grade using AWS Glue. Implement data drift detection and governance frameworks.
Science & ML Platform Partnership — Serve as the primary data infrastructure partner to Applied Scientists and ML Platform. Define data contracts and SLAs, participate in model design reviews, and ensure the Feature Store integrates seamlessly with training and inference systems.
People Leadership — Recruit, develop, and retain a high-performing team of SDEs and DEs. Set goals, manage roadmaps, and foster a culture of operational excellence.
About the team
Traffic Engineering's Bot Management organization protects Amazon's ecosystem by detecting and mitigating automated threats at scale. Our Core ML Data Infrastructure team is responsible for building and operating the foundational data infrastructure that powers bot detection, AI agent identification, and content exfiltration defense. We are building a unified, model-agnostic, production-grade ML platform that brings together training, evaluation, and inference pipelines into a cohesive system serving multiple model types across the organization.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
CAN, BC, Vancouver - 171,400.00 - 286,200.00 CAD annually
This is a high-impact leadership role. Our Science teams are building increasingly sophisticated models—and each requires different data formats, latencies, and serving patterns. Your team will be the backbone that makes this possible: ingesting billions of events from diverse source systems, building production-grade pipelines that transform raw signals into ML-ready feature groups, and operating the Feature Store that serves these features consistently across all model types. You will partner closely with Science leadership to translate model requirements into data infrastructure investments, and with ML Platform leadership to ensure seamless integration between your data layer and their training/inference systems. This role demands a leader who can navigate ambiguity across organizational boundaries, drive technical alignment between data engineering, software engineering and science teams, and build systems that scale with the rapid pace of model innovation.
Key job responsibilities
Data Infrastructure & Feature Engineering — Own the Feature Store, feature pipelines, and data serving layer. Build versioned feature groups across multiple storage backends (S3 for tabular, OpenSearch for embeddings) and production pipelines that transform disparate datasets into ML-ready features for Science teams.
Streaming & Real-Time Systems — Design and operate Apache Flink applications and live stream data processing for near real-time feature computation. Build event-driven architectures leveraging Kinesis and Kafka to support low-latency bot detection signals.
Data Pipelines & Ingestion — Own batch and near real-time pipelines spanning Trails (raw + aggregated) and Non-Trails sources (AIT, Clickstream, Customer Segmentations, OPS). Evolve pipelines from Cradle/POC to production-grade using AWS Glue. Implement data drift detection and governance frameworks.
Science & ML Platform Partnership — Serve as the primary data infrastructure partner to Applied Scientists and ML Platform. Define data contracts and SLAs, participate in model design reviews, and ensure the Feature Store integrates seamlessly with training and inference systems.
People Leadership — Recruit, develop, and retain a high-performing team of SDEs and DEs. Set goals, manage roadmaps, and foster a culture of operational excellence.
About the team
Traffic Engineering's Bot Management organization protects Amazon's ecosystem by detecting and mitigating automated threats at scale. Our Core ML Data Infrastructure team is responsible for building and operating the foundational data infrastructure that powers bot detection, AI agent identification, and content exfiltration defense. We are building a unified, model-agnostic, production-grade ML platform that brings together training, evaluation, and inference pipelines into a cohesive system serving multiple model types across the organization.
Basic Qualifications
- 3+ years of engineering team management experience
- 7+ years of working directly within engineering teams experience
- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- 8+ years of leading the definition and development of multi tier web services experience
- Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
- Experience partnering with product or program management teams
Preferred Qualifications
- Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy
- Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. As a total compensation company, Amazon's package may include other elements such as sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, basic life & AD&D insurance), Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and other resources to improve health and well-being. We thank all applicants for their interest, however only those interviewed will be advised as to hiring status.
CAN, BC, Vancouver - 171,400.00 - 286,200.00 CAD annually