Senior Manager, Storage Engineering
NVIDIA is the world leader in accelerated computing, inventing the GPU and pioneering the technologies that power modern AI, data science, autonomous vehicles, robotics, and high-performance computing. From the chips that run the world's fastest supercomputers to the software platforms that accelerate breakthroughs in science and industry, NVIDIA sits at the center of the most transformative technology shift of our time.
NVIDIA's IT Storage Engineering team architects, designs, deploys, and manages petabyte-scale storage infrastructure that serves as the foundation for some of the most demanding workloads in the industry. Our team is a critical enabler across a broad set of internal organizations — Hardware Engineering & Chip Design, Software Engineering, Manufacturing, AI/ML Research, and IT Operations — delivering reliable, high-performance storage solutions that keep NVIDIA's innovation engine running at full speed. As the Storage Deployment Manager, you will lead a team of storage engineers and collaborate across NVIDIA's global IT organization to ensure that our storage infrastructure scales with the company's relentless pace of innovation. You will own the full deployment lifecycle — from capacity planning and hardware procurement through rack-and-stack, configuration, integration, and ongoing operational health — across on-premises data centers and cloud service providers (CSPs).
What You’ll Be Doing:
- Lead petabyte-scale storage deployments across NVIDIA's on-premises data centers and major CSPs (AWS, Azure, GCP), owning the end-to-end lifecycle from design and procurement through physical installation, configuration, and production hand-off.
- Engineer and maintain automation pipelines that integrate deployed storage systems with interdependent tooling — including CMDB (asset tracking and dependency mapping), configuration management platforms, and observability stacks, ensuring a single source of truth across the entire storage fleet.
- Define and drive continuous improvement initiatives focused on data center efficiency: optimizing DC power consumption (PUE impact, drive density, power shelf utilization) and minimizing rack space footprint through high-density hardware selection and intelligent workload placement strategies.
- Develop and maintain self-service tools and dashboards that enable internal customers (EDA, Manufacturing, Software Engineering teams) to track real-time storage capacity availability, consumption trends, and projected growth, reducing friction and improving planning accuracy.
- Coordinate and partner with partner infrastructure teams — networking, compute, cloud, and Data center to present a unified capacity view, resolve cross-domain bottlenecks, and contribute to NVIDIA's holistic infrastructure capacity planning process.
- Manage vendor relationships and lead hardware refresh and EOL planning cycles across the storage portfolio, balancing performance requirements, cost efficiency, and supply chain constraints.
- Recruit, mentor, and grow a team of storage deployment engineers; establish engineering standards, runbooks, and on-call practices to ensure a high operational bar.
- Partner with architecture and security teams to evaluate new storage technologies, drive POCs, and translate findings into production-ready deployment standards.
- Own and engineer the full hardware and data lifecycle — from procurement, rack integration, and initial provisioning through capacity expansion, decommissioning, and secure data destruction — ensuring compliance, auditability, and zero unplanned data loss at every stage.
- Define and govern data lifecycle management policies in close collaboration with key stakeholders across Chip Design and Software Engineering teams, aligning retention, tiering, archival, and deletion standards to business workflows, improving overall storage efficiency, and reducing cost by eliminating stale or redundant data across the fleet.
What We Need to See:
- BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
- 12+ years of overall experience in large-scale storage architecture, operations, production engineering, or infrastructure. 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
- Deep, protocol-level knowledge of enterprise storage systems spanning block (/NVMe-oF), file (NFS, SMB/CIFS), object (S3-compatible), and high-performance parallel file systems (GPFS/IBM Spectrum Scale, Lustre).
- Hands-on experience deploying and operating solutions from multiple major storage vendors, including NetApp (ONTAP, StorageGRID), Pure Storage (FlashArray, FlashBlade), Cloudian HyperStore, and DDN (EXAScaler, A³I).
- Solid understanding of storage hardware internals — drive types and endurance profiles (NVMe, SAS, SATA, QLC/TLC/MLC NAND), controller architectures, shelf and enclosure design, cabling standards, and failure domain planning.
- Working knowledge of bare-metal server hardware (rack units, HBA/NIC selection, BMC/iDRAC/iLO, firmware management) and data center network hardware relevant to storage connectivity (ToR switches, fiber and copper interconnects, SFP/QSFP optics).
- In-depth expertise in observability tooling: building and maintaining Prometheus exporters, Grafana dashboards, and alerting rulesets for storage fleet health, performance SLOs, and capacity burn-rate tracking.
- Strong configuration management skills using Ansible (playbook authoring, role design, inventory management, Ansible Tower/AAP) for automated provisioning and day-2 operations of storage systems at scale.
- Scripting and automation proficiency in Python and/or Bash; experience integrating with REST APIs to drive CMDB updates, provisioning workflows, and reporting pipelines
- Proven track record of managing large-scale, multi-vendor storage environments (100 PB+) in a fast-paced, high-availability production setting.
Ways to Stand Out from the Crowd:
- Deep knowledge of Kubernetes storage integrations — CSI driver operations, persistent volume lifecycle management, StorageClass design for stateful workloads, and experience running storage-intensive workloads on K8s at scale.
- HPC systems expertise: experience deploying and tuning storage for HPC clusters, including parallel file system performance optimization (stripe tuning, client-side caching, OST/MDT balancing) and co-designing storage architectures with HPC compute and fabric teams.
- Familiarity with modern storage technologies (NVMe, RDMA, DPUs) and their impact on system performance, including kernel-level concepts around I/O subsystems and volumes
- Familiarity with HPC job schedulers such as IBM LSF or Slurm — understanding how job scheduling behavior (burst patterns, checkpoint I/O, scratch-space usage) drives storage design decisions, and experience implementing storage-aware scheduling policies.
- Experience building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud setups (for example AWS S3, Azure Blob, or Google Cloud Storage), as well as on-prem infrastructure
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 248,000 USD - 396,750 USD.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 31, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.