Senior Manager, Core Infrastructure Engineering
The OCI AI Infrastructure Network Operations team operates and improves the high-performance RDMA/RoCE network fabrics powering OCI’s largest AI, GPU, and HPC workloads.
As a Senior Manager, you will lead a team responsible for building, operating, and scaling these critical network fabrics and supporting systems. You will combine deep networking expertise in RDMA/RoCE, Clos fabrics, congestion control, telemetry, and performance troubleshooting with strong software engineering and people leadership.
You will drive automation, monitoring, resiliency, and operational readiness while partnering across Network Availability, Automation, Monitoring, GNOC, hardware engineering, and service teams. Your team will improve network performance and availability, resolve complex customer issues, and build fault-tolerant systems that support AI infrastructure at global cloud scale.
Qualifications
Disclaimer:Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.
Responsibilities
As a Senior Manager in the AI Infrastructure Network Operations organization, you will:
- Lead and develop a team of engineers responsible for RDMA/RoCE fabric operations, performance, automation, and troubleshooting across OCI’s AI/HPC infrastructure.
- Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems supporting hyperscale workloads.
- Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting to improve network performance and availability.
- Guide the architecture and development of operational tools, automation platforms, monitoring systems, and infrastructure services, including Infrastructure as Code (IaC).
- Drive improvements in resiliency, observability, testing, and automation while simplifying and scaling operational workflows.
- Lead operational readiness, customer escalations, NOC events, and complex production incidents, coordinating resolution across networking, software, hardware, and operations teams.
- Define team roadmaps and data-driven KPIs focused on fabric health, engineering efficiency, operational backlog, customer impact, performance, and service availability.
- Partner with Network Availability, Network Automation, Network Monitoring, GNOC, deployment, hardware, and service teams to deliver reliable infrastructure at cloud scale.
- Ensure operational planning, security, compliance, change management, staffing, on-call coverage, and service-level expectations are met.
- Drive continuous improvement in engineering practices, processes, tooling, and operational efficiency.
- Attract, mentor, and develop engineers across networking, software development, automation, and distributed systems while building a high-performing engineering organization.
- Participate in the manager on-call rotation and provide technical and organizational leadership during high-severity incidents.
Preferred Experience
- Strong background in operating or building network for large-scale cloud.
- Experience with RDMA/RoCE, GPU/HPC networking, Clos fabrics, congestion management, telemetry, and performance debugging.
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.