Principal Network Developer

Oracle•Published 23 days ago•First seen 4 hours ago

We are the AI Infrastructure - Network Operations team at OCI (Oracle Cloud Infrastructure).  We support and operate the RDMA/RoCE/InfiniBand network fabrics for OCI's largest AI and HPC customers.  These fabrics are the foundation underneath OCI's AI, GPU and HPC services, and support major tier-0 vendors in the generative AI industry. If you're running an AI workload at OCI, we're running the RDMA network underneath your workload.

A Network Operations Engineer on our team supports the design, deployment, and operations of a large-scale global Oracle cloud computing environment (Oracle Cloud Infrastructure - OCI).  Primarily focused on operation and support of RDMA/RoCE/InfiniBand network fabrics and systems, through a combination of a deep network understanding and automation skills to operate a production environment.  As OCI is a cloud-based network with a global footprint, this support will include hundreds of thousands of network devices supporting millions of servers, connected over a mix of dedicated backbone infrastructure and the Internet.

Responsibilities

     Key Responsibilities

  • Lead network lifecycle management initiatives by defining technical objectives, delivery plans, and implementation procedures for large-scale network infrastructure projects.
  • Translate high-level network architectures into detailed designs and deployment plans while ensuring scalability, reliability, and operational readiness.
  • Serve as the technical lead for moderately complex network projects, coordinating the efforts of multiple engineers across design, deployment, automation, and operational support.
  • Design, implement, and support network solutions across data center, backbone, cloud, and service provider environments.
  • Partner with service owners and infrastructure teams to ensure network solutions are fully integrated with monitoring, observability, automation, and operational support systems.
  • Act as a Tier 2 and specialized escalation point for network incidents, driving root cause analysis, corrective actions, and long-term reliability improvements.
  • Lead the investigation and resolution of complex network issues and large-scale service-impacting events.
  • Develop automation solutions, tools, and scripts to improve operational efficiency, network reliability, deployment consistency, and incident response.
  • Contribute to the design and delivery of network automation frameworks and operational tooling.
  • Collaborate closely with product teams, program managers, network leadership, and PMO organizations to align infrastructure capabilities with product and service requirements.
  • Partner with vendor engineering teams and account managers to troubleshoot issues, evaluate new technologies, and drive operational improvements.
  • Participate in hardware evaluations, RFQ/RFP processes, and adoption of new networking technologies and platforms.
  • Drive technology decisions that support business, product, and service objectives.
  • Mentor junior engineers through technical guidance, troubleshooting support, design reviews, and knowledge sharing.
  • Contribute to engineering best practices, documentation standards, operational excellence, and continuous improvement initiatives.

Preferred Skills & Experience 

  • 12+ years of experience. 
  • Bachelor’s degree (Master’s preferred) in Computer Science, Electrical Engineering or related field.
  • Strong experience with large-scale network operations, design, and troubleshooting.
  • Expertise in routing and switching technologies, including BGP, OSPF, EVPN-VXLAN, MPLS, and data center networking. Prion experience with RDMA/RoCE/InfiniBand would be a plus.
  • Experience with network automation using Python, Ansible, APIs, or similar technologies, building AI agents to automate.
  • Strong understanding of observability, monitoring, telemetry, and incident management.
  • Experience working with cloud infrastructure, hyperscale environments, or large-scale distributed systems.
  • Ability to lead technical projects and influence outcomes across multiple teams.
  • Strong written and verbal communication skills with the ability to work effectively across engineering, operations, and leadership teams.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.