Production Network Engineer
Description
Meta's global network infrastructure underpins billions of user connections and powers some of the world's most demanding AI and distributed computing workloads. The Production Network Engineering team is responsible for the design, deployment, operation, and continuous improvement of Meta's large-scale production data center networks. In this role, you will drive strategy and execution across complex network systems, lead cross-functional initiatives to improve reliability and performance, and apply deep subject matter expertise to solve problems at a scale few networks in the world match.
Responsibilities
Lead the design and delivery of large-scale production network projects spanning data center fabrics, backbone infrastructure, and edge network services Define and drive technical strategy and roadmap contributions for network reliability, capacity, and operational efficiency across multiple teams Develop and implement automation frameworks to reduce manual operational overhead and accelerate network design synthesis, build processes, and configuration management Own cross-functional incident response and post-incident review processes, driving root cause analysis and systemic improvements to reduce recurrence Identify and resolve complex network performance, routing, and reliability issues across multi-vendor, multi-protocol production environments Collaborate with network architecture, capacity planning, and software engineering teams to align infrastructure investments with evolving AI and product demand forecasts Establish and refine operational standards, runbooks, and monitoring frameworks to improve network observability and reduce mean time to detection and resolution Contribute to team-level goals by defining scalable approaches to network operations and influencing tooling and platform decisions across engineering teams Advise other engineers on network engineering best practices, operational discipline, and systems thinking across the production environment Leverage AI-integrated workflows to accelerate network analysis, anomaly detection, and documentation, sharing learnings to scale adoption across the team
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 6+ years of experience in production network engineering, including design, deployment, and operations of large-scale data center or backbone network infrastructure Experience with routing protocols and network technologies including BGP, OSPF, IS-IS, MPLS, and large-scale Ethernet fabrics in a production environment Experience developing and deploying network automation using scripting or programming languages such as Python to manage configuration, provisioning, or monitoring at scale Experience leading cross-functional network reliability or infrastructure projects, including driving incident response, root cause analysis, and systemic remediation Experience influencing technical decisions and network strategy across engineering teams through written proposals, design reviews, and stakeholder alignment Experience operating and troubleshooting networks at hyperscale, including multi-vendor spine-leaf data center fabrics and global backbone interconnects Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies Familiarity with software-defined networking principles, network operating system internals, or co-development of network management platforms with software engineering teams Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) Experience with network observability platforms, traffic engineering, and capacity modeling in high-throughput production environments Experience integrating AI tools to redesign network operations workflows and deliver measurable improvements in efficiency or reliability outcomes
Compensation: $135,000/year to $191,000/year + bonus + equity + benefits