Software Engineer II

MicrosoftApplyPublished 17 hours agoFirst seen 16 hours ago
Apply
Overview

Microsoft Azure Storage is a highly distributed, massively scalable, and ubiquitously accessible cloud storage platform. Azure storage already runs at Exascale (storing Exabytes of data) and we will scale our designs over the next decade to support Zettascale (storing Zettabytes of data).  

Azure Storage is growing its development team focused on providing the world's lowest cost cloud-based storage. We are looking for software engineers interested in helping us fulfil our vision of providing a storage service capable of satisfying the worlds ever increasing demand for cloud-based storage for many years to come on the application team. If cloud‑scale distributed systems, deep engineering challenges, and building infrastructure that powers the world’s data excite you, Azure Storage is the perfect fit. 

 
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. 


Responsibilities
    • Monitor fleet health, capacity, reliability, firmware compliance, and operational readiness. Design and implement end-to-end fleet automation platforms. Build maintainable and extensible automation frameworks that support multiple fleet types and hardware generations 

    •  Identify systemic health issues and proactively drive recovery and remediation programs.  

    •  Analyse telemetry, incident trends, and fleet-wide health indicators to identify opportunities for automation.  

    •  Lead investigations of production incidents and drive root-cause analysis.  

    •  Design and implement safe recovery workflows using related automation platforms.  

    •  Drive reduction of offline capacity, unhealthy nodes, stale firmware, and operational toil.  

    •  Improve service reliability through self-healing and automated recovery capabilities.  

    •  Create dashboards, monitoring solutions, metrics, and health signals that enable data-driven decision making.  

    • Independently use appropriate artificial intelligence (AI) tools and practices across the software development lifecycle (SDLC) in a disciplined manner. 

    • Acts as a Designated Responsible Individual (DRI) and guides other engineers by developing and following the playbook, working on call to monitor system/product/service for degradation, downtime, or interruptions, alerting stakeholders about status and initiates actions to restore system/product/service for simple and complex problems when appropriate. 

    • Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of storage fleet while also driving consistency in monitoring and operations at scale. 

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 3+ years technical engineering experience with coding in languages including, but not limited to, , C#, C++, Powershell, or Python OR equivalent experience.
  • 3+ years of experience in automation, performance, and building highly available distributed systems at scale using scripting languages such as Bash, Python, and PowerShell. Production level development in languages such as C++, C#, 

    Experience with managing and writing services on top of cloud environments such as Azure, AWS, or GCP 

    Experience with large-scale storage, Datacentre Hardware infrastructure and automation related (and not limited to)to fleet health, security, firmware consistency. 

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

Preferred Qualifications:

  • Master's Degree in Computer Science or related technical field AND 3+ years technical engineering experience with coding in languages including, but not limited to, C++, C#, Bash, Powershell, or Python
    • OR Bachelor's Degree in Computer Science or related technical field AND 5+ years technical engineering experience with coding in languages including, but not limited to, C++, C#, Bash, Powershell, or Python
    • OR equivalent experience
    • Systems Engineering experience 

    • 3+ years of experience designing, building and running large scale and highly available Cloud services or distributed systems. 

    • Experience in a cloud stack and leveraging cloud architecture, applying site reliability principles and/or demonstrating sensitivity to operational concerns 

    • Advanced server hardware component knowledge 

#azurecorejobs

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.