Microsoft Logo

Microsoft

Principal Software Engineering Manager

Posted 3 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
143K-304K Annually
Senior level
Remote
Hiring Remotely in United States
143K-304K Annually
Senior level
Lead and grow a software engineering team to design, build, and operate high-performance, scalable, observable networking systems for Azure AI/HPC infrastructure. Drive architecture, reliability, testing, automation, and cross-team collaboration to support large-scale distributed training and inference workloads.
The summary above was generated by AI
Overview

The HPC/AI (High-Performance Computing and Artificial Intelligence) organization is on a mission to build the next generation of distributed AI supercomputers - systems that deliver unprecedented computational power, scalability, and reliability to accelerate breakthroughs in artificial intelligence. Our teams design and develop world-class AI infrastructure that enables large-scale model training and inference, forming the backbone of Microsoft’s AI innovation.

As a Principal Software Engineering Manager, you will lead a team building foundational components of Azure’s AI networking infrastructure—powering some of the largest and most complex distributed training systems in the world. This is a rare opportunity to work at the intersection of AI, cloud infrastructure, and high-performance networking, driving innovation across hardware and software boundaries. With the explosive growth of generative AI and the demand for low-latency, high-bandwidth systems, your work will directly impact the scale, performance, and reliability of Microsoft’s AI platforms.
You will lead the design, development, and deployment of high-performance, scalable, and observable networking systems that connect AI accelerators at massive scale. The role requires deep technical acumen, strategic thinking, and a passion for engineering excellence. You’ll collaborate across Microsoft teams to define architecture, deliver solutions to complex infrastructure challenges, and ensure our systems meet the evolving needs of AI workloads.
If you’re passionate about building large-scale distributed systems, pushing the boundaries of AI infrastructure, and leading teams that shape the future of supercomputing, we invite you to join us on this journey to define the next era of AI at Microsoft.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. 


Responsibilities
  • Hire, manage, and grow a high-performing team of software engineers, fostering a culture of excellence, inclusion, and innovation.
  • Lead the design and development of large-scale distributed systems and services that power Azure’s AI infrastructure.
  • Drive engineering planning and execution while ensuring alignment with organizational OKRs and long-term strategy.
  • Establish lean, scalable, and efficient processes that promote innovation and engineering rigor.
  • Deliver best-in-class engineering by ensuring services and components are modular, secure, reliable, diagnosable, observable, and reusable.
  • Improve test coverage, automation, and integration testing to proactively identify and resolve reliability gaps.
  • Ensure live-site reliability and service health through robust monitoring, telemetry, and automation.
  • Collaborate across Microsoft and partner organizations to deliver cohesive, end-to-end infrastructure solutions.
  • Apply data-driven insights to optimize performance, scalability, and customer satisfaction.

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience. 

Other Requirements: 

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:  
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter. 
Preferred Qualifications:
  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.
  • 4+ years people management experience.
  • 1+ years experience building and operating networking infrastructure for hyperscale datacenters or AI clusters.
  • 1+ years hands-on experience with networking technologies in AI-specific hardware (e.g., InfiniBand, ROCE, MRC, NVLink, UALink).
  • 10+ years of professional software design and development experience in large-scale distributed systems.
#azurecorejobs

Software Engineering M5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs

2 Days Ago
Remote
United States
143K-331K Annually
Senior level
143K-331K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and grow an engineering team building core VM and container platform capabilities for Azure. Drive technical strategy, execution, incident resolution, cross-team coordination, and scalable, secure platform innovations.
Top Skills: AzureCC#C++Confidential ComputingContainersDevice DriversDistributed SystemsEdgeFirmwareJavaJavaScriptOperating SystemsPythonVirtual MachinesVirtualizationWindows
4 Days Ago
In-Office or Remote
2 Locations
166K-331K Annually
Senior level
166K-331K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and scale teams building offensive and defensive AI agents and the ML, data, and safety systems that power them. Drive strategy, production-grade engineering, safety and governance, cross-team collaboration, and hiring/leadership development to deliver secure, reliable agentic security capabilities for customers.
Top Skills: Agentic SystemsCC#C++Defensive SecurityJavaJavaScriptLarge Language ModelsMachine LearningOffensive SecurityPythonResponsible Ai
20 Days Ago
In-Office or Remote
MD, USA
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and manage engineering efforts for secure, high-performance Azure data transfer services. Define requirements, design, implement, and optimize code, own on-call responsibilities, drive reliability/observability, and collaborate with stakeholders to deliver scalable solutions.
Top Skills: AzureAzure Data TransferCC#C++JavaJavaScriptMicrosoft CloudPython

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account