Antares Capital LP Logo

Antares Capital LP

Vice President, Reliability Engineering & Technology Operations

Posted 6 Days Ago
Be an Early Applicant
In-Office
New York, NY, USA
175K-225K Annually
Senior level
In-Office
New York, NY, USA
175K-225K Annually
Senior level
Leads reliability engineering and technology operations, overseeing production operations, Azure and Kubernetes platforms, observability, incident response, automation, resiliency, disaster recovery, vendors, and operational transformation. Establishes reliability practices, SLOs, error budgets, service ownership, and performance metrics while leading engineering-minded teams and partnering with business and technology stakeholders. Drives AI-enabled operations, platform resilience, continuous improvement, and executive incident management.
The summary above was generated by AI
About Antares Capital

Antares Capital is a leading alternative credit manager and a trusted financing partner to private equity sponsors and middle-market companies. We are committed to building resilient, scalable, and modern technology platforms that support our business and clients.

As part of our continued technology transformation, we are seeking a Vice President, Reliability Engineering & Technology Operations to lead the evolution of our production operations, reliability engineering, observability, and operational automation capabilities.

This is a strategic leadership role for an engineering-minded leader who thrives at the intersection of software engineering, cloud infrastructure, platform operations, and operational excellence. The ideal candidate combines strong technical depth with exceptional execution skills and has experience building highly reliable systems while leading teams through modernization and transformation initiatives.

The Opportunity

Technology is central to Antares' growth strategy. We are investing heavily in cloud platforms, engineering excellence, AI-enabled workflows, automation, and modern operational practices.

As the leader of Reliability Engineering & Technology Operations, you will be responsible for the availability, performance, scalability, and resilience of critical business platforms. You will partner closely with Engineering, Infrastructure, Cybersecurity, Data, and Business stakeholders to ensure our systems remain secure, observable, scalable, and operationally mature.

You will help shape the future of technology operations by introducing reliability engineering practices, expanding observability, leveraging AI-driven operational capabilities, and reducing operational overhead through automation.

This role is ideal for someone who has grown through engineering, cloud, platform, infrastructure, or DevOps leadership roles and understands how to bridge engineering and operations to deliver exceptional business outcomes.

Key ResponsibilitiesReliability Engineering Leadership
  • Establish and lead Antares' Reliability Engineering function.
  • Define and implement strategies that improve system reliability, resiliency, scalability, and operational excellence.
  • Partner with engineering teams to embed reliability practices throughout the software development lifecycle.
  • Drive adoption of modern operational practices including service ownership, operational readiness reviews, SLOs, error budgets, and post-incident learning.
Technology Operations
  • Lead production operations across critical business applications and technology platforms.
  • Establish clear support models, escalation paths, ownership boundaries, and service management processes.
  • Oversee operational readiness, release support, change management, and platform health.
  • Continuously improve operational maturity through metrics, automation, process simplification, and engineering collaboration.
Azure Cloud & Kubernetes Operations
  • Provide technical leadership for cloud-based platforms running in Microsoft Azure.
  • Partner with Infrastructure and Engineering teams to optimize reliability, scalability, and operational efficiency.
  • Support containerized workloads and Kubernetes-based environments.
  • Drive best practices around cloud architecture, capacity planning, platform resilience, security, governance, and cost optimization.
  • Ensure cloud platforms are designed and operated to meet business continuity and availability objectives.
Incident Response & Problem Management
  • Serve as the executive incident leader during major production events.
  • Coordinate cross-functional teams during outages and high-severity incidents.
  • Manage communications with business stakeholders and technology leadership.
  • Establish disciplined root cause analysis processes and ensure corrective actions are executed.
  • Drive long-term reduction in recurring incidents and operational risk.
AI-Powered Operations & Automation
  • Champion an automation-first and AI-enabled approach to technology operations.
  • Identify opportunities to leverage AI for incident triage, alert correlation, knowledge management, operational analytics, and runbook execution.
  • Partner with engineering teams to develop intelligent automation and self-healing capabilities.
  • Evaluate emerging AIOps, agentic AI, and automation technologies and drive adoption where appropriate.
  • Reduce manual operational effort through scripting, workflow automation, orchestration platforms, and AI-assisted tooling.
Strategic Delivery & Organizational Leadership
  • Lead and mentor high-performing technology operations and reliability engineering teams.
  • Build strong partnerships across Engineering, Infrastructure, Cybersecurity, Architecture, Data, and Business teams.
  • Translate operational challenges into actionable roadmaps and measurable initiatives.
  • Drive accountability, execution excellence, and continuous improvement across the organization.
  • Present operational trends, risks, recommendations, and performance metrics to technology leadership.
Vendor & Partner Management
  • Manage strategic relationships with technology vendors and managed service providers.
  • Ensure vendor accountability against service commitments and contractual obligations.
  • Lead vendor escalations during service-impacting events.
  • Integrate third-party support processes into internal operational workflows.
Resiliency & Disaster Recovery
  • Partner with Engineering and Infrastructure teams to ensure systems meet recovery objectives.
  • Improve operational readiness for disaster recovery and business continuity events.
  • Support development of recovery automation, replay capabilities, and resilient platform architectures.
  • Promote documentation and operational knowledge sharing to reduce dependency on tribal knowledge.
QualificationsRequired
  • 7+ years of experience in software engineering, platform engineering, cloud engineering, infrastructure engineering, DevOps, technology operations, or related technical disciplines.
  • 3-5+ years of experience leading engineering, reliability, platform, cloud, DevOps, or technology operations teams.
  • Strong experience operating and supporting applications within Microsoft Azure environments.
  • Hands-on experience with Kubernetes and container-based platforms.
  • Experience supporting distributed systems and cloud-native architectures.
  • Strong understanding of application architecture, system dependencies, and production support models.
  • Experience leading major incident response activities and outage management processes.
  • Experience implementing monitoring, observability, and alerting solutions using tools such as Datadog, Grafana, Splunk, Dynatrace, New Relic, or similar platforms.
  • Experience building automation solutions using scripting languages, APIs, orchestration tools, and workflow platforms.
  • Strong understanding of DevOps, CI/CD, release management, and software delivery practices.
  • Demonstrated ability to define, measure, and improve KPIs related to reliability, availability, operational efficiency, and service quality.
  • Excellent communication and stakeholder management skills with the ability to influence technical and business leaders.
Preferred
  • Experience building or leading Reliability Engineering, Production Engineering, DevOps, or Platform Engineering organizations.
  • Experience implementing AI-assisted operational workflows, AIOps platforms, AI agents, or intelligent automation solutions.
  • Experience with ServiceNow, Control-M, or equivalent enterprise operational platforms.
  • Experience operating within highly regulated environments.
  • Financial services experience, including asset management, lending, banking, private credit, or investment management.

The Fine Print

  • Must have unrestricted authorization to work in the United States.
  • Must be willing to comply with pre-employment screening, including but not limited to drug testing, reference verification, and background check.
  • Must be willing to work from the Chicago or New York office.
#LI-hybrid

A reasonable estimate of the current base salary range at the time of posting is below. Base salary does not include other forms of compensation or benefits. Actual base salary within the specified range is comprised of several components, including but not limited to applicant's skill, prior relevant experience, specific degrees and certifications, job responsibilities, market considerations and the location of the position.

This role is eligible for a discretionary annual bonus (based on company, business unit and individual performance).

Our benefit offerings include medical, dental and vision coverage, employer paid short & long-term disability and life insurance, 401(k), profit sharing, paid time off, Maven family & fertility benefit, parental leave (including adoption, surrogacy, and foster placement), as well as other voluntary benefits.

Base Salary Range

$175,000 - $225,000

To learn more, visit www.antares.com. Antares is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other characteristics protected by law.

Similar Jobs

11 Minutes Ago
Easy Apply
Remote or Hybrid
14 Locations
Easy Apply
75K-95K Annually
Mid level
75K-95K Annually
Mid level
Automotive • Big Data • Insurance • Software • Transportation
Develop and maintain enterprise cybersecurity policies and standards aligned with ISO 27001, PCI-DSS, and SOC 2. Conduct vendor security and privacy risk assessments, track remediation, support audits and client questionnaires, evaluate security controls and evidence, and recommend corrective actions. Partner with IT, Engineering, vendors, and business units on security architecture, data security implementations, compliance, risk management, and organization-wide security awareness.
Top Skills: Cloud SecurityData EncryptionDevsecopsIso 27001LoggingPci-DssSecurity ArchitectureSoc 2
15 Minutes Ago
Hybrid
60K-80K Annually
Mid level
60K-80K Annually
Mid level
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Hybrid role responsible for testing and certifying fire alarm, access control, and smoke detection systems to UL standards. Manages certification projects, coordinates testing and lab technicians, communicates technical issues with clients, prepares reports, develops test programs and procedures, and supports continuous improvement and special test methods.
Top Skills: Access Control SystemsFire Alarm Control PanelsFire Alarm SystemsMS OfficeSmoke Detection SystemsUl Standards
26 Minutes Ago
Hybrid
176K-235K Annually
Expert/Leader
176K-235K Annually
Expert/Leader
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Lead and optimize a regional network of laboratories across the Americas, driving cost, productivity, quality, safety, and customer satisfaction. Oversee capacity planning, budget and footprint decisions, implement automation (Labware, Magnesis) and Lean Sigma improvements, ensure compliance and accreditations, coach managers, and align lab goals with customer operating units to deliver timely, high-quality lab services.
Top Skills: LabwareLean SigmaMagnesis

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account