Coterie Logo

Coterie

Sr Site Reliability Engineer

Posted 11 Hours Ago
Remote
Hiring Remotely in United States
140K-170K Annually
Senior level
Remote
Hiring Remotely in United States
140K-170K Annually
Senior level
Manage Azure infrastructure and AKS clusters, build GitHub Actions CI/CD pipelines, and improve Grafana-based observability and incident response. Define SLIs, SLOs, and error budgets; maintain infrastructure as code with Pulumi; troubleshoot reliability issues; perform capacity planning and performance tuning; participate in on-call support; and document operational procedures. Collaborate with development teams to deliver scalable, reliable production systems.
The summary above was generated by AI

Who we are:
Through a partnership-based approach, Coterie helps insurance professionals unlock untapped revenue in the small commercial space. With an innovative quoting platform that delivers accurate pricing and bindable quotes in less than one minute, Coterie makes small business insurance effortless.  
We are on a mission to build and foster a world-class team to bring speed, simplicity, and service to commercial insurance. We value integrity, humility, passion, and intelligence. If you want to push yourself and reshape a $200B+ market, we’re excited to talk to you!


What will the Site Reliability Engineer do?

We're looking for a Senior Site Reliability Engineer who's passionate about building and maintaining reliable, scalable infrastructure and who thrives on making systems better every day. In this role, you'll join our SRE team to help keep our platforms running smoothly, improve our observability and incident response capabilities, and partner with development teams to deliver infrastructure that supports high-quality, reliable software.

You'll play a key role in managing our cloud infrastructure, strengthening our CI/CD pipelines, and helping us get the most out of our monitoring and alerting tools, particularly Grafana. This is a great opportunity for a mid-level engineer ready to take ownership of meaningful infrastructure challenges.
Key Responsibilities:

  • Manage and maintain cloud infrastructure on Azure, including Azure Kubernetes Service (AKS) clusters and supporting resources
  • Build, improve, and maintain CI/CD pipelines using GitHub Actions to support reliable and repeatable deployments
  • Own and enhance our Grafana implementation; designing dashboards, configuring alerts, and supporting incident management workflows
  • Monitor system health, triage incidents, and drive root cause analysis to prevent recurrence
  • Collaborate with development teams to define and track SLIs, SLOs, and error budgets that align with business goals
  • Contribute to infrastructure-as-code practices using Pulumi
  • Identify and resolve reliability risks through capacity planning, performance tuning, and proactive system improvements
  • Participate in an on-call rotation to support production systems and respond to incidents
  • Document runbooks, operational procedures, and architectural decisions to support team knowledge sharing

What we are looking for:  

  • 5+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure role
  • 3+ years experience working with infrastructure as code
  • 2+ years of experience architecting CI/CD pipelines and cloud-based infrastructure
  • Strong hands-on experience with:
  • Azure Cloud services and resource management
  • Kubernetes and AKS administration, including deployments, networking, and troubleshooting
  • GitHub Actions for CI/CD pipeline development and maintenance
  • 3+ experience with Grafana or similar tooling, including dashboard creation, alerting configuration, and incident management
  • Hands-on experience with Prometheus, Loki, or other observability tools in the Grafana ecosystem
  • Proficiency in at least one scripting or programming language such as Python or Bash
  • Understanding of networking fundamentals, DNS, load balancing, and container orchestration concepts
  • Strong analytical and communication skills; able to diagnose complex system issues and clearly communicate findings
  • Demonstrated ability to collaborate across teams and contribute to a culture of reliability
  • Experience working in an agile environment with modern DevOps practices

What will make you stand out:

  • Experience working at a startup or in a fast-paced, cross-functional environment
  • Familiarity with the insurance industry or other regulated sectors
  • Familiarity with service mesh technologies (e.g., Istio)


Our interview process:

Our hiring process generally consists of 4 phases. The goal is to provide an opportunity for us to learn more about our candidates while allowing them to get to know us as well!

  • Phase 1: Qualified candidates will first meet with a member of our People Operations team for a phone interview.  This discussion is a high-level conversation to understand more about your background and interests and for us to share more about Coterie and the position.
  • Phase 2: Selected candidates will be invited to meet with our Hiring Manager for a 2nd interview via Teams video. This interview is designed to be more detail oriented and allows you to learn more about the role and expected to be 30 minutes in length.
  • Phase 3: Top candidates will be invited to participate in an experiential assessment phase, which includes a take-home coding exercise. Candidates who successfully complete the project will be invited to a 1-hour technical interview with our hiring manager and members of the engineering team.
  • Phase 4: Final candidates will receive an invite to our final interview series. This series will include 1:1 interview with senior leadership. The final series is roughly 1 hour total.


What's in it for you:

Coterie has excellent benefits for all full-time employees. We offer the following:

  • 100% remote
  • Health insurance through Aetna (we pay 100% of premiums)
  • Dental and vision insurance through Guardian (we pay 100% of premiums)
  • Basic life insurance (we pay 100% of premiums)
  • Access to flexible spending account (FSA) or health savings account (HSA) (for those using HSA eligible plans)
  • 401K plan (up 4% match with immediate vest). Must be 21 years of age or older to participate
  • Flexible PTO policy offering employees up to 4 weeks of PTO in their first 12 months. Thereafter, PTO usage aligns with company standards and typically does not exceed 5 weeks per calendar year.
  • 12 company-paid holidays each year
  • Continuing education annual stipend
  • Annual salary estimated between $140,000-$170,000 based on national data. Candidates who meet all the minimum requirements and possess additional relevant experience, as outlined in the job description, may be considered for a salary above the midpoint of the above range. Salary is based on internal equity; internal salary ranges; market data/ranges; applicant’s skills; prior relevant experience; degrees or certifications, etc. 

Work Authorization:
At this time, Coterie Insurance is unable to consider candidates who require current or future visa sponsorship. Applicants must have authorization to work in the United States without the need for sponsorship now or in the future. Falsification of an application, including work authorization status, is immediate grounds for dismissal from consideration.

Similar Jobs

Yesterday
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
2 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
3 Days Ago
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills: AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account