Nest Veterinary Logo

Nest Veterinary

Senior Engineer, Production Support

Posted 9 Days Ago
Remote
Hiring Remotely in USA
Entry level
Remote
Hiring Remotely in USA
Entry level
Lead production support for a veterinary SaaS platform by triaging incidents, diagnosing issues across Flutter, FastAPI, gRPC, and legacy integrations, and driving permanent fixes. Build monitoring, alerting, diagnostic automation, ticket taxonomy, reporting, runbooks, and SLAs. Partner with Hospital Success and Engineering to resolve customer-impacting issues, identify recurring failure patterns, and communicate technical problems clearly to clinic staff and internal stakeholders.
The summary above was generated by AI
Senior Engineer, Production Support

Engineering · Remote (US or Canada) · Full-time · Reports to CTO

About Nest

Nest builds care plan infrastructure for veterinary hospitals. Corporate veterinary groups use Nest to run preventive care memberships under their own brand, and we handle what sits underneath: enrollment, recurring billing, analytics, marketing, and seller-of-record compliance.

Care plans change how pets receive care. Pet owners pay a predictable monthly amount, visit more often, and catch problems earlier. Hospitals build steadier revenue and longer client relationships. Our mission is to make that kind of preventive care available to every pet.

About the team

Nest Engineering builds the software hospitals use to run care plans, including a Flutter desktop application used in clinics, FastAPI and gRPC services in Python and Rust on Google Cloud, and integrations with practice management systems. This role sits between engineering and Hospital Success, the team that supports hospitals day to day.

What you'll do

When something breaks in production, a front desk team may be unable to enroll a client or a pet owner's payment may fail. You will be the first technical responder for those issues, and you will build the system that catches them before a hospital has to report them: monitoring, a clear ticket taxonomy, runbooks, and response targets agreed with Hospital Success. After a year, recurring issues will be fixed at the source, the team will know from data where the platform is weakest, and hospitals will see faster and more predictable resolution.

Responsibilities
  • Serve as the primary responder for customer-impacting issues across the clinic app, backend services, and legacy integrations.

  • Use AI-assisted tools to trace a reported symptom through the API layer to its source.

  • Build and maintain monitoring and alerting using tools such as Datadog, Sentry, and GCP.

  • Write Python scripts and small tools that automate detection and repeated diagnostic work.

  • Define the ticket taxonomy and report on volume, resolution time, and recurring categories.

  • Set response and resolution SLAs with Hospital Success and maintain runbooks for common issues.

  • Bring recurring failure patterns to engineering with the data needed to prioritize a permanent fix.

  • Explain technical issues clearly to clinic staff and internal teams.

Who you are

We are looking for someone who meets the minimum requirements for this role. If you meet them, we encourage you to apply. Preferred qualifications are a bonus, not a requirement.

Minimum requirements
  • You have triaged and resolved production incidents in a SaaS environment.

  • You read backend logs and API errors comfortably and can trace an issue through a client application.

  • You know your way around the application data models and are proficient in SQL.

  • You write Python for automation and reporting.

  • You have built or configured monitoring and dashboards.

  • You explain technical problems clearly to non-technical customers.

  • You are available for escalation coverage during US business hours.

Preferred qualifications
  • Experience with GCP, Datadog, or Sentry.

  • Experience on an on-call rotation.

  • Experience setting up ticket tagging and reporting in Zendesk, Linear, or similar.

  • Experience in healthcare or another regulated industry.

How we work

Nest is a small, fully remote team. Each person owns outcomes the company depends on, and the scope of most roles grows as the company does. We write decisions down, give direct feedback, and expect problems to be raised early. The work is demanding because hospitals and pet owners rely on it every day, and it suits people who want more ownership than a larger company would offer.

Our hiring process

Our process typically includes a conversation with the hiring manager, a working session based on the actual job, and conversations with the people you would work with most. We will tell you what each stage covers before it starts, and we close the loop with everyone we interview.

Equal opportunity

Nest is an equal opportunity employer. We make hiring decisions based on the work and the person, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. If you need accommodation during the hiring process, email [email protected].

Similar Jobs

15 Days Ago
Remote
USA
Senior level
Senior level
Fintech • Financial Services
Owns Bloom Credit’s production support and incident response function, including monitoring, incident coordination, root cause analysis, documentation, metrics, and preventive improvements. Supports client-facing Operations with onboarding, offboarding, and troubleshooting tasks while automating repetitive workflows using scripts. Partners with engineering, operations, and product teams to improve system reliability, reduce support overhead, and strengthen incident response processes.
Top Skills: Ai Agentic ToolsBashDatadogPagerdutyPython
2 Days Ago
Remote
TX, USA
Senior level
Senior level
Security
Provides day-to-day production support for complex cloud application systems. Responsibilities include outage detection and resolution, system administration, infrastructure maintenance, user support, monitoring, automation, database support, asset inventory, security compliance, documentation, and continuous improvement. The role collaborates with development, operations, security, customers, and management, supports AWS and Azure environments, and participates in on-call production support.
Top Skills: Amazon CloudwatchAnsibleAWSAws BatchAws CodebuildAws CodedeployAws RdsAws Step FunctionsAzureBashChefCi/CdCloudFormationDynatraceGitGithub ActionsLinuxOracle RdbmsPowershellPuppetPythonS3SaltstackSciencelogicSQLSQL ServerUnixVpnWindows
2 Days Ago
Remote
TX, USA
Senior level
Senior level
Security
Provide day-to-day production support for complex cloud systems and scientific workstations. Diagnose and resolve outages, maintain infrastructure, manage system assets, support users, monitor performance, implement automation and continuous improvement, and ensure compliance with VA security standards. Responsibilities include system administration, cloud and database operations, documentation, training, troubleshooting, and collaboration with development, operations, and security teams. Participation in on-call production support is required.
Top Skills: Amazon CloudwatchAmazon S3AnsibleAWSAws BatchAws CloudformationAws CodebuildAws CodedeployAws RdsAws Step FunctionsAzureBashChefCi/CdDynatraceGitGithub ActionsLinux/UnixOracle RdbmsPowershellPuppetPythonSaltstackSciencelogicSQLSQL ServerWindows

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account