Retool Logo

Retool

Site Reliability Engineer (SRE)

Reposted One Month Ago
Hybrid
New York, NY, USA
164K-306K Annually
Senior level
Hybrid
New York, NY, USA
164K-306K Annually
Senior level
Operate and improve reliability across Retool Cloud, managed, BYOC, and self-hosted deployments. Automate provisioning, upgrades, migrations, and secret rotations. Build observability and safer deployment/rollback workflows, partner with product teams, and produce runbooks, docs, and migration guides to reduce customer toil and scale operations.
The summary above was generated by AI
ABOUT RETOOL

Nearly every company in the world runs on custom software for critical operations like tracking performance metrics, handling support workflows, building admin dashboards, and countless processes you might never have thought of. But most companies don't have the resources to properly invest in these tools, leading to a lot of old, clunky internal software, or worse, teams still stuck in manual and spreadsheet workflows.

AI has changed who gets to build software. The definition of "developer" now includes analysts, operators, and domain experts creating solutions directly—and the tools they reach for are multiplying by the week. That's both an opportunity and a challenge: as more people build with more AI tools, the risk of shipping ungoverned software into production grows just as fast.

At Retool, we're building the platform that makes all of it safe to ship. Build with any AI tool you want, then deploy into one place that connects to your real business data, enforces enterprise policies automatically, and lets teams create once and reuse everywhere with shared, trusted components. The cost of building software has collapsed. The cost of governing it hasn't—and that's the problem we solve.

Developers and domain experts have already automated over 100 million hours of work on our platform, freeing them to focus on creative problem-solving and strategic work that drives real business value. The people closest to the problem can now build the software to solve it, safely, and within enterprise guardrails.

Let's build the future together.
WHY WE'RE LOOKING FOR YOU:
Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system.

Retool's Core Infrastructure team owns the systems that make this possible: Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud) environments, Kubernetes and Helm deployments, Docker Compose, and the migration paths between them. It is a broad surface area, and it is one of the biggest levers we have for making Retool work for enterprise customers.

The work is not clean-room infrastructure. Customers run different clouds, different versions, different deployment models, and different levels of operational maturity. A bad upgrade experience can leave a customer many versions behind. A manual Terraform run can become the bottleneck during a launch or incident.

We are hiring SREs who want to turn that mess into leverage. You will help us reduce customer toil, automate upgrades and infrastructure changes, build reliability tooling across Retool Cloud and customer-owned environments, and make Retool easier to deploy and operate at enterprise scale. The strongest candidates are comfortable debugging Kubernetes, Terraform, AWS, Postgres, networking, and deployment problems, then stepping back and building the automation or product surface that prevents the same problem from happening again.

What you'll do:
  • Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • Build the automation that turns today's manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • Improve observability for Retool Cloud, self-hosted customers, and internal operators. We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
  • Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current
  • Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
  • Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies
  • Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
  • Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.

What we're looking for:
Infrastructure fundamentals
  • Deep experience operating production infrastructure in AWS.
  • Experience improving reliability for customer-facing SaaS systems.
  • Strong Kubernetes fundamentals.
  • Real Terraform or infrastructure-as-code experience.
  • Good operational judgment around databases, especially Postgres.

Reliability and automation
  • Experience building or operating observability systems.
  • Programming ability in a language such as Go, Python, TypeScript, Java, or Ruby.
  • A bias toward automation. If you find yourself doing the same operational task twice, you should start thinking about the interface, workflow, or tool that eliminates the third time.

Customer and team judgment
  • Clear written communication.
  • Comfort working directly with customer-facing teams and, when useful, customers themselves.

What makes SREs successful here:
You will do well here if you like infrastructure that sits close to real customer pain. Some days that means debugging a specific customer environment. Other days it means improving Retool Cloud reliability or designing the migration path so the next 25 customers do not need that same debugging session.

We value SREs who are ambitious, curious, energetic, and careful with the details. Retool moves quickly, priorities can change, and the systems are not always as clean as we want them to be. The work needs SREs who can get their hands dirty, tell the truth about tradeoffs, and leave the system better than they found it.
For candidates based in the United States, the pay range(s) for this role is listed below and represents base salary range for non-commissionable roles or on-target earnings (OTE) for commissionable roles. This salary range may be inclusive of several career levels at Retool and will be narrowed during the interview process based on a number of factors such as (but not limited to), scope and responsibilities, the candidate’s experience and qualifications, and location. 
Additional compensation in the form(s) of equity and/or commission are dependent on the position offered. Retool provides a comprehensive benefit plan, including medical, dental, vision, and 401(k). Pay and benefits are subject to change at any time, consistent with the terms of any applicable compensation or benefit plans.

The base pay range for this role is $163,710 – $306,000 per year.
Retool offers generous benefits to all employees and hybrid work location. For more information, please visit the benefits and perks section of our careers page!

Retool is currently set up to employ all roles in the US and specific roles in the UK. To find roles that can be employed in the UK, please refer to our careers page and review the indicated locations.

Retool New York, New York, USA Office

Retool NYC is in the heart of the Flatiron District, with tons of shopping and dining nearby. The 23rd Street station is just a short walk away. We have limited bike storage and 24/7 security.

Retool New York, New York, USA Office

New York, United States

Similar Jobs

Yesterday
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
Yesterday
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, resilience, security, and performance optimization of enterprise mainframe environments. Responsibilities include z/OS performance tuning, WLM and RACF administration, business continuity planning, automation, technical governance, incident resolution, stakeholder collaboration, and guidance of cross-functional engineering and operations teams. The role also evaluates cloud, DevOps, AI, and hybrid IT technologies for mainframe transformation.
Top Skills: AnsibleCsmGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRed Hat Ansible For Ibm Z CollectionsRmfSmfWlmZlinux
2 Days Ago
Hybrid
2 Locations
151K-187K Annually
Senior level
151K-187K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Design and implement reliable, scalable IT infrastructure; automate processes; monitor systems; resolve incidents; conduct performance and load testing; optimize cloud environments; maintain storage architecture; and lead continuous improvement. The role requires troubleshooting complex system issues, developing automation solutions, managing incidents, collaborating across teams, mentoring others, and supporting secure business operations across AWS, Google Cloud, and Microsoft Azure.
Top Skills: AWSGoogle Cloud PlatformAzure

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account