InstaLILY is an AI products and infrastructure company that puts execution at the frontier of enterprise AI. That work begins with Lily™, the world's first AI Forward Deployed Engineer, which learns how a business works, builds the software it needs, and goes live in days. It does not leave when the work ships; it stays and keeps the software working as the business changes. Lily runs wherever the work happens, in the cloud, on-premise, or at the edge, through InstaLILY's Small Data Center, built with NVIDIA technology. Founded in 2023 by Amit Shah and Sumantro Das, InstaLILY has raised nearly $100 million from Energize Capital, Insight Partners, and Home Depot Ventures. Headquartered in New York, with offices in San Francisco, London, and Toronto, InstaLILY serves leading companies across construction, industrial distribution, logistics, healthcare, and other operationally intensive industries. Learn more at https://instalily.ai/.
The TractionRevenue grew 5x over the past year, and Lily has driven over $200M in new annual sales for a single customer. We serve some of the largest operators in our industries, including SRS Distribution (part of The Home Depot family), United Rentals, and Henry Schein, and we work closely with the Google DeepMind and NVIDIA ecosystems.
How We WorkWe work in small teams with real ownership: clear problems, direct access to the customers whose work you're changing, and room to ship. Your code runs in live production systems inside billion-dollar operations, so you see your impact directly. People who do well here want that proximity to the work. We're growing fast, and the people who join now shape what this company becomes. Everything runs on three principles: Customers, Culture, and Code.
Instalily, a cutting-edge AI startup, is seeking a curious and highly skilled Site Reliability Engineer to help build the Internal Developer Platform (IDP) that powers our AI agent platform. We are redefining how organizations leverage AI using vertical agents, and we’re looking for engineers who think of developer experience as a product. You will build the paved roads, golden paths, and self-service tooling that allow every team at Instalily to ship AI products with speed and confidence.
As a Mid-Level Site Reliability Engineer, you will help build and own meaningful pieces of our IDP and the multi-cloud, Kubernetes-based infrastructure beneath it. You’ll partner closely with AI and Software Engineers to turn rough edges into self-service abstractions. You will benefit from world-class mentorship from highly-regarded executive leaders at Internet Retailer 100 brands.
Responsibilities- Build out Instalily’s Internal Developer Platform — including the developer portal, golden paths, and one-click developer workflows.
- Help design and stand up our Kubernetes platform; operate clusters and workloads as services migrate, focusing on networking, autoscaling, RBAC, and reliability.
- Design, implement, and maintain cloud infrastructure across AWS, GCP, and/or Azure to support the AI agent platform.
- Develop and maintain Infrastructure as Code (IaC) using OpenTofu, focusing on modularity and safe rollouts.
- Build and improve CI/CD pipelines and GitOps workflows (e.g., ArgoCD, Flux) for seamless deployment.
- Implement platform security best practices, including IAM, network segmentation, and policy-as-code.
- Implement and tune logging, monitoring, and alerting using tools such as Datadog, Prometheus, and OpenTelemetry.
- Optimize platform environments for cost, performance, and reliability; participate in on-call rotation.
- Treat developers as customers — gather feedback and measure platform adoption to iterate on the developer experience.
- Mentor more junior engineers and contribute to architectural discussions and technical reviews.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3 to 5 years of experience as a platform, cloud, infrastructure, or DevOps engineer.
- Production experience operating Kubernetes (managing clusters, upgrades, and reliability), with greenfield build-out experience being a strong plus.
- Strong hands-on experience with at least one major cloud platform (AWS, GCP, or Azure); multi-cloud experience is preferred.
- Experience contributing to or building Internal Developer Platforms (golden paths, paved roads) is a strong plus.
- Proficiency with Infrastructure as Code (OpenTofu or Terraform) and GitOps workflows (e.g., ArgoCD, Flux).
- Experience with CI/CD tools and practices such as GitHub Actions, Jenkins, or GitLab CI.
- Solid understanding of cloud networking concepts (VPCs, load balancers, DNS, CDNs, service meshes).
- Working knowledge of cloud security principles, IAM, and compliance frameworks (SOC 2, HIPAA, ISO 27001).
- A product-minded approach to internal tooling, focusing on adoption and feedback loops rather than just uptime.
- Strong problem-solving skills and ability to work in fast-paced, collaborative environments.
- Excellent communication skills to engage effectively with technical and non-technical teams.
- Interest in AI and machine learning infrastructure (GPU workloads, model serving, vector databases) is a plus.
- Proven product: Customers are live; this isn't a bet on an unproven thesis
- AI-native: In how we build, how we work, and what we sell
- Stage: Early enough to shape how the company scales
- Global: Based in New York with offices in SF and London
- Culture: Sharp, low-ego team that keeps raising the bar
- Growth: The learning curve is steep
- Salary Range: $150,000–$190,000 per year, commensurate with experience
- Equity: Stock options awards, and refreshers for top performers
- Benefits: Medical, Dental, Vision, 401K, in-office Lunch reimbursement, Wellbeing Stipend, Generous Parental leave, PTO and 10 US Federal Holidays, and more!
To ensure a focused, high-quality hiring experience, we kindly ask candidates to limit their applications to 3 open requisitions at any given time. Applying strategically to roles that best align with your skills and career goals gives you the highest chance of standing out. Have you interviewed with us in the past 12 months? We encourage you to reach out directly to your previous interviewer rather than submitting a new application.
InstaLILY is committed to providing an inclusive and barrier-free recruitment process. If you require an accommodation, please let us know, and we will work with you to meet your needs.
InstaLILY New York, New York, USA Office
New York, NY, United States
Similar Jobs
What you need to know about the NYC Tech Scene
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory



