Flowcode Logo

Flowcode

Senior SRE Engineer

Posted Yesterday
Be an Early Applicant
Hybrid
New York, NY, USA
Senior level
Hybrid
New York, NY, USA
Senior level
Senior Site Reliability Engineer responsible for improving platform availability, scalability, resilience, and observability. The role manages AWS and EKS infrastructure with Terraform, builds CI/CD pipelines using GitHub Actions, expands GitOps through ArgoCD and Helm, strengthens disaster recovery, supports incident response, and develops monitoring, alerting, dashboards, and SLOs. The engineer partners with product teams, leads infrastructure initiatives, troubleshoots Kubernetes, and supports high-availability distributed systems.
The summary above was generated by AI
Senior SRE Reliability Engineer

Location: New York, NY (Hybrid) / Remote
Department: Engineering

The Role

Flowcode is seeking a Senior Site Reliability Engineer (SRE) to work on reliability and infrastructure efforts across our platforms. This role will help grow and drive our infrastructure strategy, operational rigor and observability while building and supporting the systems and tooling required to support Flowcode’s continued growth.

As an individual contributor within our engineering organization, you will develop and operate scalable cloud infrastructure, establish best practices around deployment and reliability, and partner closely with engineering teams to ensure systems are scalable, resilient and observable. 

What You’ll DoReliability & Infrastructure
  • Improve system availability, scalability, and resilience across Flowcode's platforms
  • Own key pieces of our EKS-based infrastructure end-to-end
  • Contribute to incident response and postmortems, turning findings into durable fixes
  • Support engineering teams with infrastructure questions, escalations, and day-to-day unblocking
Cloud & Platform Engineering
  • Manage and scale our core AWS footprint (EKS, VPC, RDS) through Infrastructure as Code (Terraform)
  • Enhance disaster recovery and failover mechanisms to protect mission-critical workloads
  • Collaborate with product engineering to streamline and optimize internal developer experience
CI/CD & Deployment Automation
  • Design and scale deployment pipelines using GitHub Actions
  • Expand GitOps practices and tooling through ArgoCD
  • Facilitate secure delivery with automated validation and progressive rollout strategies
Observability & Monitoring
  • Oversee and optimize the organization's monitoring, logging, and alerting infrastructure
  • Develop high-signal metrics, tracing, and visualization dashboards while minimizing operational noise
  • Establish and monitor Service Level Objectives for managed platform components
QualificationsRequired
  • 4+ years of professional experience across SRE, DevOps, or Platform Engineering domains
  • Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs
  • Advanced Terraform or OpenTofu expertise, encompassing module architecture and production state management
  • Hands-on operational experience with GitOps workflows via ArgoCD and Helm-based deployments
  • Ability to author production-grade code in Go or Python alongside robust shell scripting
  • Mastery of core AWS services, specifically EKS, Networking/VPC, RDS, and IAM
  • Experience maintaining and scaling CI/CD automation using GitHub Actions within collaborative environments
  • Proven track record of leading infrastructure initiatives from initial design through to long-term operation
  • Background in supporting large-scale distributed systems within high-availability production environments
  • Adept at navigating interrupt-driven workflows, balancing strategic project delivery with day-to-day operational support
Preferred
  • Exposure to Crossplane or alternative Kubernetes-native solutions for infrastructure provisioning
  • Deep observability experience utilizing Datadog or Prometheus to engineer SLOs, high-signal dashboards, and intelligent alerting
  • Practical knowledge of modern secrets management frameworks and implementation
  • Experience optimizing cluster efficiency through autoscaling technologies such as Karpenter or Cluster Autoscaler

Flowcode is not for everyone. We hire with a pinhole lens — only those with the rare combination of intellectual horsepower, execution velocity, and uncompromising drive will thrive here. If you are seeking to operate at the highest levels of performance and impact, we want to meet you.

How to Apply

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

A successful candidate’s starting pay will be determined based on the role, job-related skills, experience, qualifications, work location, and market conditions. 

HQ

Flowcode New York, New York, USA Office

We're based in the heart of Soho, where there is no shortage of creativity. From galleries to boutiques, trendy bars to some of the best restaurants in the city, our office fits right in. We are located in an industrial garage, converted from a metalworks factory into a collaborative work space.

Similar Jobs

2 Hours Ago
Easy Apply
Hybrid
New York City, NY, USA
Easy Apply
129K-258K Annually
Senior level
129K-258K Annually
Senior level
Marketing Tech • Mobile • Software
Lead reliability engineering for Braze’s Ruby on Rails monolith and Go API services. Operate high-performance NGINX and Kubernetes ingress infrastructure, develop automated scaling routines, define SLIs and SLOs, conduct systems design and capacity planning, and participate in PagerDuty on-call rotations. Lead incident response, root-cause analysis, blameless retrospectives, documentation, and permanent reliability improvements while partnering with product engineering teams on resilient, highly available architectures.
Top Skills: AnsibleAWSAzureChefDatadogGCPGoGrafanaHorizontal Pod Autoscaler (Hpa)JavaKafkaKubernetesLinuxMongoDBNginxPagerdutyPostgresPrometheusPythonRedisRubyRuby On RailsTcp/IpTerraformUnix
12 Days Ago
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
5 Days Ago
Hybrid
New York, NY, USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills: AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account