crewAI Logo

crewAI

Software Engineer, Infrastructure & Reliability

Posted 22 Days Ago
Remote
Hiring Remotely in United States
Entry level
Remote
Hiring Remotely in United States
Entry level
Build and operate CrewAI’s cloud and enterprise infrastructure across AWS, Azure, and GCP. Develop CI/CD pipelines, deployment automation, observability, security controls, self-hosted installation tooling, and reliability processes. Support containers, Kubernetes, databases, queues, networking, secrets, and production workloads. Participate in incident response and on-call operations while partnering with runtime and product engineering teams to improve deployment safety, scalability, and customer production environments.
The summary above was generated by AI
About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.

The Role

You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.

This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.

What You'll Do
  • Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
  • Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
  • Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.
  • Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
  • Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
  • Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
  • Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.
  • Reduce operational toil by automating recurring workflows and making deployments boring.

RequirementsWhat We're Looking For
  • Strong infrastructure/platform engineering experience in production SaaS environments.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers.
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
  • Experience supporting enterprise/self-hosted deployments.
  • Terraform or other IaC experience.
  • SRE background: SLOs, incident review, capacity planning, load testing.
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.

Similar Jobs

2 Months Ago
In-Office or Remote
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Lead design, development, and operation of cloud infrastructure and SRE-focused systems. Own medium-to-large infrastructure projects, build resilient platforms, and drive cross-team technical delivery. Mentor engineers, define SLOs, reduce failure domains, and build tooling for automated, secure CI/CD and production reliability.
Top Skills: ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
2 Months Ago
In-Office or Remote
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Design, build, and operate resilient cloud infrastructure and platform services. Lead cross-team projects, mentor engineers, define SLOs, automate CI/CD and IaC, improve reliability, and own medium-to-large infrastructure features.
Top Skills: ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
13 Minutes Ago
Remote or Hybrid
USA
20-23 Hourly
Entry level
20-23 Hourly
Entry level
Fintech • Healthtech • HR Tech • Information Technology • Financial Services • Telehealth
Provides empathetic phone and chat support to families navigating loss, disability leave, and other difficult life events. Responsibilities include creating care plans, researching resources and service providers, answering logistical questions, documenting interactions, meeting service levels, escalating sensitive issues, maintaining knowledge resources, and sharing user insights. The role requires independent work, strong communication, organization, problem-solving, adaptability, and comfort with technology and sensitive data.
Top Skills: Google SuiteSlackZendesk

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account