Nexxa.ai Logo

Nexxa.ai

Staff DevOps Engineer

Posted 10 Days Ago
In-Office or Remote
Hiring Remotely in Toronto, ON
Senior level
In-Office or Remote
Hiring Remotely in Toronto, ON
Senior level
Own and evolve cloud, on-premises, and edge infrastructure supporting AI, data, and product workloads. Design Kubernetes platforms, GPU scheduling, CI/CD pipelines, infrastructure-as-code, observability, security, and reliability practices. Lead infrastructure projects, incident response, platform strategy, and developer enablement while balancing cost, latency, reliability, and velocity. Mentor engineers and provide technical direction for scalable, production-grade ML and industrial systems.
The summary above was generated by AI

Nexxa is building the best AI systems for heavy industries — enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

About the Role

We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.

This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.

What You'll Do
  • Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end

  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams

  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments

  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling

  • Partner with data and AI teams to support the infrastructure behind:

    • Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks)

    • Feature stores, embedding indices, and retrieval pipelines

    • Model training, evaluation, and serving infrastructure

  • Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems

  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations

  • Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations

  • Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity

  • Collaborate with engineering leadership to define infrastructure roadmap and platform strategy

  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org

Required Qualifications
  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles

  • Deep hands-on experience with:

    • Cloud platforms (AWS, GCP, or Azure) at production scale

    • Kubernetes in production, including GPU workload scheduling

    • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)

    • CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)

  • Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)

  • Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus

  • Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling

  • Proven ability to independently scope and lead infrastructure projects from design through production rollout

  • Strong incident management instincts — you can lead through an outage calmly and drive toward root cause

Preferred Qualifications
  • Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts

  • Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)

  • Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)

  • History of building internal developer platforms or self-service infrastructure tooling

  • Experience scaling infrastructure teams or setting technical direction at a Staff level

What Success Looks Like
  • You can own ambiguous, high-stakes infrastructure problems end-to-end

  • Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today

  • You bring strong technical judgment on tradeoffs between reliability, cost, and speed

  • You raise the bar for operational rigor and engineering discipline across the team

  • You help define what's next for the platform, not just execute what's known

Why Join Nexxa.ai?
  • Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies

  • Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement

  • Professional Growth: Benefit from significant opportunities for career development and advancement

  • Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions

If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.

Similar Jobs

14 Days Ago
Remote
United States
Expert/Leader
Expert/Leader
Productivity • Software
Lead design and operation of scalable AWS/GCP cloud infrastructure. Build and operate Kubernetes (EKS), Docker, Terraform/Terragrunt, Helm and GitOps workflows. Improve CI/CD with GitHub Actions and ArgoCD, implement monitoring and observability, strengthen security, disaster recovery, and developer self-service automation to improve platform reliability and productivity.
Top Skills: ArgocdAWSBetterstackCloudwatchDockerGCPGithub ActionsGitopsGoGrafanaHelmKubernetes (Eks)LinuxLokiNode.jsPostgresPrometheusRedisRuby On RailsRustServerlessService MeshTerraformTerragrunt
One Month Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Software • Automation
Build, maintain, and operate Mimica’s multi-region Kubernetes platform (GKE/EKS), infrastructure-as-code, GitOps pipelines, and observability. Support engineering teams, triage incidents, migrate IaC to Crossplane, drive FinOps, and design BYOC/single-tenant access patterns while automating work with Python tooling.
Top Skills: ArgocdAWSCrossplaneDatadogEksGCPGitopsGkeGrafanaGrafana AlloyHoneycombIamKubernetesLokiMongoDBNew RelicPostgresPythonRbacSQL ServerTemporalTerraformTerragrunt
23 Days Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Software • PropTech • Generative AI
Lead architecture and rewrites for a scalable, resilient AI-first platform. Own backend systems, data storage performance, deployment pipelines, cloud infrastructure via IaC, observability, incident response, and analytics. Mentor engineers and enable rapid customer onboarding and AI-driven automation.
Top Skills: Ai InfrastructureAWSCi/CdGCPInfrastructure As CodeLlm-Based SystemsNode.jsPulumiReactTerraformTypescript

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account