AZX Logo

AZX

Senior Software Engineer (AI Inference & Runtime Platform)

Posted 18 Days Ago
In-Office or Remote
Hiring Remotely in Seattle, WA
140K-230K Annually
Senior level
In-Office or Remote
Hiring Remotely in Seattle, WA
140K-230K Annually
Senior level
Own the AI inference and runtime platform, including open-weight model serving, Kubernetes operators and autoscaling, GPU infrastructure, stateful vector and graph stores, and secure microVM-based agent sandboxing. Build Rust, Python, and Kubernetes control-plane services; manage infrastructure as code, observability, metering, lifecycle operations, upgrades, backups, and failover. Establish security controls, open-source standards, and reliable production operations across cloud and managed GPU environments.
The summary above was generated by AI

About AZX

Our mission is to accelerate positive impact in critical industries through AI transformation. We specialize in physics-informed ML and enterprise AI solutions that directly address climate and sustainability challenges.

We’re growing quickly and already work with category-leaders in real estate (CBRE), energy (LevelTen Energy), logistics (Flexe) and utilities.

We bootstrapped profitably for our first year and are now backed by leading investors focused on AI, climate and energy.

We work on challenges in clean energy, decarbonization, climate risk, energy systems, and global economics. We’re building our company for long-term success and aim to create the ultimate place to work for those passionate about AI and making a positive impact.

About This Role:

You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. This role anchors on two systems. The first is our inference control plane — open-weight models and custom task-model zoos, hosted and operated across managed GPU clouds and customer-managed Kubernetes clusters, with scale-to-zero economics, cold-start discipline, and per-token cost accounting that stays correct even when a client disconnects mid-stream — along with the Kubernetes layer those workloads live on: operators, autoscaling, node lifecycle. The second is our agent-sandboxing platform: hardware-isolated microVMs for running untrusted, agent-generated code securely and compliantly by construction, where agents operate with least privilege, never see a credential, and a human gates anything that writes to a system of record. You'll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home — building the fork engine, the guest agent, and the multi-substrate model lifecycle. We're looking for individuals who've built this class of stack (an inference-serving or serverless-GPU platform), operated it hard at scale, or ideally both.

Responsibilities:

  • Manage the serving tier for open-weight models: engine deployment and configuration, cold-start strategy, per-model SLOs, and upgrade/canary discipline.

  • Administer the Kubernetes layer for inference and sandbox workloads: operators and CRDs, autoscaling (KEDA/Karpenter-class), GPU scheduling and sharing, and node lifecycle.

  • Own the stateful data plane end to end with restore procedures that are regularly tested.

  • Oversee the sandbox runtime and its host-side control plane: lifecycle, exec, snapshot/fork, teardown, metering, and the threat model of the isolation boundary.

  • Direct the FastAPI control-plane services, Terraform/OpenTofu, Bicep, and the dashboards

  • Run the layer the backend services team builds on, expect to debug into their services, and expect them to read your dashboards.

  • Manage the open-source posture: build to OSS standards and release as it matures, with reviewed PRs, real docs, and reproducible builds.

Core Qualifications

  • 5+ years of shipping production systems in a systems language. Rust is the house language, but polyglots are welcome — deep Go, C/C++, or Zig with genuine appetite for Rust counts. Async runtimes, memory-safety discipline, and debugging at the syscall boundary should be familiar territory.

  • Operated Kubernetes workloads that other people depended on — controllers or operators, scheduling, autoscaling, node lifecycle. You've been paged, and the experience changed how you build.

  • Strong ability to threat-model isolation boundaries (namespaces, cgroups, seccomp, hypervisors), including identifying what an untrusted guest could observe, forge, or exhaust, and applying security best practices for agentic execution — least privilege, no credentials in the sandbox, audit trails, and human approval on write actions.

  • Hands-on experience deploying or operating open-weight LLM serving infrastructure (vLLM/SGLang or similar), including packaging models into reliable, metered production endpoints.

  • Performance discipline in distributed systems: you measure before you optimize, and you can tell the story of a latency you killed with the numbers attached.

  • Practical depth in some of our core stack — Rust (tokio), Python/FastAPI, Kubernetes operators (controller-runtime/Kubebuilder/CRDs), KEDA, Karpenter, GPU device plugins/DRA — with a genuine willingness to research your way into the rest.

  • Familiarity with isolation technology (Firecracker, Kata, gVisor, or comparable), secrets management and egress control (Vault/KMS-class), and hosting stateful systems (vector stores like pgvector/Qdrant, graph stores like Neo4j) with backup and failover discipline.

  • Comfort operating across cloud and GPU substrates — AWS/Azure/GCP plus managed GPU clouds — using infrastructure-as-code (Terraform/OpenTofu, Bicep) and observability tooling (OpenTelemetry).

  • Bachelor's Degree; Master's is a plus

Why AZX!

  • Be part of a fast-growing, profitable, mission-driven company with industry-leading clients tackling the massive opportunity of AI transformation in critical industries.

  • Competitive early-stage startup compensation (based on capabilities, experience, and location)

  • Bonus eligibility

  • Health insurance with meaningful coverage for dependents

  • Flexible paid time off

  • Equity

  • Fully remote culture with a cluster of teammates in Seattle

 

Additional Information:

  • Must be able to travel 2x/year for company summits

  • Applicants must be currently authorized to work in the United States on a full-time basis.

  • We are unable to sponsor or take over sponsorship of employment visas at this time.

  • Please note that our interview process includes a written take-home assignment followed by a live two-hour technical session with our engineering team, so if that format isn't a good fit, we'd ask that you not apply

  • Please only apply to a maximum of 2 roles at a time, any applicants who apply to more then 2 roles within a 6 month period will automatically be disqualified

Next Steps:

If this job sounds like a great fit but you don’t check ALL of these qualification boxes, we’d still love to hear from you!

Similar Jobs

13 Minutes Ago
In-Office or Remote
2 Locations
215K-358K Annually
Senior level
215K-358K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads measurement and product marketing for eight enterprise AI platforms. Builds shared KPI, taxonomy, scorecard, data-quality, and value-realization frameworks; translates usage and outcome data into executive insights and investment guidance. Oversees positioning, internal launches, campaigns, enablement, adoption, and audience segmentation. Partners across product, engineering, data, communications, and business teams while building and managing teams responsible for analytics, marketing, and communications.
Top Skills: Ai PlatformsBusiness IntelligenceDashboardsData ContractsData FabricData VisualizationEvent TaxonomyExperimentationKnowledge GraphsKpi FrameworksOkr PlatformsProduct AnalyticsTelemetry
13 Minutes Ago
In-Office or Remote
2 Locations
163K-272K Annually
Senior level
163K-272K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads product strategy, roadmap, lifecycle ownership, adoption, and measurable outcomes for a greenfield enterprise AI platform. Defines platform boundaries, governance controls, evaluation, registration, and cost-tracking capabilities while translating evolving risk, privacy, security, and GxP requirements into usable product features. Partners with engineering, design, legal, compliance, risk, security, and global agent-building teams to prioritize investments, guide delivery, communicate direction, and drive platform adoption.
Top Skills: AgileAIAi AgentsAi GovernanceGxpLeanModel Evaluation
An Hour Ago
In-Office or Remote
New York, NY, USA
86K-118K Annually
Junior
86K-118K Annually
Junior
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Own end-to-end revenue and partnership accounting, including month- and quarter-end close, reconciliations, journal entries, subledger reviews, variance analysis, revenue recognition, incentive accounting, disclosures, controls, and audit support. Apply ASC 606 and US GAAP to complex arrangements while partnering with commercial, legal, tax, and data teams. Design AI and automation solutions using ChatGPT and Gemini to improve efficiency, quality, governance, and auditability.
Top Skills: Apple MacosAsc 606ChatgptGeminiGoogle SuiteOracleSlackUs Gaap

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account