General Compute Logo

General Compute

Founding Inference Engineer

Posted Yesterday
Be an Early Applicant
Hybrid
New York, NY, USA
Senior level
Hybrid
New York, NY, USA
Senior level
Own the end-to-end LLM inference serving stack for specialized ASIC hardware. Design request routing, batching, scheduling, KV-cache management, autoscaling, monitoring, alerting, and failover. Optimize throughput and cost per token, support production reliability and on-call operations, collaborate with compiler and hardware bring-up teams, shape the serving roadmap, and establish technical standards for a growing engineering team.
The summary above was generated by AI
About us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.

About the Role

Getting a model correct and fast on our silicon is only half the problem — the other half is serving it. You'll build and own the inference layer that sits between a bought-up model and a live customer request: request scheduling, batching, KV-cache management, autoscaling across our ASIC fleet, and the failure modes that only show up at real traffic and real scale.

This is a founding role on a small team, which means the scope is wide and the ownership is real: there's no separate SRE org to hand reliability to and no platform team to hand infra to. You'll design the serving architecture, then be the person paged when it breaks. The bet is that a serving stack built specifically for our hardware — not adapted from a GPU-first framework — is a durable edge, and you're the person who proves that out in production.

What You'll Do:
  • Own the inference serving stack end-to-end. Design and build the system that takes a bring-up-verified model and serves it in production: request routing, batching, scheduling, and autoscaling a single model's serving replicas.

  • Push cost-per-token down. Continuously tune batching strategy, KV-cache handling, and hardware utilization to widen the throughput advantage over GPU-based serving.

  • Build for reliability from day one. Put in place the monitoring, alerting, and failover that make a fast-moving inference stack trustworthy under real customer load — and be the one who responds when it isn't.

  • Work at the boundary with the compiler and bring-up team. Define the interface between "a model is correct and compiled" and "a model is live and fast," and push issues back to the right side of that line.

  • Shape the roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in serving — multi-tenant isolation, speculative decoding, new scheduling strategies — not just execute a spec someone else wrote.

  • Set the technical bar for the team you're helping build. Early architecture and code-quality decisions you make here will shape how the serving team operates as it grows.

What We Need From You:
  • 5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.

  • Direct experience with LLM inference serving — request batching, KV-cache management, continuous batching, or similar — in a production environment, not just research code.

  • Comfortable owning reliability: you've been on call for a system that mattered, and you design for failure rather than reacting to it after the fact.

  • Strong systems fundamentals — concurrency, networking, scheduling — deep enough to reason about performance at the hardware level, not just the application level.

  • Self-directed and comfortable with ambiguity. This is a founding role: there's no existing playbook to follow, and you'll help write it.

Nice-to-Haves:
  • Experience serving models on non-NVIDIA accelerators (TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar).

  • Familiarity with serving frameworks such as vLLM, TGI, TensorRT-LLM, or SGLang, and an opinion on where they fall short.

  • Experience running infrastructure at a small company or in a founding/early-engineer capacity before.

  • Exposure to capacity planning or fleet management for specialized hardware.

Similar Jobs

9 Minutes Ago
In-Office or Remote
United States
Senior level
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Manage the global travel and expense program, including Concur administration, expense report processing, corporate-card reconciliation, audits, policy compliance, employee support, reporting, and process improvements. Partner with Accounting, Accounts Payable, FP&A, HR, and IT to maintain accurate financial data, effective controls, system integrations, and timely reimbursements.
Top Skills: Concur ExpenseConcur TravelCorporate Card PlatformsErpHrisExcelMS OfficeNetSuiteOracleSAP
13 Minutes Ago
Hybrid
121K-205K Annually
Senior level
121K-205K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead program quality planning and execution across design, suppliers, and manufacturing. Manage lifecycle reviews, customer quality liaison activities, audits, requirements flowdown, nonconformance investigations, corrective actions, quality risks, supplier plans, and performance metrics. Ensure compliance with aerospace and quality standards while coordinating cross-functional resources, supporting production transitions, and driving improvements in program and customer outcomes.
Top Skills: ApqpAs13100As9100As9102As9145CmmiDfmeaFaiIatf 16949Iso 9001PfmeaPpapRccaSpc
14 Minutes Ago
Hybrid
43K-69K Annually
Entry level
43K-69K Annually
Entry level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Performs automated and manual testing of electronic circuits, controls, avionics systems, satellite communications equipment, aircraft control systems, and mission computers. Uses electronic measurement instruments, reads schematics and technical drawings, follows written procedures, and troubleshoots test failures. The role requires strong attention to detail, problem-solving ability, physical stamina, and willingness to learn.
Top Skills: AmmetersAutomated TestingElectrical Schematic DiagramsLogic AnalyzersManual TestingOscilloscopesTechnical DrawingsVoltmeters

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account