Cox Exponential Jobs

Founding Engineer, AI Infra

Cox Exponential

Founding Engineer, AI Infra

Posted Yesterday

Remote or Hybrid

Hiring Remotely in CA, USA

Senior level

Remote or Hybrid

Hiring Remotely in CA, USA

Senior level

Design, build, and operate end-to-end training and inference infrastructure for large language and multimodal models. Improve efficiency (memory, parallelism, kernel optimizations), ensure robust scalable training and RL pipelines, optimize low-latency/high-throughput serving (quantization, caching, speculative decoding), manage multi-GPU and multi-cloud orchestration, and productionize new algorithms with strong observability and reproducibility.

The summary above was generated by AI

About Goaly

At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.

About the Role

You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale.

Key Responsibilities

Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton).
Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability.
Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding.
Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration.
Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems.

Requirements

5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models.
Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies.
Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling.
Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production.
Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems.
Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry).

Bonus Points

Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models.
Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects.
Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.

Similar Jobs

GC AI

Forward Deployed Engineer

An Hour Ago

Remote

United States

165K-350K Annually

Mid level

165K-350K Annually

Mid level

Artificial Intelligence • Legal Tech

Embed with customers to design, build, and deploy GC AI integrations into legal workflows. Develop production-grade API/webhook integrations, troubleshoot deployments, create reference implementations and documentation, and feed product insights back to Engineering to improve the platform. Travel to customer sites up to 25% as needed.

Top Skills: APIsLlmsPythonSdksTypescriptWebhooksWorkflow Automation

PwC

Oracle HCM Cloud - Manager

3 Hours Ago

Remote or Hybrid

99K-232K Annually

Senior level

99K-232K Annually

Senior level

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI

As a Manager in Oracle HCM, you'll help clients optimize HR processes by implementing Oracle solutions, leading teams, and ensuring project success through effective problem-solving and innovation.

Top Skills: Cc&BEbsFusionHyperionOracle ApplicationsOracle Hcm CloudPeoplesoftRiceSiebel

NBCUniversal

Architect

3 Hours Ago

Remote or Hybrid

Senior level

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development

Design and implement a scalable UGC framework in Unreal Engine: data models, runtime systems, scripting model, APIs, sandboxing, performance budgets, and AI-enabled content tooling. Partner across gameplay, online, tools, and AI/ML teams, drive prototypes and documentation, and mentor engineers to ensure a cohesive, extensible platform for creators.

Top Skills: Agent-Based SystemsAsset StreamingBlueprintsC++Ci/CdEntity Component System (Ecs)LlmsLuaMultithreadingPythonRestRpcSerializationUnreal EngineVerseWorld Partitioning

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
Key Industries: Artificial intelligence, Fintech
Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory