CAST AI Logo

CAST AI

AI Solutions Architect

Posted Yesterday
In-Office or Remote
Hiring Remotely in New York, NY, USA
Entry level
In-Office or Remote
Hiring Remotely in New York, NY, USA
Entry level
Serve as the founding customer-facing engineer for Kimchi, leading demos, customer environment setup, LLM inference optimization, model routing, governance policy implementation, and serverless API integrations. Translate customer feedback into product improvements while identifying revenue expansion opportunities. Explain complex AI concepts to non-ML engineers and optimize workloads for coding, reasoning, and planning use cases.
The summary above was generated by AI
Why Cast AI?

Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. By embedding autonomous decision-making directly into Kubernetes and cloud environments, Cast AI continuously optimizes performance, reliability, and efficiency in production.

The old way doesn't work. As Kubernetes and AI environments grow, manual decisions don’t. Cast AI replaces tickets, alerts, and manual tuning with continuous automation that adapts infrastructure as conditions change. Efficiency and cost savings follow naturally from that automation.

Over 2,100 companies already rely on Cast AI, including Akamai, BMW, Cisco, FICO, HuggingFace, NielsenIQ, Swisscom, and TGS.

Global team, diverse perspectives
We're headquartered in Miami, but our impact is international. We take a global and intentional approach to diversity. Today, Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APAC, bringing a wide range of perspectives into how we build and lead. 

Unicorn momentum
In January 2026, we achieved unicorn status with a strategic investment from Pacific Alliance Ventures, the corporate venture arm of Shinsegae Group (a $50+ billion Korean conglomerate). Our valuation now exceeds $1 billion, and we're just getting started.

Join us as we build the future of autonomous infrastructure.

About this role: 

You'll be the founding customer-facing engineer for our flagship Kimchi product within our Customer Success organization. Kimchi executes multi-model LLM workflows and autonomous agents with embedded governance and cost control. It intelligently routes across frontier and open-source models while enabling teams to establish strict budget caps, enforce model access policies, configure roles, and audit TCO via deep analytical tracking. Through Kimchi Studio, developers gain desktop access to agentic workflows, helping organizations accelerate adoption without sacrificing operational governance or spend visibility.
This is a hands-on, individual contributor role where you will live at the intersection of inference optimization, autonomous coding agents, multi-model routing and LLM governance. You will have the unique opportunity to architect and execute the growth strategy that scales Kimchi adoption and creates new revenue expansion opportunities for our existing customers.

Requirements:
  • Deep hands-on experience with LLMs: you've deployed them, fine-tuned them, evaluated them, or debugged their failure modes. You care about why model A outperforms model B for a specific task.
  • You know what it's like to operate in a fast-moving environment that’s prone to change. 
  • You've built or optimized inference systems out of curiosity. You know the difference between latency, throughput, and total cost of ownership. You ballpark costs and performance implications in your head.
  • You can explain complex inference concepts to non-ML engineers and help them see why it matters to their business.
  • You can think commercially. You understand revenue expansion, can identify expansion opportunities within customer organizations, and can design strategies to capture them. You're not just a technical person, you're someone who combines technical excellence with entrepreneurial thinking about growth.
Responsibilities:
  • Lead hands-on demos of our flagship inference products - Kimchi Coding, Studio and Teleport. 
  • Set up Kimchi Studio workspaces within customer environments. Help teams configure worktrees, integrations (Slack, repos, docs), and skills so agents understand their conventions.
  • Design and drive implementation of governance policies for customers: spending limits per team/developer, cost attribution by project, model routing preferences. Make sure the hard budget caps actually work for their org.
  • Help teams integrate Kimchi's serverless inference API. Optimize model selection and routing for their specific workloads (coding, reasoning, planning tasks).
  • Translate customer feedback into product input. Highlight where the agent fails, where Studio UX gets in the way, where governance policy doesn't match their needs.
What’s in it for you?
  • Enjoy a flexible, remote-first global environment.
  • Collaborate with a global team of cloud experts and innovators, passionate about pushing the boundaries of Kubernetes technology.
  • Equity options.
  • Get quick feedback with a fast-paced workflow. Most feature projects are completed in 1 to 4 weeks.
  • Spend 10% of your work time on personal projects or self-improvement. 
  • Learning budget for professional and personal development - including access to international conferences and courses that elevate your skills.
  • Team-building budget and company events to connect with your colleagues.
  • Equipment budget to ensure you have everything you need.
  • Extra days off to help maintain a healthy work-life balance.
Hiring process
  • Screening call with Recruiter
  • Hiring Manager interview
  • 1-2 additional interviews based on the role
  • Culture Check interview with an executive

*As part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.
*Please note that Cast AI does not provide any form of visa sponsorship/work permit.
#LI-Remote

Similar Jobs

Yesterday
Remote
USA
Senior level
Senior level
Software
Designs and publishes networking architectures for large-scale GPU and AI infrastructure, including InfiniBand, RoCEv2, Kubernetes, Linux hosts, virtual machines, and multi-tenant environments. Builds reproducible automation, prototypes emerging networking technologies, validates designs through benchmarks and proofs of concept, and advises customers, partners, and internal engineering teams. Presents technical findings through workshops, conferences, documentation, and research publications.
Top Skills: AnsibleBashBgpCalicoCiCiliumCluster ApiCumulus LinuxDcqcnDevlinkDpusDynamic Resource AllocationEcmpEcnEthernetEthtoolEvpn-VxlanGateway ApiGitGoGpudirect StorageGslbHelmIb_Write_BwInfinibandIommuIproute2IpxeIronicK0Rdent AiKubernetesKubevirtLinuxMetal3MultusNccl TestsNumaNvidia BluefieldNvidia ConnectxNvidia NcclNvidia Network OperatorNvidia Quantum InfinibandNvidia Spectrum-X EthernetNvme-OfOvn-KubernetesPci PassthroughPerftestPfcPkeysPxePythonRdmaRedfishRocev2SonicSr-IovSupernicsTerraformUfmUltra EthernetVrfs
3 Days Ago
In-Office or Remote
Minnesota, USA
122K-204K Annually
Senior level
122K-204K Annually
Senior level
Insurance
Leads enterprise AI-driven solution architecture, cloud-native platforms, automated software delivery pipelines, DevOps frameworks, and infrastructure automation. Designs architecture standards, governance, reusable patterns, and AI-enabled tooling; evaluates emerging technologies; guides implementation and troubleshooting; and mentors architects and engineering teams. Requires extensive cross-disciplinary technology experience, cloud and hybrid-cloud architecture expertise, AI-assisted development, CI/CD automation, and enterprise-scale solution design.
Top Skills: AWSAzure FoundryCi/CdClaude CodeCloud ComputingCodexConfluenceDevOpsDistributed SystemsGithub CliGithub CopilotHybrid CloudInfrastructure As CodeJIRAMcp FrameworksAzureMicrosoft FabricSalesforceTerraform
Yesterday
Remote
2 Locations
111K-194K Annually
Mid level
111K-194K Annually
Mid level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software
Build and optimize conversational AI and agentic virtual agents on Genesys Cloud. Responsibilities include prompt engineering, NLU model development, intent and entity design, production conversation analysis, evaluation framework creation, performance optimization, scripting, and customer collaboration. The role focuses on improving reliability, task completion, edge-case handling, and user experience across LLM and NLU systems.
Top Skills: Conversational AiGenesys CloudLlmsNlpNluNode.jsPrompt EngineeringPython

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account