Cerebras Systems Inc. Logo

Cerebras Systems Inc.

Software Engineer, Kernel Reliability

Posted 5 Hours Ago
Remote
Hiring Remotely in United States
Entry level
Remote
Hiring Remotely in United States
Entry level
Develop kernel-centric reliability solutions for AI compute clusters and production services. Responsibilities include debugging failures, building diagnostic tools, supporting incident response, performing root-cause analysis, improving kernel and software reliability, and collaborating with systems, hardware, ASIC, and architecture teams on reliability-focused designs.
The summary above was generated by AI

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About The Role

We're looking for a deeply technical, hands-on software engineer to join our on-field Kernel Reliability team. You'll help tackle a critical challenge: improving the reliability of our advanced compute clusters and the underlying inference, training, and internal production services. In this role, you'll work close to the code and design solutions that will scale with our rapidly growing system production and software service offerings. If you have strong fundamentals in systems, debugging, and failure analysis—and enjoy building tools and solving hard reliability problems—we want to hear from you. New college graduates are welcome.

Responsibilities

  • Contribute to the technical roadmap and execution for kernel-centric reliability of our internal and customer-facing systems.

  • Partner with System and Cluster Operations teams to reduce system and service downtime after failure through tooling, analysis, and hands-on debugging support.

  • Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis.

  • Collaborate with software teams to improve the software stack—including kernels—to improve on-field debugging and failure analysis.

  • Work with ASIC and hardware architecture teams to co-design next-generation architectures with reliability and ease of debug in mind.

  • Participate in incident response, root-cause analysis, and post-mortems; drive follow-ups that measurably improve reliability over time.

Skills & Qualifications

  • We recognize great engineers come from different backgrounds. If you're excited about the role, we encourage you to apply even if you don't meet every qualification.

  • Required (or demonstrated through projects/internships/coursework):

    • Strong programming skills in C/C++ and Python.

    • Solid foundations in operating systems, computer architecture, and systems programming fundamentals.

    • Ability to debug complex issues using logs, traces, and standard debugging workflows; interest in root-cause analysis.

Preferred Skills & Qualifications

  • Exposure to parallel and distributed programming (message passing, multicore, GPU, embedded, etc.).

  • Experience building or using debug/diagnostic tools (debuggers, core dump handling, tracing, sanitizers, profilers, etc.).

  • Familiarity with debugging distributed and parallel applications (deadlocks, livelocks, race conditions, etc.).

  • Knowledge of computer architecture concepts (instruction pipelining, multithreading, networking, memory systems, etc.).

  • Operations & Monitoring: familiarity with monitoring, incident response, and post-mortem culture.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.

  2. Publish and open source their cutting-edge AI research.

  3. Work on one of the fastest AI supercomputers in the world.

  4. Enjoy job stability with startup vitality.

  5. Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Similar Jobs

36 Minutes Ago
Remote
142K-193K Annually
Mid level
142K-193K Annually
Mid level
Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
Owns post-sale relationships for a portfolio of B2B SaaS customers, driving adoption, value realization, retention, renewals, and net revenue retention. Builds strategic stakeholder relationships, develops mutual success plans, leads business reviews and renewal conversations, identifies customer risks, and coordinates cross-functional mitigation and growth strategies. Uses customer data, product usage, sentiment, and business context to prioritize actions and improve outcomes.
Top Skills: B2B Saas
An Hour Ago
Remote or Hybrid
121K-169K Annually
Entry level
121K-169K Annually
Entry level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Owns product vision and roadmap for Mastercard’s identity and payments data products. Responsibilities include prioritizing product development, creating data-backed business cases, aligning cross-functional stakeholders, defining requirements, managing risks, tracking delivery, representing customer needs, and developing future commercial opportunities. The role also involves mentoring product managers, managing globally distributed teams, and communicating with executive stakeholders.
Top Skills: Artificial IntelligenceMachine Learning
An Hour Ago
Remote
75K-85K Annually
Entry level
75K-85K Annually
Entry level
Artificial Intelligence • Software
Own the full sales cycle for net-new SMB accounts in Canada, from prospecting through negotiation and close. Build relationships with CPA, audit, advisory, and assurance firms; conduct discovery; present Fieldguide’s AI audit platform; develop territory and account plans; apply MEDDICC; maintain CRM accuracy and forecasts; collaborate cross-functionally; and attend networking events. The role requires achieving sales targets and involves up to 30% regional and national travel.
Top Skills: AICRM

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account