BJAK Logo

BJAK

Principal Machine Learning Engineer

Posted 25 Days Ago
Remote or Hybrid
Hiring Remotely in United States
Expert/Leader
Remote or Hybrid
Hiring Remotely in United States
Expert/Leader
Design, build, and operate large-scale ML systems across data, training, evaluation, inference, and deployment. Lead GPU-optimized training and inference pipelines, ensure reliability, performance, safety, and seamless integration into products. Own production deployments, monitoring, and cross-functional best practices to drive measurable model improvements and efficient research-to-production cycles.
The summary above was generated by AI
About the Role

There are over 5 billion users using basic applications today such email, notes, tasks, calendar and they're not AI-native. Our mission is to build proactive applications for anyone in the world, who are not used to complex prompting. We aim to bring intelligence to conversations, errands, organising and workflows, with minimal to no prompting.

Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. We believe products will greatly reduce hallucinations

Our objective is to organise anyone's life, allowing us all to spend time on valuable and meaningful things

As Principal Machine Learning Engineer, you own the execution layer of our intelligence, turning research and model capabilities into reliable, scalable production systems.

You will work across the model lifecycle: data, training, evaluation, inference, and deployment. This is a hands-on leadership role for someone who wants to operate at the intersection of research, systems, and product.

 
What You'll Own
  • Own the end-to-end ML systems powering our company, from data and training to evaluation, inference, and deployment.

  • Build and evolve training and fine-tuning pipelines for large models.

  • Design evaluation systems that measure capability, robustness, safety, and real-world product performance.

  • Architect high-performance inference systems, optimizing latency, GPU utilization, memory, cost, and reliability.

  • Build data pipelines and systems for high-quality real-world and synthetic training data.

  • Establish reliable production infrastructure for deploying, monitoring, and continuously improving models.

  • Partner closely with research and application engineering to turn model capabilities into product improvements.

  • Make pragmatic technical trade-offs and rapidly iterate based on real-world performance.

 
What We're Looking For
  • Experience building and shipping ML systems used in production, not just research prototypes.

  • Strong understanding of modern large-model training, fine-tuning, evaluation, and inference.

  • Strong software engineering and systems fundamentals.

  • Experience operating ML workloads at meaningful scale, particularly GPU-based systems.

  • Strong technical judgment and the ability to navigate ambiguous problems independently.

  • A bias toward experimentation, measurement, and shipping.

  • High standards for correctness, reliability, and production quality.

 
Outcomes
  • Research and models reliably translate into production-ready solutions with clear performance and quality targets.

  • ML pipelines, training loops, and inference systems are stable, efficient, and maintainable.

  • Production issues are detected, debugged, and resolved quickly, minimizing user impact.

  • Team members are supported, aligned, and able to deliver high-impact ML work with minimal friction.

  • Iterations on models and systems are measurable, safe, and improve user experience over time.

 
Tech Stack
  • Python

  • PyTorch / JAX

  • GPU-based training and inference system

 
Ideal Experience
  • You have built or shipped real ML systems used by people, not just demos.

  • You are comfortable working with large models and understanding their failure modes.

  • You write strong, production-grade code and care about system correctness.

 
How We Work

We are a small, high-talent-density, hands-on team. Engineers have broad ownership and are expected to exercise strong judgment and execute independently.

We make decisions quickly, work closely together, and balance speed with engineering fundamentals. We care less about process and more about building something exceptional.

 
Interview process

If there appears to be a fit, we'll reach to schedule 3, but no more than 4 interviews.

Applications are evaluated by our technical team members. Interviews will be conducted via virtual meetings and/or onsite.

We value transparency and efficiency, so expect a prompt decision. If you've demonstrated the exceptional skills and mindset we're looking for, we'll extend an offer to join us. This isn't just a job offer; it's an invitation to be part of a team that's bringing AI to have practical benefits to billions globally.

Similar Jobs

8 Hours Ago
Remote or Hybrid
United States
195K-210K Annually
Expert/Leader
195K-210K Annually
Expert/Leader
Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Own the MLOps foundation for a shared enterprise AI platform. Build production code, infrastructure automation, delivery pipelines, model and prompt lifecycle tooling, serving systems, observability, evaluation, and cost controls. Establish technical guardrails, governance, security, and operational standards while embedding with teams to move AI workloads into reliable production. Provide principal-level technical leadership, mentor engineers, influence architecture, and communicate trade-offs with technical and executive stakeholders.
Top Skills: Amazon BedrockAmazon Ec2Artifact RegistriesAudit LoggingAutomated Ai EvaluationAutoscalingAWSAws IamCi/CdData IngestionDataset VersioningDistributed TracingEmbeddingsGitopsIdentity ManagementIncident ResponseIndexingInfrastructure-As-CodeKubernetesLineage TrackingLoad BalancingMlopsModel ServingModel VersioningMulti-Tenant IsolationNetworkingObservabilityPolicy EnforcementPrompt VersioningPythonRag PipelinesRegression TestingSecrets ManagementService-Level ObjectivesTerraformVector Stores
14 Days Ago
In-Office or Remote
196K-309K Annually
Senior level
196K-309K Annually
Senior level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Lead the technical vision and development of scalable ML and AI evaluation systems for Atlassian’s Rovo Chat. Build agentic capabilities, datasets, simulations, automated judges, experimentation platforms, and quality signals. Identify failure modes, ship interventions, and improve customer outcomes while balancing quality, latency, reliability, safety, and cost. Provide cross-organizational technical leadership, mentorship, design guidance, and strategic direction for AI products.
Top Skills: Agentic SystemsArtificial IntelligenceExperimentation PlatformsFine-TuningGenerative AiInference SystemsLarge Language ModelsMachine LearningModel EvaluationPromptingRetrieval
20 Days Ago
Remote or Hybrid
165K-282K Annually
Expert/Leader
165K-282K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead the architecture and end-to-end development of production-grade AI/ML systems, including training, deployment, monitoring, MLOps, deep learning, NLP, ASR, and agentic AI workflows. Translate research into scalable enterprise solutions, establish engineering and governance standards, evaluate emerging tools, communicate technical tradeoffs to leadership, and mentor engineers. The role requires expertise in Python, deep learning frameworks, cloud ML platforms, ML pipelines, statistics, and production LLM or agentic systems.
Top Skills: AsrAws SagemakerAzure MlCi/CdDistributed Data SystemsGcp Vertex AiGenerative AiLarge Language ModelsMlopsNlp/NluPythonPyTorchRetrieval-Augmented GenerationTensorFlow

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account