CloudBolt Logo

CloudBolt

Senior ML Engineer

Posted 13 Days Ago
In-Office
Rockville, MD
170K-245K
Senior level
In-Office
Rockville, MD
170K-245K
Senior level
Own and develop a production recommendation engine for Kubernetes resource optimization. Design time-series forecasting models, data-quality pipelines, safety guardrails, regression tests, and production metrics for CPU, memory, GPU, and JVM workloads. Investigate customer issues, improve recommendation accuracy, maintain Python services, and guide ML strategy while collaborating with platform engineering, product, and leadership.
The summary above was generated by AI

Description

At CloudBolt we help organizations maximize the value of their cloud investments through greater visibility, governance, and cost optimization across complex cloud environments. Our platform empowers teams to make smarter cloud decisions by turning insights into action, helping businesses improve efficiency, control spending, and accelerate innovation.

As a remote-first global SaaS company, CloudBolt is committed to fostering a collaborative, inclusive, and high-performing culture where employees can do their best work and make a meaningful impact.

Learn more at www.cloudbolt.io.

As a Senior ML Engineer, you'll be building and owning the recommendation engine at the heart of our Kubernetes resource optimization product, StormForge, cutting our customers' cloud spend without putting a workload at risk. This is a critical hands-on position at the intersection of applied machine learning and production engineering, where the hardest problems are as much about data quality, guardrails, and knowing when not to recommend as they are about forecasting itself. You'll need to bring rigor and curiosity in equal measure, designing time-series models and the fallback strategies around them, proving their behavior through repeatable regression testing, and iteratively raising the ceiling on accuracy and safety as we discover and learn more about the workloads our customers run.
 

Responsibilities

  • Own the recommendation engine end to end: model selection, algorithm design, preprocessing, and the guardrails that keep recommendations safe to apply to live production workloads. 
  • Design, evaluate, and productionize time-series forecasting and statistical models (e.g., Prophet, percentile-based estimation) that right-size Kubernetes workloads across CPU, memory, GPU, and JVM heap. 
  • Build and maintain the data-quality layer: detecting and filtering anomalies, load-test windows, startup spikes, and autoscaling artifacts from production telemetry before it reaches a model. 
  • Define and continuously improve how we measure recommendation quality: regression testing against golden datasets, behavioral validation, and accuracy/safety metrics in production. 
  • Investigate and resolve recommendation quality issues reported from customer environments, tracing them through data, preprocessing, and model behavior. 
  • Serve as the team's machine learning authority: guide technical direction on ML questions, make model-vs-heuristic tradeoff calls, and clearly communicate them to platform engineers, product, and leadership. 
  • Write production-grade Python for models and pipelines alike and share ownership of the surrounding service (message consumption, metrics ingestion, caching) with the rest of the team. 
  • Prototype and validate new optimization capabilities (new resource types, new algorithms, new workload classes) from research through gradual, feature-flagged rollout.
  • Stay current on time-series forecasting and resource optimization techniques, and pragmatically evaluate which are worth adopting.

Requirements

Must have strong experience…? 

  • Master's degree or higher in a quantitative field (Computer Science, Machine Learning, Statistics, Applied Mathematics). 
  • 5+ years of software engineering experience, with at least 3 years building and operating machine learning or statistical systems in production. 
  • Expert-level Python: you write typed, tested, production-grade code, and you're fluent in numpy or similar array-based numerical computing. 
  • Hands-on experience with time-series analysis and forecasting: seasonality, trend decomposition, anomaly detection, and classical statistical methods (percentiles, distributions, smoothing), not just deep learning. 
  • Experience testing ML systems rigorously: regression testing against known-good baselines, behavioral validation, and reasoning about numerical reproducibility. 
  • Working knowledge of Kubernetes: resource requests and limits, autoscaling behavior, and what happens to a workload when it's under-provisioned (OOM kills, CPU throttling). 
  • Comfort owning a production service, not just a model: queues, caches, retries, observability, and debugging issues in customer environments from logs and metrics. 
  • Clear written and verbal communication: as an ML engineer on the team, so you must be able to explain model behavior and tradeoffs to platform engineers, product managers, and customers. 

Experience in the following is beneficial 

  • Experience with Prophet or similar forecasting libraries. 
  • Prometheus/PromQL and experience working with metrics at scale. 
  • Cloud cost optimization, capacity planning, or infrastructure efficiency background 
  • AWS (S3, Managed Prometheus). 
  • Experience being the ML domain expert on a team of generalists. 

We Offer

Our US benefits package includes: 

  • Medical/Dental/Vision coverage
  • 401k with Company Match
  • Health & Dependent Care FSA
  • Unlimited PTO
  • 11 Company Holidays
  • Volunteer/Community Engagement Day
  • Tuition Reimbursement 
  • Paid Parental Leave
  • Equity Grants
  • Home internet Reimbursement

Base Salary Range: $170,000 - $245,000 USD Annually, actual compensation will be determined based on job-related factors, including but not limited to relevant experience, skills, geographic location, internal equity, and business needs.

About CloudBolt

CloudBolt is an Equal Opportunity Employer committed to building a diverse and inclusive workplace. We celebrate diversity and do not discriminate on the basis of race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity or expression, national origin, age, disability, protected veteran status, or any other characteristic protected by applicable law.

If you require a reasonable accommodation during the application or interview process, please contact [email protected].

Similar Jobs

4 Days Ago
Hybrid
2 Locations
77K-202K Annually
Senior level
77K-202K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Supports payer modernization and health services consulting engagements by analyzing operational challenges, optimizing processes, applying business analytics, engaging stakeholders, implementing quality standards, and guiding transformation initiatives. Develops operational support frameworks, conducts market research, manages change, and advises clients on improving service delivery and business performance while mentoring junior team members.
Top Skills: GitMlops
19 Days Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
170K-286K Annually
Senior level
170K-286K Annually
Senior level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Build and operate production machine learning systems for Safety AI, including low-latency APIs, data pipelines, model serving, evaluation, monitoring, and rollout infrastructure. Process large-scale camera and telematics data, optimize cloud and edge-to-cloud execution, track model performance and drift, and partner with applied scientists, firmware engineers, platform teams, and product managers to deliver reliable safety features.
Top Skills: SparkAWSAzureC++Ci/CdComputer VisionDockerGCPGoGpusGrafanaInfrastructure As CodeJavaKubernetesMachine LearningMlflowPythonPyTorchRayRay ServeScala
Yesterday
In-Office
New York, NY, USA
115K-230K Annually
Senior level
115K-230K Annually
Senior level
Insurance
Design, deploy, and operate production machine learning systems across predictive analytics, automation, and decision support. Own the full ML lifecycle, build scalable batch and real-time pipelines, optimize model serving, and contribute to shared AI infrastructure. Provide technical leadership and mentorship while collaborating with data science, engineering, operations, and product teams. Ensure monitoring, security, privacy, compliance, explainability, governance, and reliability for AI systems in a regulated environment.
Top Skills: AirflowDbtFeature StoresJavaLlmsModel RegistriesPythonPyTorchRetrieval-Augmented Generation (Rag)Scikit-LearnSnowflakeSparkTensorFlow

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account