Weekday, Inc. Logo

Weekday, Inc.

Cloud / DevOps Engineer (Infra & IaC)

Posted 6 Days Ago
Remote
Hiring Remotely in United States
75-110 Hourly
Mid level
Remote
Hiring Remotely in United States
75-110 Hourly
Mid level
Design and evaluate cloud infrastructure, Kubernetes, IaC, AWS, and CI/CD solutions for GenAI training and inference environments. Create technical scenarios, reference solutions, evaluation rubrics, and feedback for AI-generated infrastructure outputs. Troubleshoot production Kubernetes and distributed-system failures while applying cloud architecture, security, reliability, observability, and operational best practices.
The summary above was generated by AI

This role is for one of our clients

Compensation: $75 - $110 per hour

We are seeking an experienced Cloud / DevOps Engineer (Infra & IaC) to contribute to a cutting-edge GenAI environment focused on building and improving large-scale AI training and inference infrastructure.

The ideal candidate will bring strong, hands-on expertise in Kubernetes, AWS cloud services, Infrastructure-as-Code (IaC), and CI/CD. You will apply your real-world infrastructure engineering experience to evaluate technical workflows, create high-quality reference solutions, identify gaps in AI-generated outputs, and help establish rigorous standards for cloud and DevOps reasoning.

This is a full-time engagement requiring 40 hours per week, Monday through Friday.


RequirementsKey Responsibilities
  • Collaborate with research and engineering teams to identify knowledge gaps and improve AI model performance across cloud infrastructure, DevOps, Kubernetes, and Infrastructure-as-Code domains.
  • Design realistic and technically challenging tasks covering Kubernetes troubleshooting, AWS service integration, infrastructure automation, and production operations.
  • Develop accurate, detailed reference solutions for complex infrastructure engineering scenarios.
  • Review and evaluate AI-generated technical solutions for correctness, reliability, scalability, security, and adherence to production best practices.
  • Provide clear, structured written feedback highlighting technical gaps, incorrect assumptions, and opportunities for improvement.
  • Create detailed evaluation criteria, rubrics, and benchmarks for assessing Kubernetes troubleshooting, IaC architecture, AWS integrations, and CI/CD reasoning.
  • Develop scenarios involving cluster failures, infrastructure automation, deployment workflows, service integrations, and operational reliability.
  • Work closely with other technical subject matter experts to maintain consistency, accuracy, and quality across evaluation datasets.
  • Translate practical production experience into structured guidance that can be used to improve AI-generated infrastructure solutions.
Core Qualifications
  • 4+ years of professional experience in Cloud Infrastructure, DevOps, Site Reliability Engineering, Platform Engineering, or a closely related field.
  • Strong hands-on experience managing Kubernetes in production environments, including diagnosing, troubleshooting, and resolving cluster failures and operational issues.
  • Experience with Kubernetes beyond simply writing manifests or consuming managed Kubernetes control planes.
  • Proven production experience with Infrastructure-as-Code, particularly Terraform and/or AWS CDK.
  • Strong practical knowledge of AWS cloud services, including production integration with services such as:
    • AWS Lambda
    • API Gateway
    • DynamoDB
  • Experience designing, implementing, and maintaining CI/CD pipelines for production workloads.
  • Strong understanding of cloud architecture, infrastructure automation, deployment strategies, observability, reliability, and operational best practices.
  • Demonstrated career progression with increasing ownership and responsibility in infrastructure, DevOps, or platform engineering.
  • Ability to commit reliably to 40 hours per week during standard weekdays.
  • Excellent written and verbal communication skills, with the ability to explain complex technical concepts and engineering decisions clearly.
  • Strong analytical and troubleshooting abilities, particularly when diagnosing distributed systems and infrastructure failures.
Preferred Skills
  • Experience working with large-scale cloud infrastructure or highly distributed systems.
  • Familiarity with Kubernetes networking, security, storage, scaling, and cluster lifecycle management.
  • Experience implementing infrastructure security and reliability best practices.
  • Knowledge of AWS architecture patterns and cloud-native application design.
  • Experience with GitOps, containerization, monitoring, logging, and observability platforms.
  • Familiarity with modern DevOps and platform engineering methodologies.
  • Experience reviewing or evaluating technical documentation, engineering solutions, or AI-generated outputs.
What You’ll Contribute

In this role, your production infrastructure expertise will help establish high-quality standards for AI systems working with complex Cloud, DevOps, Kubernetes, AWS, and IaC problems.

You will play a key role in transforming practical engineering knowledge into structured tasks, reference solutions, evaluation frameworks, and high-quality technical feedback that can improve the capabilities of next-generation AI models.

Equal Opportunity

We are committed to providing equal employment opportunities to all qualified candidates. Employment decisions are made without regard to legally protected characteristics, and reasonable accommodations are available throughout the hiring and engagement process upon request.

Similar Jobs

7 Minutes Ago
Easy Apply
Remote
United States
Easy Apply
148K-222K Annually
Expert/Leader
148K-222K Annually
Expert/Leader
Fintech • Social Impact • Financial Services
Leads OppFi’s lifecycle and brand marketing strategy across acquisition, retention, repayment, servicing, and collections. Owns omnichannel communications through Braze, lifecycle KPIs, brand standards, creative direction, customer insights, and go-to-market strategy. Manages and mentors lifecycle marketing, design, and marketing technology teams while partnering cross-functionally with Product, Data, Engineering, Operations, Legal, Finance, and executives. Drives experimentation, optimization, customer experience, and responsible marketing within a regulated FinTech environment.
Top Skills: BrazeFigmaJIRAMonday.Com
An Hour Ago
Remote or Hybrid
New York City, NY, USA
265K-332K Annually
Senior level
265K-332K Annually
Senior level
Consumer Web • Healthtech • Professional Services • Social Impact • Software
Lead the Ranking & Relevance team of software and machine-learning engineers. Own retrieval, ranking, recommendation, and personalization systems that match patients with therapists. Drive the transition from filter-based matching to outcome-aware machine learning, establish shared marketplace objectives, improve experimentation and model monitoring, and build evaluation and drift-detection capabilities. Hire and develop engineers, set technical standards, and partner with product, data science, payer, and provider engineering teams in a sensitive healthcare domain.
Top Skills: Ai Coding ToolsFeature PipelinesInformation RetrievalLearning-To-RankMachine LearningOnline ExperimentationPersonalization SystemsRecommendation Systems
An Hour Ago
Easy Apply
Remote
United States
Easy Apply
181K-214K Annually
Expert/Leader
181K-214K Annually
Expert/Leader
Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Leads Motive’s internal communications strategy, narrative, and cadence across company-wide forums, executive messaging, video, Slack, intranet, email, and other channels. Advises executives, supports crisis and change communications, aligns internal and external messaging, measures employee sentiment, and improves engagement programs. The role also partners with Corporate Communications, Legal, and People teams and mentors communications talent while remaining an individual contributor.
Top Skills: Ai-Enabled ToolsEmailIntranetSlackVideo Communications

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account