Metasys Logo

Metasys

DevOps Engineer Internship

Reposted 2 Days Ago
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Build and maintain automated infrastructure and CI/CD pipelines using Terraform, Docker, Traefik, and Makefile. Implement observability (Prometheus, Grafana, Loki, Tempo, OpenTelemetry), backups (pgBackRest/Postgres15), and cloud/Linux administration. Automate deployment and monitoring for AI agent services and support monorepo workflow with SRE and DevSecOps teams.
The summary above was generated by AI
Overview: Infrastructure Automation and CI/CD

The DevOps Engineer is responsible for automating, streamlining, and maintaining the infrastructure and deployment pipelines for our entire integrated platform. You'll ensure rapid, reliable, and consistent delivery of our e-commerce storefront, internal supply chain tools (MES, WMS, OMS), and cutting-edge AI agent services, primarily utilizing Infrastructure-as-Code (IaC) and robust CI/CD practices.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will build and maintain the fully automated platform that underpins our entire tech stack.

  • Infrastructure-as-Code (IaC): Design, implement, and manage infrastructure provisioning across all environments using Terraform for our Oracle Cloud Free VMs (or equivalent cloud resources). Ensure infrastructure is auditable, repeatable, and secure.

  • CI/CD Pipeline Management: Set up and maintain the Continuous Integration and Continuous Deployment (CI/CD) pipelines, primarily driven by Makefile and automated testing, for the Node.js/NestJS modular monolith and Next.js frontend applications.

  • Containerization & Orchestration: Manage application containerization using Docker. Define deployment strategies, service discovery, and traffic routing using Traefik for our containerized services.

  • Observability Implementation: Implement, manage, and optimize the comprehensive logging, monitoring, and alerting system using our selected stack: Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. Ensure end-to-end tracing is functional across the complex business flow (MES → WMS → OMS).

  • Resilience & Backups: Collaborate with the SRE team to implement high-availability features and maintain automated backup solutions, including pgBackRest for our PostgreSQL 15 database.

  • Workflow: Maintain the Monorepo structure for streamlined code management and deployment separation across applications (web / admin / API) and domain packages.

Required Technologies & Tools

Candidates must possess mandatory expertise in our core infrastructure and automation stack:

  • Infrastructure-as-Code: Expert proficiency in Terraform.

  • Containerization: Expert proficiency in Docker and deployment strategies (e.g., Traefik, orchestration concepts).

  • CI/CD: Hands-on experience building and maintaining complex pipelines (Makefile, Jenkins/GitHub Actions/GitLab CI concepts).

  • Observability: Strong implementation experience with Prometheus, Grafana, Loki, and OpenTelemetry.

  • Cloud & Linux: Experience with Linux administration and managing cloud resources (Oracle Cloud or equivalent).

AI Agent Focus

You will ensure the scalable and monitored deployment of the AI layer.

  • Deployment Automation: Automate the packaging and deployment pipelines for resource-intensive AI agent services and LLM fine-tuning environments.

  • Resource Monitoring: Set up specific monitoring and alerts to track the performance, resource consumption, and cost of the AI agent compute demands.

Success Metrics & Career Path

Performance will be measured by:

  • Deployment Frequency: Reduction in lead time and increased frequency of stable deployments.

  • Infrastructure Stability: Reliability of provisioned infrastructure (minimal unplanned downtime).

  • Observability Coverage: Completeness and reliability of monitoring, logging, and tracing across all production services.

Mentorship Structure: Reports to the Solution Architect or Head of Technology, working closely with the SRE, DevSecOps, and Backend engineering teams to build a robust platform.

Similar Jobs

31 Minutes Ago
Remote or Hybrid
45K-85K Annually
Junior
45K-85K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handles inbound calls and warm leads to understand customers’ insurance needs, recommend appropriate coverages, and convert prospects into policyholders. The role includes paid training and Property & Casualty licensing, customer communication, sales closing, and brand representation. Employees work remotely, follow assigned evening and weekend schedules, maintain required home-office and internet standards, and remain in their resident state for at least one year.
Top Skills: Cable InternetDsl InternetFiber InternetPc
4 Hours Ago
Remote or Hybrid
83K-139K Annually
Senior level
83K-139K Annually
Senior level
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Manages complex commercial energy claims from initial report through resolution, including coverage analysis, liability assessment, reserves, litigation, settlements, and high-dollar exposures. Coordinates defense counsel, experts, legal, underwriting, reinsurance, and claims teams. Advises leadership on coverage and legal developments, mentors claims examiners, oversees litigation strategy, and supports claims process improvement while maintaining regulatory and confidentiality standards.
4 Hours Ago
Remote or Hybrid
CA, USA
185K-327K Annually
Senior level
185K-327K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Design, build, scale, and operate high-volume payment APIs and distributed systems serving global traffic. Lead complex engineering projects from requirements through production, improve reliability, performance, observability, fault tolerance, and security, and participate in on-call and incident response. Collaborate across product and engineering teams, use AI-assisted development tools responsibly, and mentor engineers through technical reviews and documentation.
Top Skills: AWSGoKafkaKotlinTypescript

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account