Voltus Logo

Voltus

Senior Software Engineer, Infrastructure

Posted 3 Hours Ago
Remote
Hiring Remotely in USA
160K-190K Annually
Senior level
Remote
Hiring Remotely in USA
160K-190K Annually
Senior level
Own and evolve Voltus’s infrastructure platform, including Kubernetes and Nomad workloads, AWS architecture, identity and secrets, stateful systems, observability, infrastructure as code, CI/CD, and developer tooling. Lead business-critical migrations with minimal downtime, establish reliable deployment and rollback practices, mentor engineers, participate in on-call, and build secure infrastructure supporting AI-assisted development.
The summary above was generated by AI
About Us

Our mission is to be the Distributed Energy Platform that fulfills the promise of the energy transition. The Voltus platform connects any distributed energy resource (DER) to any energy market across the US and Canada, providing flexibility, reliability, and resilience to the grid while generating revenue for our partners. By partnering with industry leaders and the DER ecosystem, we are building the decarbonized, distributed, decentralized, and digitized energy system of the future.

Voltus is a remote and virtual company, distributed across the US, Canada, and abroad, with team members in New York, San Francisco, Boston, Toronto, Chicago, Los Angeles, Washington DC, and many other cities.

About the Role

Are you interested in building the technical foundation of the worldwide transition to clean energy? Do you enjoy working with a highly motivated and talented team to deliver mission critical software?

Voltus is growing our Infrastructure team to help deploy, manage, troubleshoot, and enhance the Platform and tooling that every Voltus engineer and our dispatch-critical systems depend on. As a Senior Software Engineer on Infrastructure, you will own significant pieces of our core Platform, mentor other engineers, and help set the technical direction for how we build, deploy, and operate services at Voltus. You will join a team of [N] engineers supporting roughly [M] engineers across the org, so the leverage on your work is high and so is the responsibility.

Our Platform runs on HashiCorp's Nomad, Consul, and Vault in AWS. We are specifically looking for someone who brings deep Kubernetes experience to this team, because we expect to run Kubernetes workloads alongside Nomad and that is depth we do not have in house today. You will not be the only person who knows containers, but you may well be the one who knows Kubernetes best, and we want you to shape how we adopt and operate it.

You will develop the tooling that lets every engineering team ship reliably, build the testing and observability that keeps our real-time systems healthy, and help shape how we safely fold AI-assisted development and tooling into our day-to-day engineering.

When Voltus commits capacity to a grid operator, we are obligated to deliver it. That shapes how this team works: changes are planned around live dispatch windows, sequenced so each step can be rolled back, and owned by the person who made them, including on call.

Voltus is a fully distributed company and this role is remote. You can work from anywhere, but you must overlap with the Eastern time zone for at least 4 hours of the workday.

What You'll Do

    Own core Platform services and major migrations end to end, from proposal to production. Our team regularly leads multi-month migrations of stateful, business-critical systems with no customer-visible downtime.

    Run our containerized workloads and the delivery path that ships them — orchestration and scheduling, GitOps-style deploys, progressive rollout and rollback, and service mesh. Bring Kubernetes practice to a team that runs Nomad today, and help us decide what belongs where rather than adopting a second scheduler for its own sake.

    Architect and operate our AWS foundation. Multi-account structure and governance, IAM and cross-account access, VPC and Transit Gateway, PrivateLink, DNS and certificates, and the egress paths our dispatch traffic reaches grid operators over. On those paths, a changed address or an expired certificate is a market outage.

    Treat identity, secrets, and encryption as first-class infrastructure — workload identity, least privilege, SSO and OIDC, machine-to-machine credentials, Vault, and key ownership and rotation. As we open more of production to cloud and AI access, you make sure every human and workload has exactly the access it needs and no more.

    Operate the stateful systems everything else sits on: database upgrades and replication, message brokers in the critical path, caches and time-series stores, and the unglamorous part — backup coverage and restores you have actually tested.

    Instrument deeply for observability. Distributed tracing, meaningful metrics and SLOs, and the testing frameworks that keep dispatch comms and market message flows healthy. Where monitoring has sprawled into overlapping tools with alerts that live only in a UI, consolidate it into something defined in code that an on-call engineer can actually reason about under pressure.

    Build the infrastructure as code and developer tooling the whole org depends on, using Terraform, GitHub, Buildkite, Docker, Nomad, and our internal tools. That includes bringing older infrastructure under code: importing what was built by hand, detecting drift, and making what is in code match what is actually running.

    Help build the infrastructure that AI runs on. AI-assisted development is central to where we are going: you will build the guardrails that let engineers and AI tools reach internal systems securely, and use those tools yourself to move faster and to understand and document large systems quickly.

What We're Looking For

  • Strong production software development experience in Go and/or Python. You build and maintain real services and tooling, write tests, and care about code quality, not just scripts.

  • Roughly 6+ years of professional software engineering, with several in DevOps / SRE operating production systems. You have owned deployments and been on the hook for reliability and on-call.

  • A track record of owning meaningful infrastructure projects end to end, ideally including a migration of a stateful or business-critical system with minimal disruption. You plan around operational windows, sequence work so each step has a rollback, and would rather phase a migration over weeks than take one clever shortcut.

  • Real depth in AWS, beyond launching resources in a single account. You have worked with multi-account organizations, IAM and cross-account access, VPC and network design, DNS, secrets management, and encryption key management, and you understand how those pieces constrain each other. You know the difference between provisioning and configuration management, and you default to least privilege.

  • Deep, hands-on Kubernetes experience in production. You have operated real clusters, not just deployed to someone else's: upgrades, networking and ingress, RBAC, resource management and autoscaling, and debugging a workload that is misbehaving under load. EKS and the surrounding ecosystem (ArgoCD or Flux, Helm, operators, multi-tenant clusters) is a strong plus. This is the clearest gap on our team and a large part of why we are opening this role.

  • Strong infrastructure-as-code skills (Terraform or similar). You have worked in a codebase that did not start out fully covered and improved it.

  • Real depth in monitoring and observability, and not only as a consumer of one managed vendor. You instrument systems thoughtfully, define metrics, alerts, and dashboards as code, and can introduce distributed tracing where none exists. Comfort operating the open-source stack (Prometheus, Grafana, and log and search clusters such as Elasticsearch or OpenSearch) matters here: capacity, retention, index lifecycle, and upgrades, not just querying it.

  • Hands-on operational experience with stateful systems: relational databases, message brokers, or both. You have done a version upgrade, a replication change, or a restore under pressure and know why backups you have not tested do not count.

  • Comfort operating and improving systems you did not build. Some of our most important work involves reading unfamiliar code, mapping undocumented dependencies, and making careful changes to systems whose original authors are not available to ask.

  • You communicate clearly, write good documentation, mentor teammates, and can drive cross-team coordination on shared infrastructure.

  • You are genuinely excited about AI-assisted development and the role of AI in modern infrastructure. You either use these tools today (Claude Code, MCP, agents) or are eager to, and you want to help build the platform that supports them safely.

Nice to Have

  • Experience with the HashiCorp stack (Nomad, Consul, Vault).

  • Experience introducing Kubernetes to an organization that did not already run it, or operating it alongside another scheduler.

  • Experience running self-hosted CI and build tooling (Jenkins, ArgoCD, artifact repositories).

  • Experience with OpenTelemetry, or with consolidating several overlapping monitoring tools into one coherent stack.

  • Experience testing event-driven workflows, and with managed streaming platforms such as MSK.

  • Experience with AWS Organizations and Control Tower, service control policies, or leading an account restructuring or consolidation.

  • Experience owning an identity migration or SSO consolidation (Auth0, Okta, Cognito, Keycloak, or similar).

  • Experience building internal developer tooling or CLIs, and improving parity between local, dev, and production environments.

  • Ability to read Java or C++ well enough to debug and modify services in those languages, even if you would not choose to write them.

  • Hands-on experience with AI developer tooling (Claude Code, MCP servers, agent workflows) or with running model infrastructure (e.g. AWS Bedrock).

  • Interest in the energy industry and the clean energy transition.

**Please include a link to your GitHub account in your application (in the “links” section). Applications without a GitHub account will not be considered**
 
Please note that at this time, we do not sponsor visas or transfers for new hires. Voltus teammates need to be authorized to work from their home location (in the US or Canada, unless otherwise indicated on the role description).
 
Additionally, while Voltus is an all-remote workplace, we have limitations on where employees are able to work for regulatory and security reasons. We expect that Voltans are working primarily from their home country. Working while traveling to other countries must be approved as per our Global Remote Travel Policy.
 
At Voltus, we are proud to be an equal opportunity employer because we recognize that a diverse organization begins with a diverse candidate pool. This means we do not tolerate discrimination of any kind and are committed to providing equal employment opportunities regardless of your gender identity, race, nationality, religion, age, sexual orientation, veteran status, disability status, or marital status.

Similar Jobs

15 Days Ago
Remote
Michigan, USA
112K-239K Annually
Senior level
112K-239K Annually
Senior level
Fintech • Financial Services
Designs, provisions, deploys, and operates cloud infrastructure using IaC and container orchestration. Manages Terraform modules, EKS/Kubernetes, Helm charts, Istio, CI/CD pipelines (GitHub Actions), Secrets Manager, and IAM. Collaborates with engineering teams on secure, scalable deployments, participates in incident response and postmortems, conducts infrastructure code reviews, documents systems, and may participate in on-call rotations to support production reliability.
Top Skills: Aurora PostgresAWSAws Secrets ManagerCi/CdCloudFormationCloudwatchDynatraceEksGithub ActionsHelmIamIstioKubernetesLambdaNode.jsRdsSplunkTerraformTypescript
22 Days Ago
Easy Apply
Remote
United States
Easy Apply
173K-255K Annually
Senior level
173K-255K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead delivery for the Batch Infrastructure team: design, build, and operate reliable, scalable compute platforms for scheduled and on-demand batch workloads. Collaborate with product/design/analytics, define technical plans, ensure availability (monitoring/on-call), set code and design standards, and mentor engineers to improve delivery and quality.
Top Skills: AirflowAWSFlinkFlyteKotlinKubernetesLuigiMySQLPrefectPythonSparkTemporal
7 Hours Ago
Remote or Hybrid
CA, USA
160K-322K Annually
Senior level
160K-322K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Deploys and validates complete AI infrastructure software stacks on multi-node GPU systems, then converts implementations into technical documentation, automation, demos, training, and reference architectures. Tests prerelease software, evaluates interoperability and operational resilience, supports partners and field teams, collaborates with engineering and open-source communities, and recommends product improvements based on customer feedback. The role presents solutions through briefings, workshops, webinars, events, and internal training, with some travel required.
Top Skills: Ai InferenceAi TrainingAPIsBare-Metal ProvisioningBluefield DpuCertificate ManagementCi/CdCloud-NativeConfiguration ManagementContainersDgx CloudDocaEthernetGitopsGpu SystemsHelmHpcIdentity ManagementInfinibandInfrastructure As CodeKubernetesLinuxMulti-TenancyNvidia Ai EnterpriseNvidia DgxNvidia DsxObservabilityPythonSecrets ManagementShell ScriptingSlurmStorageTelemetry

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account