Hitachi Logo

Hitachi

SRE/DevOps Engineer - 67533

Reposted 4 Hours Ago
Be an Early Applicant
In-Office
Toronto, ON
Mid level
In-Office
Toronto, ON
Mid level
Monitor and support enterprise applications across cloud and on-premises environments. Perform incident triage, analyze logs and metrics, execute operational runbooks, validate Kubernetes health, troubleshoot Linux and networking issues, and escalate complex incidents. Support cloud infrastructure, API gateways, WAFs, Kafka, and databases while maintaining documentation and communicating with stakeholders. Contribute to application onboarding, automation, observability, production support, and continuous improvements to system reliability.
The summary above was generated by AI

Function

Cloud & Data EngineeringOur Company

We’re Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world’s potential. We’re people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what’s now to what’s next. We make it happen through the power of acceleration.

Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don’t expect you to ‘fit’ every requirement – your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.

Job descriptionMeet Our Team

Join our Site Reliability Engineering (SRE) Operations team, where reliability, automation, and operational excellence are at the heart of everything we do. We ensure the stability, availability, and performance of enterprise applications running across modern cloud-native and hybrid platforms, including Kubernetes, APIs, cloud services, databases, Kafka, and API gateways.

As an L1 SRE Operations Engineer, you'll be the first line of defense, monitoring production environments, responding to alerts, executing operational runbooks, and partnering with senior engineers to maintain highly available and resilient platforms. This is an excellent opportunity for professionals looking to build hands-on experience in cloud operations, DevOps, and Site Reliability Engineering.

What You'll Be Doing
  • Monitor enterprise applications, infrastructure, dashboards, logs, and alerts across cloud and on-premises environments.
  • Perform first-level incident triage by analyzing alerts, collecting logs and metrics, and determining whether issues are application or platform related.
  • Execute standardized operational runbooks for incident resolution, deployments, maintenance activities, and routine operational tasks.
  • Monitor and support Kubernetes environments by validating pod health, deployments, namespaces, logs, and service endpoints.
  • Troubleshoot infrastructure and application issues using Linux utilities, networking tools, and monitoring platforms.
  • Escalate complex incidents to L2/L3 engineering teams with complete diagnostic information to accelerate resolution.
  • Support API gateways, web application firewalls (WAF), Kafka platforms, databases, and cloud infrastructure across AWS, Azure, and GCP.
  • Maintain accurate incident documentation, operational records, and knowledge base updates while identifying opportunities to improve runbooks and automation.
  • Collaborate with development, platform engineering, and infrastructure teams during incident response and production support.
  • Assist with onboarding new applications into the operational support framework while ensuring monitoring, alerting, and operational readiness.
  • Contribute to continuous improvement by identifying repetitive manual activities suitable for automation.
  • Provide timely and professional communication to stakeholders during production incidents and operational events.
What You'll Bring to the TeamRequired Qualifications
  • 2–5 years of experience in IT Operations, NOC, SRE, DevOps, or Infrastructure Support.
  • Working knowledge of Kubernetes administration and day-to-day cluster operations.
  • Good understanding of Linux administration and command-line troubleshooting.
  • Familiarity with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
  • Experience with observability and monitoring tools such as Prometheus, Grafana, Splunk, ELK Stack, Datadog, Argos, or AIOps platforms.
  • Ability to execute operational runbooks and follow structured incident response procedures.
  • Experience using Kubernetes CLI (kubectl) to verify pod health, deployments, namespaces, and application logs.
  • Basic scripting knowledge in Python, Bash, or PowerShell for operational automation.
  • Understanding of networking fundamentals including DNS, HTTP/HTTPS, TCP/IP, firewalls, WAF, proxies, connectivity troubleshooting, and diagnostic tools such as ping, curl, netstat, and traceroute.
  • Strong analytical and troubleshooting skills using structured problem-solving techniques such as 5 Whys and Fishbone Analysis.
  • Excellent documentation, communication, and stakeholder management skills.
Preferred Qualifications
  • Experience working with API gateways such as Apigee or Gloo API Gateway.
  • Basic knowledge of SQL and NoSQL databases with the ability to validate database connectivity.
  • Familiarity with messaging platforms such as Apache Kafka.
  • Experience with ITSM and incident management tools including ServiceNow, Jira, xMatters, or similar platforms.
  • Exposure to automation and self-service operations initiatives.
  • Experience using AI-assisted operational tools or chatbots for runbook search, log summarization, and incident analysis.
  • Understanding of cloud-native application architectures, CI/CD pipelines, and production support best practices.
  • Passion for continuous learning, operational excellence, and improving system reliability through automation.
About us

We’re a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.

Fostering innovation through diverse perspectives

Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.

We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.

How we look after you

We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We’re also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We’re always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you’ll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.

We’re proud to say we’re an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.

Hitachi Monroe, New Jersey, USA Office

Monroe, United States

Similar Jobs

57 Minutes Ago
Hybrid
96K-143K Annually
Senior level
96K-143K Annually
Senior level
Gaming
Build and own large-scale data ingestion services and SDKs, translate business requirements into technical designs, improve performance and reliability, provide production support, and drive engineering best practices for near real-time analytics pipelines.
Top Skills: AWSCi/CdDevOpsGoJavaNoSQLPythonRedshiftRestful ApisSQL
An Hour Ago
Remote or Hybrid
Junior
Junior
HR Tech • Information Technology • Professional Services • Sales • Software
Manage the full sales cycle for small and midsize businesses, from prospecting and pipeline generation through product demonstrations, relationship nurturing, forecasting, and closing. The role targets companies with 10–250 employees, engages key decision-makers, maintains sales metrics in Salesforce, and develops networks with influencers, consultants, and partners. This hybrid position requires at least three days per week in the Toronto office.
Top Skills: HrisSaaSSalesforce CRM
An Hour Ago
Remote or Hybrid
95K-119K Annually
Mid level
95K-119K Annually
Mid level
HR Tech • Information Technology • Professional Services • Sales • Software
Own the full sales cycle for mid-market HR technology customers across Canada, from prospecting and pipeline generation through product demonstrations, relationship development, forecasting, negotiation, and closing. Collaborate with Sales Engineers and Business Development representatives, target decision-makers, manage sales metrics in Salesforce, and represent HiBob at trade shows, events, and conferences. The role requires bilingual French and English communication and focuses on companies with 250–1,000 employees.
Top Skills: HrisSaaSSalesforce CRM

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account