Cellebrite Logo

Cellebrite

Site Reliability Engineer

Posted An Hour Ago
Be an Early Applicant
Hybrid
New York City, NY, USA
Mid level
Hybrid
New York City, NY, USA
Mid level
Own production health for Cellebrite’s cloud platform through incident response, monitoring and observability improvements, rolling upgrades, change management, and on-call support. Investigate production issues, coordinate incident response, improve alerting and runbooks, collaborate with support and R&D teams, and apply AI/ML tools to anomaly detection and incident triage.
The summary above was generated by AI
Description

About Cellebrite: 

Cellebrite’s (Nasdaq: CLBT) mission is to enable its global customers to protect and save lives by enhancing digital investigations and intelligence gathering to accelerate justice in communities around the world. Cellebrite’s AI-powered Digital Investigation Platform enables customers to lawfully access, collect, analyze and share digital evidence in legally sanctioned investigations while preserving data privacy. Thousands of public safety organizations, intelligence agencies and businesses rely on Cellebrite’s digital forensic and investigative solutions—available via cloud, on-premises and hybrid deployments—to close cases faster and safeguard communities.

To learn more, visit us at www.cellebrite.com, https://investors.cellebrite.com/investors and find us on social media @Cellebrite. 


What is your mission? 

As a Site Reliability Engineer, you’ll be a key technical owner of production health for Cellebrite’s cloud platform the platform investigators and public safety agencies depend on every day. You’ll work closely with our TCS support team and R&D to triage, investigate, and resolve production issues quickly and thoroughly, while helping modernize how we monitor the platform, roll out changes safely, and respond when things break. The focus of this role is keeping production healthy and continuously raising the bar on how we operate it — not building new infrastructure from scratch. Your work directly supports Cellebrite’s mission to protect lives, accelerate justice, and preserve data privacy. 

Responsibilities: 

Incident Response & Production Health 

  • Own the full incident lifecycle detect, triage by severity/impact, investigate, and drive to resolution engaging TCS, R&D, and DevOps as needed. 
  • Act as a technical responder during major incidents and grow into an incident-commander role coordinating the response across teams in real time. 
  • Lead blameless post-incident reviews and make sure corrective actions actually get implemented, not just documented. 
  • Track recurring issues and support-ticket trends; distinguish patterns that need a permanent fix from one-off noise, and route the former to R&D. 

AI-Driven Monitoring & Observability 

  • Evolve production monitoring toward AI-assisted operations anomaly detection, cross-signal correlation across logs/metrics/traces, and LLM-based triage assistants that cut time-to-diagnosis. 
  • Continuously tune dashboards, alert thresholds, and routing so real issues surface fast and noise doesn’t drown them out. 
  • Evaluate and pilot AI/ML-based observability tooling, and champion adoption of tools like Copilot or log-analysis assistants across the team. 

Rolling Upgrades & Change Management 

  • Plan and execute rolling upgrades, version updates, and patches to production services and infrastructure with zero or minimal downtime. 
  • Partner with R&D and DevOps to define safe rollout and rollback strategies for deployments, and validate system health post-upgrade. 
  • Participate in an on-call rotation to help maintain production uptime SLOs, with upgrade windows and rollback readiness built into the runbooks you own. 

Cross-Team Enablement 

  • Partner with TCS on production tickets, providing deeper technical investigation when issues exceed their level; partner with R&D on code-level fixes, deployments, and architectural input. 
  • Own and continuously improve runbooks and the known-issues knowledge base so TCS resolves more independently over time. 
  • Maintain clear, current documentation of production architecture, known issues, and resolution paths for both TCS and R&D.
Requirements

Requirements: 

  • 3–5 years of experience in a production support, SRE, or operations role. 
  • Solid working experience with AWS (troubleshooting and operating existing infrastructure; deep infra design experience is not required). 
  • Experience with monitoring/observability tools (e.g., Datadog, CloudWatch, Grafana) able to read dashboards, tune alerts, and investigate from logs/metrics/traces. 
  • Working knowledge of Linux system administration and basic networking troubleshooting. 
  • Familiarity with containerized environments (Kubernetes) — enough to check pod health, logs, and restart/rollback safely. 
  • Comfortable with scripting (Python or Bash) to automate repetitive troubleshooting or reporting tasks. 
  • Experience driving rolling upgrades, patching, or zero-downtime deployment processes in a production environment. 
  • Experience collaborating across support and engineering teams — comfortable working with an outsourced support team (e.g., TCS) on one side and R&D engineers on the other, translating between the two. 
  • Strong communication skills in English, written and verbal — clear handoffs and documentation matter as much as fixing things. 
  • Nice to have: exposure to Terraform or other IaC (read/troubleshoot, not necessarily author), basic CI/CD familiarity, and genuine interest in applying AI/ML to monitoring, alerting, and incident triage — not just using AI coding assistants. 

Cellebrite Parsippany, New Jersey, USA Office

7 Campus Dr, Parsippany, NJ, United States, 07054

Similar Jobs

Yesterday
In-Office
New York, NY, USA
147K-234K Annually
Senior level
147K-234K Annually
Senior level
Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
Leads reliability, scalability, performance, and security for large-scale cloud systems. Designs AWS infrastructure with Terraform, automates CI/CD and operational workflows, establishes SLOs, manages incident response and disaster recovery, develops monitoring and observability solutions, and builds internal tools. Partners with engineering teams on architecture and reliability practices, conducts code reviews, mentors SREs, and supports compliance, vulnerability management, and security integration.
Top Skills: Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayChaos EngineeringCi/CdDastDatadogDistributed SystemsDockerEvent-Driven SystemsGitlabGitopsGrafanaIamInfrastructure As CodeJavaKubernetesLlmsMicroservicesNew RelicNode.jsOwasp Top 10PythonSastServerless ArchitecturesSplunkTerraform
12 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
13 Days Ago
Hybrid
2 Locations
151K-187K Annually
Senior level
151K-187K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Design and implement reliable, scalable IT infrastructure; automate processes; monitor systems; resolve incidents; conduct performance and load testing; optimize cloud environments; maintain storage architecture; and lead continuous improvement. The role requires troubleshooting complex system issues, developing automation solutions, managing incidents, collaborating across teams, mentoring others, and supporting secure business operations across AWS, Google Cloud, and Microsoft Azure.
Top Skills: AWSGoogle Cloud PlatformAzure

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account