About Cellebrite:
Cellebrite’s (Nasdaq: CLBT) mission is to enable its global customers to protect and save lives by enhancing digital investigations and intelligence gathering to accelerate justice in communities around the world. Cellebrite’s AI-powered Digital Investigation Platform enables customers to lawfully access, collect, analyze and share digital evidence in legally sanctioned investigations while preserving data privacy. Thousands of public safety organizations, intelligence agencies and businesses rely on Cellebrite’s digital forensic and investigative solutions—available via cloud, on-premises and hybrid deployments—to close cases faster and safeguard communities.
To learn more, visit us at www.cellebrite.com, https://investors.cellebrite.com/investors and find us on social media @Cellebrite.
What is your mission?
As a Site Reliability Engineer, you’ll be a key technical owner of production health for Cellebrite’s cloud platform the platform investigators and public safety agencies depend on every day. You’ll work closely with our TCS support team and R&D to triage, investigate, and resolve production issues quickly and thoroughly, while helping modernize how we monitor the platform, roll out changes safely, and respond when things break. The focus of this role is keeping production healthy and continuously raising the bar on how we operate it — not building new infrastructure from scratch. Your work directly supports Cellebrite’s mission to protect lives, accelerate justice, and preserve data privacy.
Responsibilities:
Incident Response & Production Health
- Own the full incident lifecycle detect, triage by severity/impact, investigate, and drive to resolution engaging TCS, R&D, and DevOps as needed.
- Act as a technical responder during major incidents and grow into an incident-commander role coordinating the response across teams in real time.
- Lead blameless post-incident reviews and make sure corrective actions actually get implemented, not just documented.
- Track recurring issues and support-ticket trends; distinguish patterns that need a permanent fix from one-off noise, and route the former to R&D.
AI-Driven Monitoring & Observability
- Evolve production monitoring toward AI-assisted operations anomaly detection, cross-signal correlation across logs/metrics/traces, and LLM-based triage assistants that cut time-to-diagnosis.
- Continuously tune dashboards, alert thresholds, and routing so real issues surface fast and noise doesn’t drown them out.
- Evaluate and pilot AI/ML-based observability tooling, and champion adoption of tools like Copilot or log-analysis assistants across the team.
Rolling Upgrades & Change Management
- Plan and execute rolling upgrades, version updates, and patches to production services and infrastructure with zero or minimal downtime.
- Partner with R&D and DevOps to define safe rollout and rollback strategies for deployments, and validate system health post-upgrade.
- Participate in an on-call rotation to help maintain production uptime SLOs, with upgrade windows and rollback readiness built into the runbooks you own.
Cross-Team Enablement
- Partner with TCS on production tickets, providing deeper technical investigation when issues exceed their level; partner with R&D on code-level fixes, deployments, and architectural input.
- Own and continuously improve runbooks and the known-issues knowledge base so TCS resolves more independently over time.
- Maintain clear, current documentation of production architecture, known issues, and resolution paths for both TCS and R&D.
Requirements:
- 3–5 years of experience in a production support, SRE, or operations role.
- Solid working experience with AWS (troubleshooting and operating existing infrastructure; deep infra design experience is not required).
- Experience with monitoring/observability tools (e.g., Datadog, CloudWatch, Grafana) able to read dashboards, tune alerts, and investigate from logs/metrics/traces.
- Working knowledge of Linux system administration and basic networking troubleshooting.
- Familiarity with containerized environments (Kubernetes) — enough to check pod health, logs, and restart/rollback safely.
- Comfortable with scripting (Python or Bash) to automate repetitive troubleshooting or reporting tasks.
- Experience driving rolling upgrades, patching, or zero-downtime deployment processes in a production environment.
- Experience collaborating across support and engineering teams — comfortable working with an outsourced support team (e.g., TCS) on one side and R&D engineers on the other, translating between the two.
- Strong communication skills in English, written and verbal — clear handoffs and documentation matter as much as fixing things.
- Nice to have: exposure to Terraform or other IaC (read/troubleshoot, not necessarily author), basic CI/CD familiarity, and genuine interest in applying AI/ML to monitoring, alerting, and incident triage — not just using AI coding assistants.
Cellebrite Parsippany, New Jersey, USA Office
7 Campus Dr, Parsippany, NJ, United States, 07054
Similar Jobs
What you need to know about the NYC Tech Scene
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory



