We are seeking a highly skilled Lead Site Reliability Engineer (AI & Cloud Operations) to drive the reliability, scalability, automation, and operational excellence of our cloud-native platforms and AI-powered solutions. This role will serve as a technical leader responsible for building resilient infrastructure, implementing modern DevOps and SRE practices, and enabling enterprise AI capabilities through automation, observability, and operational intelligence.
The ideal candidate combines deep expertise in AWS cloud technologies, Kubernetes, infrastructure automation, CI/CD, and incident management with hands-on experience supporting AI/ML and Generative AI platforms. This individual will partner closely with software engineering, data engineering, machine learning, security, and product teams to establish highly available systems, streamline deployments, optimize platform performance, and accelerate innovation through AI-driven operations.
- Lead the design, implementation, and continuous improvement of Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient cloud platforms.
- Architect, deploy, and support AWS-based infrastructure and services, including containerized and serverless environments.
- Build, maintain, and optimize CI/CD pipelines and Infrastructure as Code (IaC) solutions to accelerate and standardize deployments.
- Develop automation solutions, operational tooling, and self-healing capabilities using Python, Shell, and modern DevOps technologies.
- Manage Kubernetes and container platforms, ensuring performance, scalability, and operational stability.
- Establish and enhance observability through monitoring, logging, alerting, and performance management tools to proactively identify and resolve issues.
- Lead incident response, root cause analysis, problem management, and service reliability improvement initiatives.
- Partner with engineering, data, AI/ML, and security teams to support enterprise applications, analytics platforms, and cloud-native solutions.
- Design and implement AI-driven operational capabilities, including intelligent monitoring, automated remediation, predictive analytics, and chatbot-enabled support workflows.
- Support MLOps and AI platform operations, including model deployment, monitoring, governance, and lifecycle management.
- Define and track reliability metrics, service-level objectives (SLOs), and operational KPIs to drive continuous improvement.
Mentor and provide technical leadership to engineering teams while promoting best practices in reliability, automation, cloud operations, and AI-enabled innovation.
Qualifications
- 5+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Operations.
- Strong hands-on experience with AWS services including EC2, EKS, ECS, Lambda, S3, RDS, IAM, CloudWatch, and VPC.
- Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, or similar platforms.
- Strong scripting and automation skills using Python, Shell, or similar languages.
- Experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
- Expertise in containerization and orchestration technologies (Docker, Kubernetes).
- Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK Stack.
- Understanding of analytics platforms, data pipelines, and operational data analysis.
- Strong troubleshooting, problem-solving, and incident management skills.
- AI & Automation Experience
- Experience implementing AI/ML or Generative AI solutions within enterprise environments.
- Familiarity with AI platforms such as Azure OpenAI, AWS Bedrock, Amazon SageMaker, OpenAI APIs, LangChain, or NVIDIA AI ecosystem.
- Experience building AI-assisted operational workflows, chatbots, intelligent monitoring, predictive analytics, or automated remediation solutions.
- Understanding of MLOps concepts, model deployment, monitoring, and governance.
WHAT WE BELIEVE
At Perficient, we promise to challenge, champion, and celebrate our people. You will experience a unique and collaborative culture that values every voice. Join our team, and you’ll become part of something truly special.
We believe in developing a workforce that is as diverse and inclusive as the clients we work with. We’re committed to actively listening, learning, and acting to further advance our organization, our communities, and our future leaders… and we’re not done yet.
Perficient, Inc. proudly provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, genetic information, marital status, amnesty, or status as a protected veteran in accordance with applicable federal, state and local laws. Perficient, Inc. complies with applicable state and local laws governing non-discrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. Perficient, Inc. expressly prohibits any form of unlawful employee harassment based on race, color, religion, gender, sexual orientation, national origin, age, genetic information, disability, or covered veterans. Improper interference with the ability of Perficient, Inc. employees to perform their expected job duties is absolutely not tolerated.
Disability Accommodations:
Perficient is committed to providing a barrier-free employment process with reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or accommodation due to a disability, please contact us.
Applications will be accepted until the position is filled or the posting is removed.
The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $111,300 to $144,600. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.
Disclaimer: The above statements are not intended to be a complete statement of job content, rather to act as a guide to the essential functions performed by the employee assigned to this classification. Management retains the discretion to add or change the duties of the position at any time.
#LI-RS1
About UsPerficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world’s most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at perficient.com.
Perficient New York, New York, USA Office
111 Broadway, New York, NY, United States, 10006
Similar Jobs
What you need to know about the NYC Tech Scene
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory


