Ensure availability, performance, scalability, and reliability of production systems. Monitor systems with observability tools, define SLIs/SLOs, lead incident management and RCAs, automate operations via IaC, and support CI/CD pipeline stability and improvements.
- Job Summary
- The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies - Technical Skills
- Cloud Platforms: Azure
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies
- Technical Skills
- Cloud Platforms: AWS / Azure / GCP
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. Explore Life at Zensar and join us to Grow. Own. Achieve. Learn. to be the best version of yourself.
We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Similar Jobs
Edtech • Software
Leads cloud platform architecture, reliability, modernization, automation, cost optimization, and operational excellence across Azure-based SaaS infrastructure with some AWS. Manages and develops cloud and database engineers, establishes observability and incident-response practices, improves CI/CD, supports highly available systems, and partners with engineering, product, and security teams. The role remains hands-on with Infrastructure-as-Code while overseeing hiring, performance management, capacity planning, and technical strategy.
Top Skills:
Arm TemplatesAWSAzure Kubernetes ServiceBicepC#Ci/CdFinopsInfrastructure-As-CodeAzurePowershellPythonSreTerraform
Marketing Tech • Software
Leads Carousell Group’s platform engineering, infrastructure, SRE, and data engineering organizations. The role sets technical direction for scalable distributed systems, reliability, observability, developer productivity, AI-native engineering adoption, cloud cost optimization, and infrastructure strategy. It manages engineering leaders and distributed teams, develops senior technical talent, drives multi-year initiatives, and partners with executives across markets to balance platform quality, reliability, velocity, and cost.
Top Skills:
Agentic EngineeringAi-Native EngineeringCloud InfrastructureData EngineeringDeveloper ToolingDistributed SystemsInfrastructureNetworkingObservabilityPlatform EngineeringSre
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills:
Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.



