Designs and operates cloud-native infrastructure, automated CI/CD pipelines, infrastructure-as-code, Kubernetes platforms, observability systems, and secure configuration management. The role focuses on high availability, disaster recovery, cluster scaling, fault tolerance, incident response, and root-cause analysis. Candidates need substantial DevOps or systems engineering experience, advanced Kubernetes and Terraform expertise, Linux and scripting knowledge, and a mandatory AWS, Azure, or Kubernetes certification.
This is a remote position.
DevOps / Site Reliability Engineer (SRE)
Job Details
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 5 to 9 years
- Relevant Experience Required: 4+ years of dedicated experience in infrastructure automation, cloud orchestration, and CI/CD pipelines
- Mandatory Certification: AWS Certified DevOps Engineer - Professional, Microsoft Certified: DevOps Engineer Expert, or Certified Kubernetes Administrator (CKA)
Job Summary
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, automate, and scale our cloud-native infrastructure pipelines. The ideal candidate will bridge the gap between software development and systems operations, building highly available deployment pipelines, writing infrastructure-as-code (IaC), and optimizing cluster scaling to maximize application uptime, system reliability, and performance.
Key Responsibilities
- Design, build, and optimize automated CI/CD pipelines across cloud platforms using industry-standard automation servers (e.g., GitHub Actions, GitLab CI, Jenkins, ArgoCD).
- Architect and manage infrastructure-as-code (IaC) templates using Terraform or OpenTofu to provision secure, modular, and repeatable multi-environment architectures.
- Orchestrate containerized production workloads, configuring cluster scaling, service meshes, network routing policies, and deployment strategies on Kubernetes (EKS/AKS/GKE).
- Implement automated monitoring, logging, and alerting systems utilizing observability tools (e.g., Prometheus, Grafana, Datadog, ELK stack) to actively track platform performance metrics.
- Drive system high-availability and fault tolerance efforts, designing disaster recovery plans, automated load balancing parameters, and self-healing cluster scripts.
- Manage centralized configuration and secret management systems, securely vaulting database credentials, API tokens, and certificate profiles (e.g., HashiCorp Vault, AWS Secrets Manager).
- Participate in on-call rotations and lead incident root-cause analysis (RCA), systematically diagnosing runtime infrastructure failures, performance bottlenecks, and resource leaks.
Requirements
- 5 to 9 years of core systems engineering or software development experience, with 4+ dedicated years actively designing, building, and operating cloud-native production platforms.
- Strong technical mastery of Kubernetes cluster administration, Terraform automation layouts, Linux system internals, shell scripting (Bash, Python, or Go), and network protocols.
- Deep structural understanding of microservices design, caching mechanics, database scaling limits, and cloud provider API governance.
- Mandatory certification: AWS DevOps Professional, Azure DevOps Expert, or CKA.
Preferred Qualifications
- Prior experience implementing DevSecOps controls (e.g., integrating SAST/DAST tools directly into container build phases).
- Experience with GitOps methodologies and progressive delivery mechanisms (e.g., Canary or Blue/Green deployments using Flagger or Istio).
Benefits
- 12+ years of total IT software engineering or operational management background, with 6+ dedicated years acting as a CISO, Director of Security, or Principal Enterprise GRC Advisor.
- Strong visionary mastery of modern security trends, zero-trust target states, risk calculation paradigms, and multi-cloud information landscape parameters.
- Deep communication execution skills, with a proven history of negotiating security budgets, steering board panels, and handling high-pressure public communication events.
- Mandatory certification: CISM or CISSP.
Preferred Qualifications
- Certified in the Governance of Enterprise IT (CGEIT) or Certified in Risk and Information Systems Control (CRISC) credential.
- Prior experience steering complex post-merger information platform integrations or stabilizing security posture profiles during major corporate equity restructurings.
Similar Jobs
Information Technology
Develop Python applications, APIs, automation tools, and platform capabilities while supporting CI/CD, Docker, Kubernetes, cloud and on-premise deployments. Implement observability, monitoring, alerting, and SRE practices including SLIs, SLOs, error budgets, incident response, disaster recovery, and root cause analysis. Troubleshoot production systems, automate recurring operational work, maintain documentation, and collaborate across engineering, security, platform, product, operations, and external teams.
Top Skills:
AWSAzureAzure DevopsClaude CodeConfluenceDnsDockerGCPGitGithub ActionsGithub CopilotGitlab CiGrafanaHttp/HttpsJenkinsJIRAJSONKubernetesLinux/UnixPower BIPrometheusPythonRest ApisServicenowTcp/IpTerraformWindows
Fintech • Analytics
Join a talent community for Site Reliability/DevOps roles focused on designing, building, and operating resilient cloud infrastructure. Responsibilities include architecting Kubernetes-based platforms, implementing IaC and CI/CD (Terraform, Helm, GitHub Actions, Azure DevOps, Jenkins), improving reliability via SLOs/SLIs, observability, chaos engineering and automation, collaborating with cross-functional teams, and mentoring junior engineers.
Top Skills:
Azure DevopsChaos EngineeringCi/CdDockerGithub ActionsHelmInfrastructure As CodeJenkinsKubernetesMicroservicesMonitoringObservabilityPerformance TestingTerraform
Information Technology
Design, operate, and automate AWS-based production and development infrastructure for eCommerce/enterprise platforms. Implement CI/CD, infrastructure-as-code, observability, security, disaster recovery, and large-scale automation. Support microservices, databases, caching, virtualization, and application servers while troubleshooting, performance tuning, and onboarding new tools.
Top Skills:
AkamaiAndroidApacheAppdynamicsAptitude/DpkgAwkAWSBashCassandraCdnChefDatadogDynatraceEc2Elastic CloudElkGeodnsGlobal Traffic ManagementGraphiteGroovyHadoopHaproxyHbaseHelmHypervisorIptablesJavaJavaScriptJbossJenkinsJettyJSONKeycloakKubernetesLdapLinuxMemcachedMicroservicesMongoDBMySQLNagiosNessusNetappNew RelicNfsNginxNmapNtpObjective-COktaOpen DirectoryOraclePerlPHPPuppetPythonRackspace CloudRedisRestRubyService MeshSoftlayerSplunkSsl/TlsTerraformTomcatVarnishVdiVirtualizationWeblogicXMLYum/Rpm
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.


.jpg)
