Designs and evolves an Azure-based observability and data quality platform. Builds reliable, automated infrastructure using IaC, CI/CD, Python, and Shell scripting; integrates AKS, Kafka, Databricks, Flink, Fivetran, and Azure SQL Server; develops monitoring and dashboards with Grafana, Prometheus, Power BI, or Databricks; applies SRE and FinOps practices to improve reliability, performance, data quality, incident response, compliance, and cloud cost efficiency.
DevOps Engineer
12 months contract
India, Remote
Growth through diversity, equity, and inclusion. As an ethical business, we do what is right — including ensuring equal opportunities and fostering a safe, respectful workplace for each of us. We believe diversity fuels both personal and business growth. We're committed to building an inclusive community where all our people thrive regardless of their backgrounds, identities, or other personal characteristics.
What You’ll Be Doing:
- Platform Design, Development & Evolution: Architect, build, and continuously evolve the core M&O platform and services, leveraging modern technologies and best practices to provide comprehensive observability and data quality functions.
- Ensuring System Reliability & Performance: Actively maintain and enhance the stability, availability, and performance of critical applications and data infrastructure by integrating Site Reliability Engineering (SRE) principles directly into the platform's design and operation.
- Proactive Issue Detection & Resolution: Develop and integrate intelligent systems within the platform to proactively identify, diagnose, and trigger automated or semi-automated resolution for technical issues, performance bottlenecks, and anomalies across all operational and data systems.
- Advanced Data Quality Platform Implementation: Build and integrate capabilities within the M&O platform for continuously measuring, monitoring, and reporting on all critical data quality dimensions (Timeliness, Consistency, Completeness, Accuracy, Validity, Uniqueness) across diverse supply chain pipelines, including Warehousing, Transformation, SAP, Manufacturing, and other critical data sources.
- Centralized Insights & Dashboarding: Develop and manage a unified dashboarding interface (e.g., Grafana, Power BI, Databricks) within the platform to visualize key performance indicators, system health, operational metrics, financial insights (FinOps), and granular data quality metrics for various stakeholders.
- Automating Infrastructure & Operations: Drive the platform's automation capabilities through Infrastructure as Code (IaC), Continuous Integration/Continuous Delivery (CI/CD) pipelines, and extensive scripting (Python, Shell) for provisioning, deployment, and operational workflows.
- Managing Core Data & Cloud Technologies: Integrate and optimize essential technologies like Fivetran, AKS, Kafka, Azure SQL Server, Databricks, and Flink into the M&O platform, ensuring seamless operation and data flow within the Azure cloud environment.
- Optimizing Cloud Resources & Costs: Embed FinOps practices and reporting into the platform to monitor cloud resource utilization and spending, identifying opportunities for cost reduction and ensuring efficient allocation of infrastructure investments.
- Fostering a Culture of Continuous Improvement: Leverage platform data and SRE practices to continuously analyze operational incidents, performance trends, and data quality issues, driving ongoing enhancements to both the platform and the systems it monitors.
- Enabling Data-Driven Decision Making & Compliance: Ensure the M&O platform provides accurate, timely data and insights, derived from both monitoring and high-quality data across all domains, to support informed business decisions, while also guaranteeing compliance with regulatory requirements for data integrity and auditability.
What We’re Looking For:
- 5+ years of experience in DevOps, Cloud Engineering, or a similar role.
- Strong hands-on experience with Microsoft Azure and cloud infrastructure.
- Experience with Azure Kubernetes Service (AKS) and containerized environments.
- Hands-on experience with CI/CD pipelines and Infrastructure as Code (IaC).
- Strong scripting and automation skills using Python and Shell scripting.
- Experience with Grafana, Prometheus, and observability and monitoring solutions.
- Strong understanding and practical experience with Site Reliability Engineering (SRE) principles.
- Experience with Kafka, Azure SQL Server, Databricks, Flink, and Fivetran.
- Experience with infrastructure automation, system monitoring, performance optimization, and incident management.
Similar Jobs
AdTech • Digital Media • Information Technology • Marketing Tech • News + Entertainment • Social Media • Software
Own and improve cloud infrastructure, Kubernetes platforms, Infrastructure as Code, CI/CD, GitOps, observability, security, reliability, and cost efficiency. Build automation and self-service tooling, lead complex production incident investigations, improve resilience and disaster recovery, evaluate architecture, and mentor engineers. The role requires autonomous platform ownership, strong systems judgment, and thoughtful use of AI-assisted engineering tools.
Top Skills:
Amazon RedshiftArgocdAWSBashCdnCi/CdDatadogDnsDockerGithub ActionsGitopsInfrastructure As CodeKubernetesLinuxLoad BalancingMySQLOpentofuPythonRedisTerraformTlsWaf
Other
The DevOps Engineer will manage on-premises IT infrastructure, focusing on security, performance, and system availability. Responsibilities include network management, firewall configuration, system administration, and user support.
Top Skills:
DevOpsFirewallsInfrastructure ManagementLanNetwork SecurityServer AdministrationVirtualization TechnologiesVpnWan
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Leads industrial engineering for semiconductor assembly and test manufacturing, including capacity and resource planning, labor optimization, operational excellence, cost reduction, manufacturing analytics, and factory systems. Develops capacity models, staffing standards, dashboards, and automation initiatives while supporting budgets and capital planning. Builds and mentors an industrial engineering team and collaborates globally to standardize manufacturing practices and improve efficiency, utilization, cycle time, and cost performance.
Top Skills:
Advanced AnalyticsAIManufacturing SimulationMesExcelPower BISQLTableau
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.

.png)

.jpeg)