Support reliability, monitoring, and modernization of Kubernetes-based microservices platforms. Build Datadog observability solutions including dashboards, alerts, APM, metrics, logs, and tracing; integrate tooling with AWS and CI/CD pipelines; automate operational tasks with Python; manage agents, integrations, API keys, and access controls; and lead platform maintenance and improvements for reliability, scalability, and performance.
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a Senior Site Reliability Engineer to support platform reliability, monitoring, and modernization across Kubernetes-based microservices environments with a strong observability focus. You will build and maintain Datadog solutions including dashboards, alerts, APM, metrics, logging, and tracing, integrate observability tooling into AWS and CI/CD pipelines, and automate monitoring and operational tasks using Python. The role blends software engineering (60–70%) with site reliability engineering (30–40%) and requires JST timezone overlap.
WHAT YOU WILL DO
- Support platform reliability, monitoring, and continuous improvement across internal systems.
- Work in Kubernetes-based environments.
- Build and maintain observability solutions, with a focus on Datadog.
- Configure dashboards, alerts, APM, metrics, logging, and tracing.
- Monitor containerized and microservices-based applications.
- Integrate observability tools into AWS environments.
- Integrate observability into CI/CD pipelines.
- Automate monitoring and operational tasks using scripting (Python preferred).
- Install and configure Datadog agents and integrations.
- Manage API keys and secure configurations.
- Manage user roles and access controls within observability platforms.
- Lead maintenance efforts and platform improvements while driving reliability, scalability, and performance.
MUST HAVES
- Strong proficiency in Python, JavaScript (Node.js), or Java.
- Hands-on experience with API integrations (designing, consuming, and integrating).
- Strong experience working in Kubernetes environments (deployment, operations, monitoring).
- Experience with Datadog (preferred) or similar tools (Prometheus, Grafana).
- Ability to configure dashboards, alerts, and APM (tracing, metrics, logging).
- Experience monitoring containerized/microservices architectures.
- Hands-on experience with AWS.
- Experience integrating observability tools into cloud environments.
- Experience integrating observability into CI/CD pipelines.
- Ability to automate monitoring and operational tasks using scripting (Python preferred).
- Upper-intermediate English level.
NICE TO HAVES
- Experience owning and operating an internal engineering platform.
- Demonstrated ownership of reliability, scalability, and performance.
- Proven ability to proactively lead maintenance efforts and platform improvements (not just reactive support).
- Familiarity with Golang.
- Experience with additional observability tools such as New Relic, Dynatrace, Elastic, or Splunk Observability.
PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location
Similar Jobs
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Support regulatory liquidity reporting by applying predefined rules and mappings, performing variance analysis and reconciliations, documenting business/functional requirements, coordinating BA/IT/testing streams, and contributing to governance, controls and audit readiness within regulatory change programmes.
Healthtech • Social Impact • Telehealth
As a Care Coordinator, you will guide older adults and their caregivers through the intake process, providing support and emotional guidance during their journey into care.
Top Skills:
CrmsEhr PlatformsGoogle Workspace
Digital Media • Information Technology • News + Entertainment
Lead design and implementation of AI Ops platform and cloud-based software solutions (AWS/GCP). Develop and maintain code (Golang, Python, Java), IaC, LLM/agent integrations, APIs/GraphQL, and automation. Mentor developers, create technical documentation and architecture diagrams, collaborate with QA and stakeholders, troubleshoot automation/operations issues, and drive continuous improvement.
Top Skills:
Agent FrameworksAgileAi OpsAmazon SqsAutomated Testing FrameworksAWSAws BedrockAws Step FunctionsEsb PlatformsEvent ManagementGitGithub ActionsGoGCPGraph DatabasesGraphQLInfrastructure As Code (Iac)ItsmJavaLlmsLow-Code/No-Code PlatformsMongoDBPrompt EngineeringPythonRRest Apis
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.


.png)
