Support and improve production reliability, scalability, and performance of Mastercard applications. Implement observability, automate operational tasks, assist deployments and monitoring, triage incidents, perform root-cause analysis, and collaborate with developers to embed reliability and operational best practices.
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Site Reliability Engineer I
Job Description Summary
Job Title
Site Reliability Engineer I
Who is Mastercard?
At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team - one that makes better decisions, drives innovation, and delivers better business results.
About the Role
The Business Operations team is seeking a highly motivated and experienced Site Reliability Engineer I (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor.
The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.
We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications.
Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle.
As part of the Business Operations team, you will: • Follow clearly defined guidelines to complete routine tasks within the Site Reliability Engineering area by applying basic/theoretical knowledge and area best practices. • Support day-to-day system maintenance tasks and help gather requirements for technical solutions. • Identify automation that can be used and scripting efforts with guidance from senior engineers. • Learn troubleshooting techniques and assist in resolving basic system issues. • Gain familiarity with system applications and infrastructure, maintenance standards, and best practices used within the organization. • Support deployment and monitoring activities, gaining experience in operational aspects of software delivery. • Seek guidance from senior team members to improve technical skills and understanding of project requirements • Troubleshoot and resolve basic to moderate system issues, escalating more complex problems as needed. • Participate in team meetings and knowledge-sharing sessions to enhance technical understanding.
Role qualifications:
The ideal candidate will understand core concepts, terminology, and principles of the following skills, apply them in routine situations with support, and build confidence through learning and practice. They may require guidance in moderately complex or unfamiliar scenarios.
• Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement.• Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar.• Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability.• Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency.• Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand• DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations.• Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.• Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage.• IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs.• Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Site Reliability Engineer I
Job Description Summary
Job Title
Site Reliability Engineer I
Who is Mastercard?
At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team - one that makes better decisions, drives innovation, and delivers better business results.
About the Role
The Business Operations team is seeking a highly motivated and experienced Site Reliability Engineer I (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor.
The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.
We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications.
Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle.
As part of the Business Operations team, you will: • Follow clearly defined guidelines to complete routine tasks within the Site Reliability Engineering area by applying basic/theoretical knowledge and area best practices. • Support day-to-day system maintenance tasks and help gather requirements for technical solutions. • Identify automation that can be used and scripting efforts with guidance from senior engineers. • Learn troubleshooting techniques and assist in resolving basic system issues. • Gain familiarity with system applications and infrastructure, maintenance standards, and best practices used within the organization. • Support deployment and monitoring activities, gaining experience in operational aspects of software delivery. • Seek guidance from senior team members to improve technical skills and understanding of project requirements • Troubleshoot and resolve basic to moderate system issues, escalating more complex problems as needed. • Participate in team meetings and knowledge-sharing sessions to enhance technical understanding.
Role qualifications:
The ideal candidate will understand core concepts, terminology, and principles of the following skills, apply them in routine situations with support, and build confidence through learning and practice. They may require guidance in moderately complex or unfamiliar scenarios.
• Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement.• Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar.• Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability.• Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency.• Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand• DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations.• Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.• Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage.• IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs.• Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard's security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Mastercard Gurugram, Haryana, IND Office
Mehrauli Gurgaon Road, Gurugram, Gurugram, India, 122002
Similar Jobs at Mastercard
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead ML engineering for Operational Intelligence: design and productionize multimodal transformer and agentic systems, build scalable RAG/Graph-RAG and LLMOps/MLOps pipelines, implement inference services and memory/state subsystems, apply traditional and deep learning methods, enforce Responsible AI, and drive research-to-production innovation.
Top Skills:
AutogenAWSAws NeptuneAzureBertClipCrewaiDatabricksFeature StoreGCPGraph-RagHugging FaceKnowledge GraphLanggraphLlavaMlflowNeo4JPower BIPythonPyTorchRagSQLT5TableauTensorFlowVector StoreWhisper
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead a team to architect and deliver production-grade generative AI: multi-agent systems, multimodal transformers, RAG/Graph-RAG, LLM fine-tuning and LLMOps on cloud platforms. Build Python inference services, manage agent state/memory, implement responsible AI and governance, and integrate traditional ML for interpretable hybrid systems while researching and productionizing frontier models.
Top Skills:
AnthropicAutogenAws BedrockAws NeptuneAws SagemakerAzure OpenaiBertClipCrewaiDatabricksFalconFastapiFeature StoresGcp Vertex AiGeminiGpt-4OHugging FaceLangchainLanggraphLlamaLlavaLoraMistralMlflowNeo4JOpenaiOpensearchPeftPgvectorPineconePythonPyTorchRlhfSQLT5TensorFlowWeights & BiasesWhisper
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Develop and maintain scalable, highly available enterprise applications and portals using Java/Spring Boot. Design cloud-native microservices, implement REST APIs, ensure operational excellence through monitoring, automation, incident management, capacity planning, and collaborate with global cross-functional teams to meet security, compliance, and performance goals.
Top Skills:
Ai Pair ProgrammingAWSAzureDevOpsHibernateItsmJava 8Java J2EeMicroservicesNoSQLOpenapi SpecsPostgresRest ApisSnowflakeSpring Boot 3.X
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.

