
Senior Enterprise Operations Engineer-2
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Senior Enterprise Operations Engineer-2Job Description:L2 Cloud Operations Engineer – Role Description
We are seeking a highly skilled L2 Cloud Operations Engineer to operate, scale, and continuously improve enterprise-grade multi-cloud platforms. This role combines expertise in cloud engineering, DevOps, SRE, automation, CI/CD troubleshooting, observability, security operations, and AIOps-driven incident management.
You will be responsible for ensuring complex incident resolution, platform reliability, DevOps enablement, Kubernetes operations, automation, and operational excellence across AWS, Azure, and hybrid environments.
Key Responsibilities
1. Cloud Operations & Platform Reliability (SRE Focus)
- Take ownership of production cloud platform stability, availability, performance, and reliability.
- Lead complex incident triage, deep troubleshooting, and resolution across cloud and container platforms.
- Implement SRE best practices such as SLIs, SLOs, SLAs, error budgets, and resilience engineering.
- Drive root cause analysis (RCA) and eliminate recurring failures through engineering solutions.
- Provide advanced troubleshooting for CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, Azure DevOps, AWS CodePipeline).
- Diagnose and resolve pipeline failures related to infrastructure, container builds, artifact storage, security scanning, and deployment orchestration.
- Enable self-service pipelines, progressive delivery, canary deployments, and rollback automation.
- Collaborate with engineering teams to improve deployment velocity, quality, and reliability.
- Operate and support large-scale AWS, Azure, and hybrid cloud platforms.
- Perform advanced Kubernetes operations including:
- Cluster management
- Node scaling
- Networking troubleshooting
- Pod failures
- Storage & ingress debugging
- Manage cloud networking, compute, storage, IAM, load balancing, and security services.
- Operate and enhance enterprise observability platforms across metrics, logs, traces, and events.
- Implement AIOps-driven alert correlation, noise reduction, anomaly detection, and predictive analytics.
- Build automated remediation workflows and self-healing systems.
- Drive MTTR reduction through automation-first operations.
- Provide cloud security incident triage and remediation.
- Implement DevSecOps automation for vulnerability scanning, compliance checks, and policy enforcement.
- Support cloud security posture management (CSPM), IAM governance, secrets management, and certificate lifecycle management.
- Build and maintain infrastructure as code (IaC) using Terraform, CloudFormation, ARM, and Bicep.
- Develop automation scripts and tools using Python, Bash, PowerShell, and Go.
- Implement event-driven automation and auto-healing frameworks.
- Serve as the technical escalation point for L1 teams.
- Lead incident management, problem management, and change governance using ITIL best practices.
- Maintain high-quality SOPs, runbooks, and knowledge articles.
Candidates must hold at least two of the following certifications:
- Certified Kubernetes Administrator (CKA)
- AWS Certified Solutions Architect – Associate
- AWS DevOps Engineer – Professional / SysOps Administrator – Associate
- Microsoft Azure Solutions Architect Expert
- Azure DevOps Engineer Expert
Required Technical Skills
Cloud & Platform
- AWS: EC2, EKS, VPC, IAM, ALB/NLB, S3, RDS, CloudWatch, Lambda
- Azure: AKS, VNET, IAM (Entra), Load Balancer, App Services, Monitor
- Multi-cloud architecture & operations
- Kubernetes (EKS, AKS, OpenShift)
- Helm, Kustomize, ArgoCD, Flux
- Docker, container build pipelines
- Service mesh (Istio / Linkerd – good to have)
- Jenkins, GitHub Actions, GitLab CI, Azure DevOps
- Terraform, CloudFormation, ARM/Bicep
- Artifact management: Nexus, Artifactory
- GitOps workflows
- Dynatrace, Datadog, Prometheus, Grafana, Splunk, ELK
- Event correlation, anomaly detection, predictive analytics
- AIOps platforms & auto-remediation frameworks
- IAM, secrets management, PKI & certificates
- CSPM: Prisma Cloud, Wiz, Defender
- Security incident response
- Vulnerability scanning & remediation
- Python, Bash, PowerShell, Go (any two)
- REST APIs, event-driven automation
- Serverless automation workflows
- SLIs, SLOs, SLAs
- Error budgets
- Chaos engineering
- Capacity planning & forecasting
- Performance engineering
- Strong problem-solving mindset
- Excellent troubleshooting skills
- High ownership & accountability
- Stakeholder communication
- Ability to work in 24x7 global operations model
Corporate Security Responsibility
Abide by Mastercard’s security policies and practices;
Ensure the confidentiality and integrity of the information being accessed;
Report any suspected information security violation or breach, and
Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
