Home

Mallikarjuna - SITE RELIABILITY ENGINEERING LEAD | PLATFORM ENGINEERING | KUBERNETES & OBSERVABILITY
[email protected]
Location: Houston, Texas, USA
Relocation: Yes
Visa: H1B
Resume file: P Mallikarjun_Palleboina_SRE_Lead_Resume_1787925733626.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
MALLIKARJUN PALLEBOINA
SITE RELIABILITY ENGINEERING LEAD | PLATFORM ENGINEERING | KUBERNETES & OBSERVABILITY
Plano, TX (949) 570 3917
PROFESSIONAL SUMMARY
Site Reliability Engineering Lead with 16+ years operating fault-tolerant, high-availability platforms across telecom, financial services, and e-commerce. Owns SLO/SLI/SLA and error-budget governance end to end, commands P0/P1 incident response for 24/7 production Guardian, and runs Kubernetes fleets at enterprise scale. Consistently converts reliability practice into measurable outcomes: 99.99% SLO attainment, 40% MTTR reduction, and 60% of Tier-1 operational toil automated through Python and AI-driven tooling. CKA-certified; multi-cloud across AWS, GCP, and Azure.
SELECTED IMPACT
99.99%+ availability SLO sustained across production, staging, and disaster-recovery environments including a national device-launch peak at 5 baseline traffic.
40% MTTR reduction through alert correlation, automated runbook execution, and structured incident command across P0/P1 events.
60% of Tier-1 incident tasks automated using Python agentic AI workflows, redirecting SRE capacity from ticket handling to engineering work.
50% faster CI/CD delivery via standardized GitLab pipelines with automated quality, security, and blue/green rollback gates.
Zero SLA breaches through the iPhone 17 launch window, with supply-chain infrastructure holding at 3 5 baseline traffic.
CORE TECHNICAL SKILLS
Reliability Engineering: SLO/SLI/SLA design and error-budget governance, SLA enforcement, reliability roadmapping, capacity planning, chaos engineering (LitmusChaos), toil elimination, MTTD/MTTA/MTTR/MTBF optimization
Kubernetes & Containers: CKA-certified EKS, AKS, OpenShift; cluster lifecycle and upgrades, node pool scaling, cordon/drain, ETCD backup and restore, RBAC, NetworkPolicy, Ingress, HPA/VPA autoscaling, CoreDNS, Helm, OPA/Gatekeeper, CrashLoopBackOff / OOMKill / ImagePullBackOff triage
Observability: OpenTelemetry(traces, metrics, logs), Prometheus (PromQL, Alertmanager, federation), Grafana, Splunk (SPL), Dynatrace (PurePath APM, Davis AI), Datadog, AppDynamics, ELK/Kibana, CloudWatch, Zabbix, AlertSite
Cloud Platforms: AWS (EKS, EC2, S3, RDS, IAM, VPC, ALB, Route53, CloudFront, CloudWatch), GCP, Azure AKS; FinOps cost optimization
Infrastructure as Code: Terraform, Ansible, Chef, CloudFormation
CI/CD & GitOps: GitLab CI/CD, Jenkins, GitHub Actions, ArgoCD, blue/green and canary deployments, automated rollback, compliance drift detection, SAST/DAST gates.
AI/ML for Operations (AIOps): Agentic AI workflows, Model Context Protocol (MCP), RAG pipelines, LangGraph, AI-driven anomaly detection, predictive alerting, ML-based alert correlation, AI-assisted root cause analysis
Incident & Service Management: Incident command for P0/P1 war rooms, 24/7 on-call (PagerDuty), blameless post-mortem facilitation, escalation management, ServiceNow (Incident, Problem, Change, Release, Knowledge), ITIL v4, Status.io, Confluence, Jira
Security & Networking: HashiCorp Vault/Consul, CyberArk, Fortify SAST, WAF authoring (Akamai, Cloudflare, Imperva), DDoS and bot mitigation, F5 BigIP, Nginx, Apigee, Okta, DNS/SSL/TLS
Languages: Python, Bash/Shell, Java/J2EE, SQL, JavaScript
PROFESSIONAL EXPERIENCE
T-MOBILE Plano, TX (Hybrid) | Oct 2024 Present
Site Reliability Engineering Lead Supply Chain Digital Platform and Digital Finance.

Reliability owner for the supply-chain digital platform (inventory, order management, and fulfillment) supporting nationwide retail and direct sales channels. Lead 24/7 production operations and incident command; mentor a team of SRE and junior SRE engineers.
SLO / SLI / Error-Budget Governance
Own the SLO/SLI/SLA framework across six supply-chain platform services: defined SLIs for request success rate, p50/p95/p99 latency, error rate, and availability; set 99.95 - 99.99% targets; publish weekly error-budget burn reports to engineering and executive stakeholders.

Sustained 99.99%+ availability across production, staging, and disaster-recovery environments through continuous SLI tracking, error-budget alerting, and automated rollback triggers wired directly into GitLab pipelines.
Built multi-window burn-rate dashboards in Grafana and Dynatrace on the Google SRE error-budget model (1-hour and 6-hour windows), enabling deploy freezes before budget exhaustion rather than after.
Reduced SLA violation risk 40% by tightening ITIL Incident Problem Change governance in ServiceNow, correlating recurring incidents to underlying problems and driving permanent fixes to closure.
Incident Command & 24/7 Production Guardian
Command P0/P1 bridge calls as incident lead mobilizing Dev, Infrastructure, Network, Security, and Business responders within SLA response windows, driving parallel investigation tracks, and broadcasting live status to internal and external stakeholders via Status.io.
Served as single point of contact for the iPhone 17 and 17 Pro launch war rooms, holding supply-chain digital infrastructure at zero disruption under peak traffic 3 5 above baseline.
Own the blameless post-mortem program end to end: facilitate 5-Why and timeline-reconstruction sessions within 48 hours of P0/P1 resolution, document contributing factors in Confluence, and track corrective actions to closure with named owners measurably reducing repeat-incident rate.
Cut mean time to detect and diagnose by collapsing alert storms into single actionable incidents through correlation rules authored in Splunk and Dynatrace; reduced overall alert fatigue 40%.

Kubernetes Operations & Platform Engineering

Manage full EKS cluster lifecycle: in-place rolling and blue/green node-group upgrades, On-Demand and Spot node pool scaling, taint/toleration management, cordon and drain procedures, and ETCD snapshot backup/restore validated in DR exercises.
Diagnose and resolve production Kubernetes incidents on P1 on-call CrashLoopBackOff from misconfigured liveness probes and OOMKill, ImagePullBackOff from registry auth and digest mismatches, Pending pods from resource-quota exhaustion and affinity conflicts, and CoreDNS/CNI networking failures.
Certified zero downtime under 5 traffic surge by running chaos and burst testing on EKS with LitmusChaos simulating pod kills, node failures, and network partitions, then tuning HPA/VPA replica bounds against observed behavior.
Reduced Kubernetes incident MTTR 40% by deploying OpenTelemetry Collector as a DaemonSet alongside kube-state-metrics, Node Exporter, Prometheus ServiceMonitor CRDs, and per-namespace Grafana dashboards.
Hardened cluster security posture: ETCD encryption at rest, least-privilege RBAC, OPA/Gatekeeper admission policies, NetworkPolicy segmentation, and secrets delivery via HashiCorp Vault Agent Injector.
Automation & AI-Driven Operations
Automated 60% of recurring Tier-1 incident tasks by architecting agentic AI workflows in Python (MCP, RAG pipelines, LangGraph) for triage and compliance validation freeing SRE capacity for engineering work.
Eliminated 50% of manual operational toil by automating runbook execution, health-check validation, certificate rotation, log triage, and ticket enrichment in Python and Bash measured against the Google SRE 50% toil cap.
Deployed AI-driven anomaly detection and ML-based alert correlation (Dynatrace Davis AI plus custom models) to surface degradation ahead of SLO burn, shifting the team from reactive firefighting to predictive operations.
Cut CI/CD deployment effort 50% by standardizing GitLab pipelines with automated quality gates, blue/green rollback, compliance drift detection, and SAST/DAST security gates; provisioned DEV through PROD AWS infrastructure via Terraform and Ansible.
Environment: Kubernetes/EKS, AWS, Terraform, Ansible, GitLab CI/CD, ArgoCD, Python, Bash, OpenTelemetry, Prometheus, Grafana, Splunk, Dynatrace, AppDynamics, Datadog, LitmusChaos, HashiCorp Vault, PagerDuty, ServiceNow, Akamai, Cloudflare, Okta, CyberArk, MCP, RAG, LangGraph, Status.io
SS&C TECHNOLOGIES (INTRALINKS) Hyderabad, India (Hybrid) | Mar 2020 Oct 2024
Senior Site Reliability Engineer Global Financial SaaS Platform
Reliability and production support for a global financial SaaS platform hosting M&A and private-equity data rooms for enterprise clients.
SLO Governance & Incident Command
Designed the platform's founding SLO/SLI framework: instrumented API success rate, p99 latency, and data-room availability via Prometheus recording rules and Grafana burn-rate dashboards; sustained 99.9 99.95% SLO targets across all critical services.
Established and enforced error-budget policy published bi-weekly burn reports to product and engineering leadership and halted non-critical deploys when 50% of budget was consumed in a rolling 30-day window.
Commanded P0/P1 war rooms for the global platform: led bridge calls, drove parallel RCA tracks across application, infrastructure, CDN, and security layers, and delivered live status directly to enterprise clients and internal leadership.
Facilitated blameless post-mortems for every P0/P1 within 48 hours 5-Why analysis, contributing-factor mapping, and corrective actions tracked to named owners in Confluence driving down recurrence of known failure modes.
Reduced false-positive alert rate 40% through composite thresholds, Alertmanager silence and inhibition rules, and alert dependency chaining improving on-call quality of life and SLO signal fidelity.
Platform, Observability & Security Engineering
Owned end-to-end traffic architecture and served as incident bridge lead across the full request path: Internet Okta (IdP) Akamai (WAF/CDN) F5 Nginx Apigee platform services diagnosing failures at every layer during live incidents.
Authored managed and custom Akamai WAF rules (bot mitigation, rate limiting, API abuse controls) and Cloudflare-integrated Splunk security dashboards achieving 45 60% reduction in malicious traffic and 40% reduction in security alert noise.
Maintained Kubernetes control-plane and node components (ETCD, kube-controller-manager, kubelet, kube-proxy, CoreDNS, ingress controllers); resolved pod failures, resource-quota exhaustion, PV binding failures, and service-discovery breakdowns during P1 incidents.
Architected the Kubernetes observability stack Prometheus with kube-state-metrics and Node/Blackbox Exporters, per-namespace Grafana dashboards, and Dynatrace OneAgent DaemonSet for full-stack APM tracing to pod level; improved monitoring efficiency 60%.
Resolved JVM heap and thread-dump performance issues on containerized Java workloads and tuned resource requests/limits to eliminate recurring OOMKill events.
Provisioned AWS infrastructure (EC2, EKS, S3, VPC, RDS, ALB, IAM, Route53) via Terraform; operated HashiCorp Vault, Consul, and Nomad for secrets management, service discovery, and workload orchestration.
Environment:
Kubernetes, AWS, Terraform, Python, Prometheus, Grafana, Splunk, Dynatrace, Datadog, AlertSite, Akamai, Cloudflare, Okta, F5, Nginx, Apigee, HashiCorp Vault/Consul/Nomad, Jenkins, ServiceNow, PostgreSQL, Status.io

KONY LABS (NOW TEMENOS INFINITY) India | May 2019 Mar 2020
Technical Lead SRE / DevOps
Stood up the SRE practice for a multi-tenant mobile-banking platform: defined SLO/SLI baselines for platform APIs, implemented error-budget tracking, established PagerDuty on-call rotations, and authored incident escalation playbooks.
Commanded P1 bridge calls and led blameless post-mortems for global enterprise banking customers, publishing RCA documentation to Confluence.
Built the observability layer (Splunk, Grafana, Prometheus, Dynatrace, AlertSite, Zabbix) enabling proactive SLI alerting; provisioned AWS infrastructure via Terraform and enforced Fortify SAST as a CI quality gate.
SEARS HOLDINGS India | Jul 2017 May 2019
Senior Technical Associate SRE / Production Support
Managed L0 L4 incident response and P1 war-room participation for a high-volume e-commerce production ecosystem; led post-mortem analysis and annual disaster-recovery drills.
Designed Grafana, Prometheus, and Dynatrace dashboards with SLI-aligned alerting for server availability and Control-M batch job success rates.
Drove Incident, Problem, Change, and Service Request resolution in ServiceNow under ITIL; applied Fortify SAST scanning and OS security patching pre-release.
EARLIER EXPERIENCE
Cognizant Technology Solutions Production Support Lead, US Insurance (Mar 2015 Jun 2017). Ran 24/7 production Guardian and P1 escalation bridges for a major US insurance carrier; owned PROD batch-cycle monitoring, JProfiler/JMeter performance profiling, Fortify SAST scanning, and annual DR exercises.
HCL Technologies Senior Consultant, Global Logistics Platform (Jun 2012 Mar 2015). Supported 24/7 production across 36 countries; diagnosed JVM performance issues via heap and thread dumps, enhanced Java/J2EE REST microservices, and led post-deployment API regression testing.
Mphasis (an HP Company) / Magna Infotech Software Engineer, Production Support (Dec 2009 Apr 2012). Built Java/J2EE web services on WebLogic across telecom, automotive, and banking clients; secured services with WS-Security and Fortify scanning while participating in 24/7 support rotations.
CERTIFICATIONS & EDUCATION
Certified Kubernetes Administrator (CKA) The Linux Foundation | Credential ID: LF-4ky283goge
Sun Certified Java Programmer (SCJP) Java Platform, J2SE
Bachelor of Engineering, Computer Science & Engineering Anna University, Chennai, 2007
AWARDS & RECOGNITION
Best Performer of the Year awarded three consecutive years (FY 2022, FY 2023, FY 2024) for sustained production reliability and incident leadership.
Keywords: cprogramm continuous integration continuous deployment artificial intelligence machine learning business intelligence sthree ffive hewlett packard Idaho Texas

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7680
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: