Home

Nidhin Chandran - DevOps/SRE/CloudEngineer
[email protected]
Location: Bentonville, Arkansas, USA
Relocation: Yes
Visa: H1B
Resume file: Nidhin_AWSDevops_1788760509684.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
PROFESSIONAL EXPEREINCE
Walmart, USA May 2025 Present
Role: DevOps Engineer
Spearheading containerized deployments with Docker and Kubernetes, optimizing scaling, rolling updates, and self-healing capabilities resulting in improved reliability and 30% faster issue resolution for production workloads.
Designed and maintained CI/CD pipelines integrating Jenkins, GitHub Actions, and Artifactory, streamlining build, test, and release workflows while cutting deployment times by 40%.


Implemented OPA Gatekeeper policies to enforce security and compliance standards across Kubernetes clusters, reducing misconfigurations and audit findings by 25%.
Implemented security hardening measures following CIS benchmarks and organizational security policies, including SELinux configuration, firewall rules, and access controls.
Collaborated with developers, QA, and cloud engineers to automate infrastructure provisioning on Linux-based systems, improving environment consistency and eliminating 20+ hours of manual configuration per month.
Championed proactive monitoring and logging strategies with Prometheus, Grafana, and ELK Stack, ensuring high availability and reducing mean-time-to-detect (MTTD) by 35%.
Manage and maintain VMware vSphere infrastructure for Linux virtual machines [Edge store cluster].
Partnered with cross-functional teams to enhance release governance and delivery practices, improving stakeholder confidence and accelerating feature rollout cycles for customer-facing applications.
Configure and troubleshoot network interfaces, bonding, and teaming on Linux servers.
WPP, San Francisco, CA August 2024 March 2025
Role: Sr. DevOps Engineer
Description: My role was pivotal in automating CI/CD pipelines, managing hybrid cloud infrastructures, optimizing containerized microservices, and implementing advanced monitoring and security solutions. By driving infrastructure automation and enabling seamless deployments, I have reduced deployment times by 40%, minimized downtime to 99.99% availability, and improved operational efficiency delivering measurable cost savings of $3 million annually.
Responsibilities
Architected and managed scalable, secure, and highly available cloud infrastructures on AWS and Azure (EC2, VPC, RDS, IAM, S3, Route53, CloudFront, AKS, EKS), automating deployments with Terraform and Helm for consistency across environments.
Deployed and orchestrated enterprise-grade Kubernetes clusters (EKS, AKS) and Dockerized applications with dynamic scaling, rolling updates, and self-healing, ensuring high availability and resiliency of microservices.
Conducted root cause analysis for database server performance issues, implementing Linux kernel tuning and system resource optimization.
Implemented security hardening measures following CIS benchmarks and organizational security policies, including SELinux configuration, firewall rules, and access controls.
Engineered CI/CD pipelines with Jenkins, GitLab, SonarQube, Nexus, and automated testing workflows, cutting deployment time by 40% while improving release quality and governance.
Configured S3 bucket policies, IAM roles, and access controls following least privilege principles
Troubleshoot network connectivity issues between servers, applications, and databases.
Implemented observability solutions using ELK Stack, AWS CloudWatch, and Prometheus/Grafana, establishing real-time alerting, system health monitoring, and performance tuning across cloud workloads.
Implemented error handling, logging, and monitoring for Glue workflows using CloudWatch
Created Glue crawlers for automatic schema discovery and metadata cataloging in AWS Glue Data Catalog
Designed disaster recovery and HA strategies with Pacemaker, Corosync, Rsync, and multi-region cloud setups, achieving near-zero downtime for mission-critical services.
Automated infrastructure provisioning and configuration via Ansible, Terraform, and Proxmox, integrating version-controlled workflows in Git repositories to enable peer reviews and infrastructure-as-code practices.
Strengthened DevOps & Linux operations by managing artifacts (Artifactory, Maven, S3), securing SSH access, analyzing network traffic with tcpdump, and optimizing VMware/Proxmox/KVM environments for cost-effective scaling.
Environment/Tools: Git, GitHub, Jenkins, Maven, SonarQube, Nexus, Kubernetes, EKS, AWS Fargate, AWS ECS, Prometheus, Grafana, ELK Stack, Terraform, AWS SageMaker, PyTorch

COMCAST (Remote) August 2022 June 2024
Role: Sr DevOps Engineer Responsibilities
Deployed and managed LXC containers on Proxmox and Kubernetes clusters in production, ensuring lightweight virtualization, high availability, and efficient resource utilization across DevOps workflows.
Participated in production support and on-call rotations, resolving deployment and infrastructure issues.
Collaborated with security teams to implement best practices for access control, RBAC, and secure pipelines.
Supported API lifecycle management including versioning, security policies, throttling, and subscription management in APIM.
Support zero-downtime deployments, blue/green deployments, and rollback strategies.
Orchestrated EKS upgrades and disaster recovery strategies with blue-green/canary deployments, enabling seamless application updates, fault tolerance, and zero-downtime resilience.
Configured AWS IAM roles, accounts, and policies to enforce least-privilege access, while leveraging EC2, S3, RDS, Route 53, ELB, and Auto Scaling Groups to build secure, scalable, and highly available infrastructures.
Developed and maintained end-to-end CI/CD pipelines using Jenkins, Terraform, and JUnit testing, reducing manual intervention, accelerating release cycles, and improving code reliability.
Managed Docker and container images in Nexus Artifactory, hosted Maven repositories, and automated dependency management to streamline builds and deployment lifecycles.
Implemented infrastructure monitoring and logging with AWS CloudWatch, Nagios, and the ELK Stack, including long-term log storage in S3, enabling proactive alerts and compliance tracking.
Deployed and optimized databases including MySQL, PostgreSQL, and MongoDB clusters in containerized and virtualized environments, improving scalability and performance.
Centralized Terraform state management in S3 and Terraform Cloud for collaborative infrastructure-as-code practices, ensuring consistency and auditability across teams.

Environment/Tools: Linux, AWS,Docker, Kubernetes, Jenkins, Nexus Artifactory, Nagios, ELK stack, Ansible, GitLab, Route53, IAM, CloudWatch, S3, EC2, Security Groups, TCP/IP, DNS, Shell/Bash Scripting

GBS PLUS PVT LTD, India April 2018 July 2022
Role: Sr Linux Engineer | SRE Responsibilities
Administered Linux systems (Red Hat, Debian, Ubuntu) by installing, upgrading, and managing packages using YUM, RPM, and APT; performed advanced kernel tuning, resource allocation, and automated backups with Crontab to optimize performance and reliability.
Implemented Linux server hardening with SELinux, AppArmor, Fail2ban, Firewalld, and iptables; enforced PCI DSS/GDPR compliance through secure access controls, password-less SSH authentication, and regular vulnerability scans with Nessus, Qualys, and Lynis.
Monitored cloud resource utilization using AWS Cost Explorer, CloudWatch, and third-party tools
Executed quarterly OS patching cycles for RHEL servers, ensuring 99.9% uptime and maintaining strict PCI-DSS compliance requirements through coordinated maintenance windows.
Installed and managed application and web servers including WebSphere, WebLogic, Apache, HTTPD, and Tomcat; handled SSL certificate integrations (Certbot, GoDaddy) and domain management via cPanel for secure and scalable deployments.
Built and maintained CI/CD pipelines with Jenkins, automating build, test, and deployment workflows; engineered VM images for VMware vSphere, OpenShift, and Workstation, streamlining infrastructure provisioning and reducing manual intervention.
Deployed monitoring and observability stacks using Prometheus, Grafana, Datadog, CloudWatch, and Kafka Manager; established real-time dashboards, alerts, and pod-level analysis that improved system reliability by 25% and cut response times by 40%.
Deployed and managed EC2 instances with auto-scaling groups for compute-intensive data processing workloads
Implemented custom Linux server builds with organization-specific configurations, security hardening, and application-specific tuning parameters.
Integrated monitoring with incident management systems like PagerDuty, automated response playbooks, and ticket resolution scripts; reduced MTTR significantly and standardized RCA processes through detailed post-incident reviews.
Orchestrated end-to-end streaming pipelines by integrating Kafka with Splunk, enabling seamless data ingestion and real-time analytics; monitored throughput, latency, and stream health with Grafana and Kafka Manager for continuous optimization.
Configured AWS CloudWatch detailed monitoring and custom metrics for performance tracking
Manage and optimize Cloudera Hadoop ecosystems (HDFS, Hive, Spark, YARN, etc.)
Implemented AWS VPC architectures with proper subnet segmentation, security groups, and network ACLs
Techvantage Systems, India November 2014 March 2018
Role: Sr. System Administrator Responsibilities
Installed, configured, and maintained Windows Server environments, ensuring stable and secure operations.
Performed server patching and updates using Windows Server Update Services (WSUS) and automated tools to maintain compliance and reduce vulnerabilities.
Managed Active Directory (AD), including creating and managing users, groups, and organizational units (OUs) to maintain efficient directory structures
Configured role-based access control (RBAC) to enforce least privilege access policies across user accounts, enhancing security.
Administered and maintained Red Hat Enterprise Linux servers across production, development, and testing environments in a 24x7 high-availability infrastructure.
Implemented and maintained Group Policy Objects (GPOs) to enforce organizational security and compliance requirements, such as password policies and application restrictions
Deployed and maintained System Center Configuration Manager (SCCM) for operating system imaging, application deployment, and patch management, streamlining endpoint management.
Designed and managed Active Directory (AD) infrastructure, including domains, forests, and trust relationships, ensuring seamless authentication and resource access across the organization.
Created and maintained standardized system images using Sysprep for consistent deployment across diverse hardware and virtual environments.
Performed system performance tuning and capacity planning for Linux servers, optimizing CPU, memory, disk I/O, and network utilization to support business-critical applications.
Environment/Tools: Windows Server 2008 R2, Windows Server 2012, CentOS, Ubuntu Server, VMware ESXi, Task Manager, Resource Monitor, Performance Monitor, YUM, Hyper-V.
Keywords: continuous integration continuous deployment quality analyst sthree active directory bay area California

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7699
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: