| Manasa - Senior Data Engineer |
| [email protected] |
| Location: Grayson, Georgia, USA |
| Relocation: |
| Visa: |
| Resume file: Manasa B - Sr. Data Engineer_1786562220070.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
Manasa Badugu
Sr. Data Engineer [email protected] | +1 704-791-3158 PROFESSIONAL SUMMARY: Senior Data Engineer with 10 years of experience designing, developing, and optimizing scalable data engineering solutions across Azure and AWS cloud platforms. Extensive experience building enterprise data lakes, data warehouses, ETL/ELT pipelines, and real-time data processing frameworks using Azure Data Factory, Azure Synapse Analytics, Azure Databricks, PySpark, Apache Spark, AWS Glue, Amazon Redshift, Apache Kafka, and Delta Lake. Strong expertise in developing high-performance data transformation pipelines for processing structured, semi-structured, and unstructured data from multiple enterprise data sources. Hands-on experience designing scalable cloud-based data architectures that support business intelligence, advanced analytics, regulatory reporting, and machine learning initiatives. Proficient in Python, SQL, PySpark, Spark SQL, Azure Data Factory, Databricks, Snowflake, Redshift, Delta Lake, and Apache Kafka for developing reliable and efficient data integration solutions. Experienced in implementing dimensional data models including Star Schema and Snowflake Schema, creating optimized fact and dimension tables to improve reporting performance. Strong knowledge of data ingestion techniques, including batch and streaming data processing, Change Data Capture (CDC), incremental loading, and API-based integrations. Skilled in optimizing Spark workloads through partitioning, caching, query tuning, indexing strategies, and resource optimization to improve processing efficiency and reduce execution time. Experience integrating data from relational databases, cloud storage, REST APIs, flat files, enterprise applications, and third-party systems into centralized analytics platforms. Hands-on experience implementing data quality validation, reconciliation processes, exception handling, and monitoring frameworks to ensure data accuracy and consistency across enterprise environments. Experience working with cloud storage technologies including Azure Data Lake Storage (ADLS Gen2), Azure Blob Storage, Amazon S3, and Delta Lake for large-scale data management. Strong understanding of data governance, metadata management, security best practices, role-based access control (RBAC), encryption, and regulatory compliance within cloud data platforms. Experienced in modernizing legacy ETL systems by migrating on-premises data pipelines to cloud-native architectures while improving scalability, maintainability, and operational efficiency. Hands-on experience supporting CI/CD implementations using Azure DevOps, Git, and automated deployment pipelines for version control and release management. Collaborated with enterprise data integration teams on legacy ETL modernization initiatives, supporting migration planning, data validation, reconciliation, and operational readiness for large-scale analytics platforms. Partnered with business leaders, analytics teams, and solution architects to support AI/ML enablement, advanced analytics adoption, and enterprise customer analytics initiatives across cloud data platforms. Strong analytical, problem-solving, and communication skills with a consistent track record of delivering high-quality, scalable, and reliable data engineering solutions in Agile development environments. TECHNICAL SKILLS: Cloud Platforms: Azure (Azure, (Data Factory (ADF), Synapse Analytics, Data Lake Storage Gen2, Blob Storage, Event Hubs, Stream Analytics, Monitor, Log Analytics, Machine Learning), AWS( Glue, Lambda, Amazon S3, Amazon Redshift, Amazon EMR, Amazon SageMaker, Amazon CloudWatch) Data Engineering & Big Data: ETL, Data Pipelines, Data Integration, Data Ingestion, Data Transformation, Apache Spark, PySpark, Hadoop, HDFS, Apache Hive, HiveQL, Apache Sqoop, Spark SQL, Delta Lake, Apache Airflow, Apache Oozie, Kafka, Streaming Data Processing, Batch Processing Programming & Databases: Python, SQL, Bash Shell Scripting, Unix/Linux, REST APIs, SQL Server, Oracle, Azure SQL Database, Data Warehouse, Data Modeling, Dimensional Modeling, Star Schema, Snowflake Schema, Stored Procedures, Query Optimization Machine Learning & Analytics: Feature Engineering, Data Preprocessing, Predictive Analytics, ML Data Pipelines, Model Deployment Support, Pandas, NumPy, Scikit-learn, Power BI, Tableau Data Architecture & Governance: Data Lake Architecture, Cloud Data Architecture, Hybrid Data Architecture, Data Lineage, Metadata Management, Data Quality, Data Validation, Data Profiling, Data Cleansing, Data Governance, Data Security, Data Encryption, IAM, RBAC DevOps & Operations: Azure DevOps, CI/CD Pipelines, Docker, Deployment Automation, Infrastructure Automation, Production Support, Monitoring, Troubleshooting, Root Cause Analysis, Performance Tuning, SLA Management Development Practices: Agile Scrum, Code Reviews, Modular Development, Exception Handling, Technical Documentation, Solution Design, Knowledge Sharing PROFESSIONAL EXPERIENCE: Expion Health, Dallas TX | September 2025 Present Sr. Data Engineer Responsibilities: Redesigned enterprise ETL workflows across Azure Data Factory, Databricks, and Synapse Analytics, supporting 1TB+ daily data ingestion while improving pipeline reliability, scalability, and operational performance. Built PySpark-based feature engineering pipelines to prepare, transform, and optimize machine learning datasets for predictive analytics and advanced data science initiatives. Automated application and infrastructure deployment processes using Azure DevOps CI/CD pipelines, improving release reliability, deployment efficiency, and environment consistency. Managed production monitoring and support activities using Azure Monitor and Log Analytics, troubleshooting pipeline failures, analyzing system issues, and ensuring SLA compliance for business-critical workloads. Automated recurring data processing activities through ADF scheduling triggers and Databricks job orchestration, reducing manual operational dependencies and improving workflow efficiency. Supported real-time analytics implementations by integrating event-driven streaming data pipelines with downstream reporting, analytics, and operational platforms. Architected and modernized enterprise-scale data engineering solutions using Azure Data Factory (ADF) and Azure Synapse Analytics to support scalable batch processing and near real-time data workloads. Designed and implemented Azure Data Lake Storage Gen2 (ADLS Gen2) architectures with structured data zones, including raw, curated, and consumption layers to improve data organization, accessibility, and governance. Developed distributed data processing frameworks using PySpark on Azure Databricks to transform and analyze large-volume datasets across cloud-based environments. Integrated data from 20+ enterprise source systems, consolidating millions of transactional records into centralized data platforms supporting critical reporting and analytics initiatives. Implemented real-time data ingestion solutions using Azure Event Hubs and Azure Stream Analytics to enable low-latency processing and operational analytics capabilities. Improved Azure Synapse Analytics performance by optimizing SQL workloads, distributed query execution, storage configurations, and analytical data processing strategies. Troubleshot application integration, data pipeline, authentication, authorization, and cloud platform issues, performing root cause analysis and implementing corrective actions to maintain reliable enterprise operations. Partnered with business, application, platform, and security teams to understand integration requirements and deliver scalable, secure, and reliable cloud-based solutions. Collaborated with business stakeholders, enterprise architects, data analysts, and data science teams to convert business requirements into scalable Azure-based data engineering solutions. Developed Python-based data validation, reconciliation, and exception-handling frameworks to improve enterprise reporting accuracy and downstream analytics reliability. Integrated batch and streaming pipelines with enterprise analytics platforms to enhance customer and operational analytics availability across business teams. Collaborated with application teams and enterprise architects to understand application architecture, data flows, integration dependencies, and access requirements, ensuring seamless connectivity between applications and cloud data platforms. Supported application onboarding by coordinating data integration requirements, connectivity setup, authentication, authorization, and access provisioning for applications connecting to enterprise data platforms. Implemented and supported cloud security policies, RBAC controls, encryption standards, and access management practices to protect enterprise data and application integrations. Collaborated with architects, data analysts, and analytics stakeholders to support advanced analytics enablement and scalable cloud data platform adoption. Contributed to enterprise data governance initiatives by documenting validation rules, reconciliation procedures, lineage information, and operational support processes. Leveraged Delta Lake architecture within Azure Databricks to enable ACID transactions, schema enforcement, time travel/data versioning, and reliable data management for analytical workloads. Implemented enterprise security and governance controls using Azure Active Directory (AAD), Role-Based Access Control (RBAC), encryption standards, and secure authentication mechanisms to protect cloud data assets. Designed and optimized dimensional data models using Star Schema and Snowflake Schema methodologies to support enterprise BI, reporting, and analytical processing requirements. Applied advanced data engineering practices including modular development, reusable frameworks, code optimization, performance tuning, exception handling, and scalable cloud architecture design. Provided technical mentorship to junior data engineers through code reviews, troubleshooting assistance, design guidance, and knowledge sharing focused on engineering best practices. Created and maintained detailed technical documentation covering solution architecture diagrams, ETL workflows, metadata management processes, and data lineage to support governance, compliance, audits, and long-term maintainability. Supported machine learning initiatives by developing scalable feature engineering pipelines and assisting data science teams with ML model deployment and operationalization workflows. Charles Schwab, Pheonix AZ | March 2024 August 2025 Sr. Data Engineer Responsibilities: Developed cloud-native data integration solutions using AWS Glue and AWS Lambda to automate ETL processes, streamline data movement, and orchestrate serverless workflows across AWS services. Designed, implemented, and optimized Amazon Redshift data warehouse environments, including schema design, workload management, query optimization, and performance tuning for analytical reporting. Created machine learning data pipelines to support feature engineering, training dataset generation, and end-to-end model lifecycle processes, improving data readiness for data science initiatives. Developed RESTful APIs and reusable data service endpoints to securely expose processed datasets for downstream applications, business reporting, and analytics platforms. Established data quality frameworks incorporating rule-based validations, anomaly detection, and automated integrity checks to improve data accuracy and minimize reporting discrepancies. Performed feature extraction, data preprocessing, transformation, and normalization using Python libraries including Pandas, NumPy, and Scikit-learn to support predictive analytics and AI-driven solutions. Optimized Apache Spark applications and complex SQL queries through performance tuning, reducing execution time and improving resource utilization across distributed computing environments. Developed high-performance distributed data processing pipelines using Apache Spark (PySpark) on AWS EMR to process large-scale datasets for enterprise analytics and reporting workloads. Built and managed scalable Amazon S3 data lake architectures to store, organize, and maintain multi-terabyte structured and semi-structured datasets for enterprise data platforms. Engineered scalable ETL pipelines using Python and SQL to ingest, transform, and integrate data from diverse internal and external source systems into centralized analytical repositories. Collaborated with enterprise ETL teams on legacy DataStage modernization and cloud migration initiatives, supporting source-to-target mapping validation, reconciliation testing, and production cutover activities. Partnered with analytics and business stakeholders to enable cross-channel customer interaction analytics through curated datasets, API-driven data services, and near real-time event-processing pipelines. Supported AI/ML initiatives by engineering scalable feature-engineering and training datasets and assisting data science teams with model operationalization and business adoption activities. Provided technical recommendations on data integration patterns, streaming architectures, cloud data services, and analytics platform optimization to support enterprise modernization objectives. Containerized data engineering applications with Docker to ensure consistent deployment, portability, and reliable execution across development, testing, and production environments. Monitored cloud infrastructure and data pipeline health using AWS CloudWatch, investigated production issues, performed root cause analysis, and implemented corrective actions to improve system reliability. Delivered curated, analytics-ready datasets and reporting layers that enhanced Tableau dashboard performance while enabling faster and more accurate business insights. Actively participated in Agile Scrum ceremonies, including sprint planning, backlog refinement, daily stand-ups, sprint reviews, and release activities to deliver iterative data engineering solutions. Provided production support through on-call rotations by troubleshooting pipeline failures, resolving data processing issues, and minimizing the impact of downstream reporting disruptions. Maintained comprehensive technical documentation covering ETL processes, machine learning pipeline architecture, metadata management, data lineage, and operational procedures to support governance and knowledge sharing. Implemented secure data access and governance controls using AWS IAM roles, encryption mechanisms, and policy-based security configurations to meet enterprise compliance and security standards. BayOne Solutions, India | February 2019 July 2023 Data Engineer Responsibilities: Developed end-to-end ETL workflows using Python and SQL to extract, transform, and load structured and semi-structured data from multiple enterprise applications into centralized data repositories. Leveraged Hadoop technologies, including HDFS and Apache Hive, to store, manage, and query large-scale datasets supporting business intelligence, reporting, and analytical initiatives. Performed data profiling, cleansing, normalization, and validation using Python libraries such as Pandas and NumPy to improve data accuracy, consistency, and overall data quality. Designed and implemented batch data ingestion pipelines with Azure Data Factory (ADF) to securely move data between on-premises systems and Azure cloud platforms. Built scalable data transformation pipelines using Apache Spark and PySpark to process high-volume datasets, improving processing efficiency and minimizing execution time for enterprise workloads. Diagnosed and resolved production issues involving ADF pipelines, Spark jobs, and Hive queries by performing root cause analysis and implementing long-term performance improvements. Utilized Apache Sqoop to efficiently migrate and synchronize large datasets between SQL Server databases and Hadoop environments for reporting, analytics, and archival purposes. Automated recurring operational tasks and batch processing activities using Bash Shell scripting, Azure Data Factory scheduling triggers, and Cron jobs, improving workflow efficiency and reducing manual intervention. Developed and optimized complex HiveQL and Spark SQL queries with joins, aggregations, filtering, and advanced transformations to support large-scale analytical and reporting requirements. Assisted with integrating on-premises Hadoop platforms with Azure cloud services, contributing to hybrid data architecture implementations and cloud migration initiatives. Managed raw, staging, and curated datasets using Azure Blob Storage and Azure Data Lake Storage (ADLS Gen1), ensuring scalable, secure, and cost-effective data storage solutions. Supported daily ETL production activities by monitoring scheduled workflows, troubleshooting pipeline failures, and ensuring reliable data processing across Hadoop and Azure ecosystems. Designed and maintained dimensional data models using Star Schema and Snowflake Schema within Azure SQL Database and SQL Server to support enterprise reporting and business intelligence solutions. Prepared, transformed, and validated analytics-ready datasets for Power BI dashboards, enabling accurate reporting and data-driven business decisions. Recro, India | March 2015 January 2019 Junior Data Engineer Responsibilities: Performed data preparation activities using Python with Pandas and NumPy, including cleansing, normalization, validation, and transformation to improve data quality and usability. Monitored daily ETL workflows, reviewed execution logs, and resolved data pipeline issues in Linux/Unix environments to maintain reliable data processing operations. Developed Bash Shell scripts to automate batch processing, file movement, scheduled jobs, and routine operational tasks, reducing manual effort and improving process consistency. Applied dimensional data modeling techniques, including Star Schema and Snowflake Schema, to design efficient data warehouse structures that supported enterprise reporting requirements. Assisted in developing batch-based data ingestion processes using Apache Hive and Apache Sqoop to load large datasets from relational databases into Hadoop HDFS efficiently. Designed and maintained Apache Oozie workflows to schedule, orchestrate, and automate ETL jobs, ensuring timely execution of recurring data integration processes. Collaborated with senior engineers on ETL enhancement and modernization activities to improve scalability and operational efficiency. Wrote and optimized advanced Oracle SQL queries, stored procedures, functions, and database views to improve reporting performance and support business intelligence initiatives. Supported the design and maintenance of ETL pipelines using Python and SQL to integrate data from multiple enterprise systems into centralized reporting platforms. Conducted comprehensive data validation, reconciliation, and quality assurance checks to ensure data accuracy, completeness, and consistency across source and target systems. Created and maintained technical documentation, including ETL workflow documentation, source-to-target mapping specifications, and data dictionaries to support knowledge transfer, maintenance, and compliance. Worked closely with business analysts, data architects, and senior data engineers to understand business requirements and contribute to the delivery of scalable and reliable data engineering solutions. EDUCATION: Masters of Science in Business Analytics Trine University, Arizona. 2024 Bachelors of commerce in Accounting and Finance Osmania University, 2014 Keywords: continuous integration continuous deployment artificial intelligence machine learning business intelligence sthree Arizona Texas |