| Deepa - Senior Data Engineer |
| [email protected] |
| Location: Bridgeton, North Carolina, USA |
| Relocation: Open |
| Visa: GC |
| Resume file: DEEPA_1785254485705.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
DEEPA
Senior Data Engineer +1 6234715460 Email: [email protected] Professional summary Senior Data Engineer with 12+ years of experience in data engineering, analytics, and reporting across Finance, Healthcare, and Insurance domains. Strong expertise in Python, SQL, PySpark, and cloud platforms, with a focus on data validation, data quality, and analytics reporting. Experienced in building data pipelines, performing data profiling, and ensuring data accuracy, along with developing interactive dashboards using Tableau, Power BI, and Databricks SQL. Proven ability to translate complex datasets into actionable insights to support business and sustainability initiatives. Strong leadership and mentoring skills; led data engineering teams and advocated best practices in code review, version control, and DevOps processes. In-depth knowledge of data governance, security and implementing data encryption, fine-grained access controls, and compliance with regulations (HIPAA, PCI-DSS, GDPR) across data solutions. Designed and implemented modern Lakehouse architectures leveraging Delta Lake for ACID-compliant data lakes and Unity Catalog for unified governance and data discoverability. Hands-on experience with Databricks Lakehouse Federation to enable federated queries across disparate data sources, improving data accessibility without duplicating data. Proficient in Databricks Delta Live Tables (DLT) for building automated, reliable data pipelines with quality checks, and using Databricks Asset Bundles to version and deploy data pipeline assets. Expertise in real-time and streaming data processing using Azure Event Hubs and Azure Databricks Structured Streaming to deliver timely data for analytics (e.g., financial transactions, IoT sensor data). Skilled in performance tuning of large-scale Spark jobs and SQL queries, optimizing data processing workflows to reduce runtime and cost (achieved up to 40% improvement in some use cases). Extensive experience in data warehousing and BI: designed star, snowflake schema data models, and developed enterprise data warehouses on Azure Synapse, Snowflake for analytics and reporting. Strong SQL development skills in T-SQL and PL/SQL; created complex stored procedures, user-defined functions, and optimized queries for data transformation and retrieval. Proven track record of modernizing legacy ETL processes (SSIS/Informatica) to cloud-based architecture, reducing maintenance overhead and increasing scalability. Implemented robust monitoring, logging, and alerting data pipelines using Azure Monitor and Log Analytics, enabling proactive issue detection and quick resolution to maintain data SLAs. Experience with Infrastructure-as-Code tools like Terraform and ARM/Bicep templates to provision Azure data infrastructure and enforce environment consistency in CI/CD pipelines. Adept at integrating data from diverse sources (relational databases, APIs, flat files, streaming data) into centralized data lakes, ensuring data quality and consistency through validation and cleansing. Collaborated closely with data scientists, analysts, and business stakeholders to deliver datasets for machine learning models, BI dashboards, and regulatory reports. Familiar with healthcare data standards (e.g., HL7/FHIR) and insurance data models, enabling effective handling of domain-specific data formats and ensuring regulatory compliance. Solid understanding of financial data management, including handling of PII and sensitive financial transactions with appropriate anonymization and audit trails. Demonstrated ability to implement data solutions under strict compliance and governance frameworks, aligning with corporate policies and industry standards at Fortune 100 companies. Strong problem-solving and analytical skills, with a history of troubleshooting complex data pipeline issues and optimizing system performance. Familiar with Dataiku DSS concepts and similar workflow-based tools, with the ability to quickly adapt and build end-to-end data pipelines. Technical Skills Data Analysis & Validation Data Profiling, Data Quality, Data Validation, Data Cleansing Azure Services: Azure Data Factory (ADF), Azure Databricks, Azure Synapse Analytics, Azure Data Lake Storage Gen2 (ADLS Gen2), Azure SQL Database, Azure Blob Storage, Azure Functions, Azure Key Vault, Azure Event Hubs, Azure Stream Analytics, Azure Monitor, Logic Apps, Azure DevOps, Microsoft Fabric Databases: Snowflake, Microsoft SQL Server (2012/2014/2016), Azure SQL DB, Azure Synapse Analytics (SQL Pools), Oracle 11g/12c, Redshift, MongoDB, Cosmos DB, PostgreSQL, MySQL, MS Access, MS Excel Big Data Technologies: Apache Hadoop (HDFS, MapReduce), Apache Spark (Scala, PySpark, Spark SQL, Spark Streaming), Delta Lake, Delta Live Tables (DLT), Kafka, Hive, Pig, Sqoop, Flume, Oozie, Zookeeper, Unity Catalog Hadoop Distribution: Cloudera, Hortonworks, Apache Hadoop Languages: Python, SQL, T-SQL, PL/SQL, Scala, Java, HiveQL, Shell Scripting, Bash, PowerShell Web Technologies: HTML, CSS, JavaScript, JSP, XML, RESTful APIs, SOAP Build Automation tools: Ant, Maven, Jenkins, Azure DevOps Pipelines, YAML Version Control: Git, GitHub, Bitbucket Methodology: Agile, Scrum, Waterfall IDE &Build Tools, Design: Visual Studio, Eclipse, IntelliJ IDEA, PyCharm, VS Code Operating Systems: Windows (XP/7/8/10), Linux, Ubuntu, CentOS, UNIX, macOS Visualisation Tableau, Power BI, Databricks SQL Others Analytical Thinking, Documentation, Business Rules Definition Certifications Microsoft Certified: Azure Data Engineer Associate (DP-203) Snowflake SnowPro Advanced: Data Engineer Databricks Certified Data Engineer Professional Professional Experiences Bank of America | Chicago, IL | Apr 2025 Present Senior Data Engineer Responsibilities: Conduct data profiling, validation, and quality checks to ensure accuracy and completeness of large-scale financial datasets. Develop interactive dashboards and reports using Power BI and Databricks SQL to support business insights and reporting needs. Translate complex financial and transactional datasets into actionable insights for risk analysis and decision-making. Document data definitions, transformation logic, and business rules to ensure transparency and consistency. Support testing and validation of data models built on Databricks to ensure data reliability. Collaborate with business stakeholders and analysts to deliver reporting solutions aligned with business goals. Implement automated data quality checks and validation frameworks within ETL pipelines. Built real-time streaming ETL pipelines using Databricks Delta Live Tables (DLT) to process digital banking transactions, card payments, and ATM activity, enabling near real-time fraud and risk analytics. Implemented advanced data quality controls within Delta Live Tables, including schema validation, constraint enforcement, and automated anomaly alerts to ensure financial data accuracy. Established centralized data governance using Databricks Unity Catalog, defining catalogs, schemas, and fine-grained access policies to manage secure access to sensitive financial datasets. Implemented column-level masking, encryption, and tokenization for confidential financial fields such as account numbers, customer identifiers, and transaction details. Enabled Databricks Lakehouse Federation to perform federated analytics across external enterprise systems including Oracle, SQL Server, and third-party financial APIs without duplicating datasets. Configured governance policies within Unity Catalog to enforce consistent data security controls across internal and federated datasets. Implemented Databricks Asset Bundles to package notebooks, jobs, and pipeline configurations into reusable deployment artifacts. Built end-to-end CI/CD pipelines using Azure DevOps integrated with Databricks Asset Bundles, automating unit testing, integration testing, and pipeline deployments across environments. Managed pipeline configurations as code using YAML-based infrastructure definitions, enabling automated validation and version-controlled deployments. Implemented CI/CD quality gates including code reviews, pipeline testing, and automated data validation using Great Expectations to maintain production-grade data reliability. Designed secure enterprise data pipelines integrating Azure Key Vault for secrets management and encryption standards to protect sensitive financial datasets. Implemented data anonymization and tokenization strategies to protect customer personally identifiable information (PII) in analytics and development environments. Enabled enterprise audit logging using Azure Log Analytics and Unity Catalog audit trails to support internal compliance reviews and regulatory audits. Optimized Spark workloads using Delta Lake file compaction, Z-Ordering, broadcast joins, and caching strategies, improving data processing performance for multi-terabyte financial datasets. Tuned Azure Databricks cluster configurations include autoscaling policies, spot instance utilization, and workload optimization, reducing infrastructure costs while maintaining SLAs. Implemented enterprise monitoring frameworks using Azure Monitor and Databricks metrics, enabling proactive detection of pipeline failures and system bottlenecks. Engineered scalable batch and streaming analytics pipelines processing millions of financial transactions daily across mobile banking, ATM networks, and card payment systems. Built curated analytics datasets in Delta Lake that unify data across retail banking, credit card, mortgage, and investment products for enterprise analytics. Collaborated with fraud analytics teams to integrate machine learning models into streaming pipelines using Databricks MLflow, improving fraud detection accuracy and reducing false positives. Implemented event-driven architectures using Azure Event Hubs and Spark streaming, enabling near real-time transaction monitoring for AML and fraud detection. Implemented full data lineage tracking and dataset versioning to support regulatory investigations, financial audits, and compliance reporting. Designed and implemented an enterprise Customer 360 data platform consolidating data across multiple banking systems to support customer behavior analytics and cross-sell opportunities. Implemented master data management logic within Customer 360 pipelines to unify customer identities across multiple financial systems. Enforced consumer privacy policies and regulatory compliance controls, including opt-out management and anonymization for restricted customer datasets. Enabled self-service analytics for business users by integrating curated Delta Lake datasets with BI platforms such as Power BI and Tableau. Implemented regulatory compliance checks within pipelines including data retention policies, archival automation, and regulatory reporting validations for CCAR, AML, and liquidity reporting. Collaborated with risk and compliance teams to ensure data pipelines align with Basel III, BCBS 239, and enterprise risk management frameworks. Led a team of data engineers delivering complex data platform initiatives within Agile development environments, conducting code reviews and mentoring engineers. Developed reusable engineering frameworks including Spark ETL libraries, CI/CD templates, and enterprise monitoring dashboards to standardize data engineering processes. Worked on workflow orchestration and data pipeline design for large-scale data platforms. Built ETL/ELT pipelines and performed data wrangling for structured and unstructured data. Drove adoption of modern data engineering capabilities including Unity Catalog governance, Delta Live Tables pipelines, and Lakehouse Federation through proof-of-concepts and engineering enablement. Communicated complex data platform architectures and project progress to both technical and executive stakeholders, aligning data initiatives with enterprise risk analytics and regulatory reporting strategies. Environment: Azure Data Factory, Azure Databricks, Delta Lake, Delta Live Tables (DLT), Unity Catalog, Azure Data Lake Storage Gen2, Azure Synapse Analytics, Azure Event Hubs, Azure Key Vault, Azure Monitor, Azure Log Analytics, Azure DevOps, Databricks Asset Bundles, Databricks Lakehouse Federation, Databricks MLflow, PySpark, Spark SQL, Python, SQL, Snowflake, Oracle, SQL Server, Power BI, Tableau, Great Expectations, Git, CI/CD, YAML, Agile, Linux, Windows. Pennsylvania Department of Transportation |PA |Aug 2023 Mar 2025 Azure Data Engineer Responsibilities: Performed data profiling and validation on large datasets to ensure data accuracy and consistency. Created dashboards and reports using Power BI and SQL-based tools for operational and analytical insights. Transformed complex transportation datasets into meaningful insights for business users. Documented data flows, assumptions, and business rules for improved data governance and reuse. Assisted in testing and validating data pipelines and data models developed in Databricks. Implemented data quality checks and validation processes within ETL workflows. Built automated data ingestion frameworks using Azure Data Factory to extract data from relational databases, APIs, flat files, and streaming sources into Azure Data Lake Storage Gen2. Developed robust ETL and ELT pipelines using Azure Databricks with PySpark and Spark SQL, transforming raw datasets into analytics-ready Delta Lake tables. Implemented Delta Lake architecture with bronze, silver, and gold layers to ensure progressive data refinement and support reliable analytics workloads. Developed incremental data ingestion pipelines using watermarking and Change Data Capture strategies to process daily data updates efficiently. Built scalable data transformation frameworks using PySpark notebooks and Databricks workflows, enabling high-performance distributed data processing. Integrated Snowflake data warehouse with Azure-based data pipelines to support enterprise analytics and optimized query performance through clustering and caching techniques. Implemented Databricks Unity Catalog to manage metadata, enforce access controls, and maintain data lineage across enterprise data platforms. Developed real-time data ingestion pipelines using Azure Event Hubs and Spark Structured Streaming to process streaming events and operational datasets. Implemented CI/CD pipelines using Azure DevOps, enabling automated deployment of Azure Data Factory pipelines and Databricks notebooks across development, testing, and production environments. Developed automated data validation frameworks using PySpark unit tests and data quality checks to ensure integrity and accuracy of data transformations. Collaborated with analytics teams and data scientists to deliver curated datasets supporting analytical models and reporting platforms. Optimized Spark processing workloads by applying partitioning strategies, broadcast joins, coaching, and query tuning, improving performance for large-scale datasets. Implemented monitoring and alerting frameworks using Azure Monitor and Log Analytics to track pipeline execution and quickly detect processing failures. Designed secure data pipelines by integrating Azure Key Vault for secrets management and encryption mechanisms. Migrated legacy on-premises ETL workflows and data warehouse systems to Azure cloud-based data platforms. Built reusable ETL frameworks and parameterized pipelines to simplify onboarding of new data sources. Enabled self-service analytics by integrating curated datasets with BI tools including Power BI and Tableau. Implemented data governance policies including data retention management, dataset classification, and secure access policies. Designed scalable data lake storage structures with standardized naming conventions and metadata tagging. Implemented robust error handling and retry mechanisms within pipelines to improve reliability and reduce pipeline failures. Developed streaming ETL proof-of-concepts to support near real-time analytics for operational datasets. Worked closely with infrastructure teams to configure secure networking, private endpoints, and VNet integrations for data services. Optimized cloud resource utilization by tuning cluster configurations and scheduling workloads efficiently. Participated in Agile Scrum ceremonies including sprint planning, backlog grooming, and sprint reviews to deliver incremental data engineering solutions. Provided production support for data pipelines and resolved issues to maintain high data availability. Created detailed technical documentation including pipeline design diagrams, data lineage documentation, and operational runbooks. Mentored junior engineers on Azure data engineering practices, PySpark development, and Databricks workflows. Built reusable PySpark utility libraries for logging, validation, and error handling to standardize ETL development. Continuously evaluated new Azure and Databricks capabilities to improve data platform scalability, reliability, and performance. Environment: Azure Data Factory, Azure Databricks, PySpark, Spark SQL, Delta Lake, Azure Data Lake Storage Gen2, Snowflake, Azure Event Hubs, Spark Structured Streaming, Azure DevOps, Azure Key Vault, Azure Monitor, Azure Log Analytics, Databricks Workflows, Python, SQL, Power BI, Tableau, Git, CI/CD, Agile, Linux, Windows. Cullen Frost Bankers, Inc San Antonio, TX | May 2019 July 2023 Data Engineer Responsibilities: Conducted data validation and cleansing to ensure high-quality healthcare and retail datasets. Developed dashboards using Power BI and Tableau to track key performance metrics. Translated large datasets into actionable insights for healthcare analytics and reporting. Documented data definitions and transformation logic for business and compliance needs. Supported testing and validation of data models and reporting datasets. Implemented large-scale data transformation pipelines using Azure Databricks with PySpark, integrating prescription records, clinic visit data, and insurance claim datasets into curated analytics tables. Designed and implemented Delta Lake based data processing layers to structure raw operational datasets into analytics-ready datasets for reporting and downstream data science workloads. Implemented data privacy and governance controls ensuring compliance with HIPAA and internal healthcare data protection policies, including PHI masking and secure data access controls. Designed and developed a healthcare analytics data mart using Azure Synapse Analytics, enabling reporting on prescription fulfillment, medication adherence, and patient care trends. Optimized ETL pipelines by implementing Spark performance tuning techniques such as caching, broadcast joins, and partition pruning, reducing batch processing runtimes by approximately 25%. Built real-time streaming pipelines using Azure Event Hubs and Spark Structured Streaming to process POS transactions and online order events for near real-time analytics. Integrated pharmacy data with retail sales datasets to generate insights on healthcare services and store-level performance trends. Implemented CI/CD pipelines using Azure DevOps to automate deployment of Azure Data Factory pipelines and Databricks notebooks across multiple environments. Developed automated data validation frameworks including schema checks, row count validations, and anomaly detection to ensure high-quality data across ETL pipelines. Migrated legacy ETL processes including SSIS packages and batch scripts to modern Azure cloud-based data pipelines using Data Factory and Databricks. Collaborated with analysts and business stakeholders to design new datasets and data pipelines supporting emerging healthcare analytics initiatives. Implemented secure credential management using Azure Key Vault and managed identities to protect sensitive secrets used within data pipelines. Created comprehensive documentation including pipeline architecture, data dictionaries, and operational runbooks to improve maintainability and knowledge sharing. Developed Power BI dashboards and curated datasets providing insights into prescription trends, patient demographics, and store performance metrics. Implemented cost optimization strategies including data lifecycle management, archival policies, and automated cluster shutdown schedules to reduce cloud infrastructure costs. Enabled self-service analytics capabilities by building logical views and curated data models within Synapse Analytics for business users. Participated in enterprise data governance initiatives including data classification, access policy definition, and data catalog integration to improve data trust and compliance. Provided production support for critical data pipelines including nightly prescription data loads, implementing monitoring and alerting frameworks to ensure reliable pipeline execution. Environment: Azure Data Factory, Azure Databricks, PySpark, Spark SQL, Delta Lake, Azure Data Lake Storage Gen2, Azure Synapse Analytics, Azure Event Hubs, Spark Structured Streaming, Azure DevOps, Azure Key Vault, Azure Monitor, Azure Log Analytics, Power BI, Python, SQL, Git, CI/CD, Agile, HIPAA Compliance, Linux, Windows. Walgreens | Dallas, TX Duration: Jan 2014 -OCT 2018 Data Engineer Responsibilities: Developed and maintained data processing scripts using Python with libraries such as Pandas and Numpy, ensuring efficient manipulation and analysis of large datasets. Managed and optimized Oracle databases, implementing effective schemas and queries to ensure high performance and reliability for enterprise applications. Implemented data integration solutions using SQLAlchemy to facilitate seamless connections and data manipulation between various databases and applications. Administered and optimized MySQL databases, ensuring data integrity, performance, and security to support transactional and analytical workloads. Designed and developed RESTful APIs using Flask (with Jinja2 templating and WSGI), enabling secure data services and microservice-based architectures. Automated build, testing, and deployment processes using Jenkins/Travis CI integrated with GitHub, streamlining CI/CD workflows and improving efficiency. Leveraged open-source big data processing frameworks like Spark and MapReduce to perform large-scale data processing tasks, optimizing performance and resource utilization for big data analytics. Keywords: continuous integration continuous deployment business intelligence database microsoft mississippi procedural language Illinois Pennsylvania Texas |