Home

Sri Datta - senior data engineer
[email protected]
Location: Dallas, Texas, USA
Relocation: Yes
Visa: GC
Resume file: Sri_Datta_Senior_Data_Engineer__1788285238124.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
NAME: SRI DATTA
Mobile: +1 737 334 8425
Senior Data Engineer
Email: [email protected]
LinkedIn: LinkedIn
Professional Summary:
Senior Data Engineer with 10+ years of experience building and supporting cloud data platforms, batch and streaming pipelines, data warehouses, and lakehouse solutions across AWS, Azure, and Google Cloud Platform.
Hands-on experience with Python, PySpark, Scala, SQL, Apache Spark, Databricks, Snowflake, BigQuery, Redshift, Azure Synapse Analytics, Hadoop, and relational and NoSQL databases.
Built end-to-end ETL and ELT pipelines for structured, semi-structured, and streaming data using Airflow, Cloud Composer, Azure Data Factory, AWS Glue, Informatica PowerCenter, Talend, Matillion, SSIS, and Sqoop.
Strong background in Spark performance tuning, partitioning, caching, data-skew analysis, query-plan review, cluster troubleshooting, and workload optimization across EMR, Dataproc, and Databricks.
Designed dimensional models, Star and Snowflake schemas, Data Vault structures, governed data lakes, warehouses, and Delta Lake lakehouse layers for analytics and reporting.
Implemented practical data quality, security, monitoring, recovery, and CI/CD controls using PyTest, Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, CloudWatch, Azure Monitor, RBAC, and encryption.
Worked closely with analysts, data scientists, application teams, security, QA, and business stakeholders in Agile delivery and production-support environments.
Core Competencies:
Expertise: Enterprise Data Engineering | Cloud Architecture | Data Platform Modernization | Lakehouse and Warehouse Architecture | Data Modeling
Expertise: ETL/ELT | Batch Processing | Streaming Analytics | Change Data Capture | API Integration | Data Migration | Database Engineering
Expertise: Data Governance | Data Quality | Metadata and Lineage | Security and Compliance | CI/CD | Infrastructure as Code | DataOps
Expertise: Technical Leadership | Solution Design | Code Reviews | Cross-Functional Delivery | Agile/Scrum | Production Operations
Technical Skills:
Cloud Platforms and Services: AWS | Azure | Google Cloud Platform | GCP | Amazon S3 | EC2 | IAM | VPC | Lambda | AWS Glue | EMR | Kinesis | Step Functions | Redshift | Redshift Spectrum | DynamoDB | CloudFormation | EKS | Azure Data Factory | ADF | ADLS Gen2 | Azure Databricks | Azure Synapse | Azure Event Hubs | Azure SQL Database | Cosmos DB | Azure Key Vault | Azure Monitor | Azure Log Analytics | Azure DevOps | AKS | BigQuery | Cloud Storage | Cloud SQL | Dataflow | Pub/Sub | Cloud Composer | Data Catalog | BigQuery ML | BQ-ML | Looker | Google Data Studio
Lakehouse, Warehouse and Analytics Platforms: Databricks | Delta Lake | Unity Catalog | MLflow | Auto Loader | Snowflake | Snowpipe | Streams and Tasks | Time Travel | Zero-Copy Cloning | Secure Data Sharing | Snowpark | Apache Iceberg | Teradata | Exadata
Programming and Distributed Processing: Python | PySpark | Pandas | Scala | SQL | T-SQL | PL/SQL | Shell Scripting | Apache Spark | Spark SQL | Spark Structured Streaming | Hadoop | HDFS | Hive | MapReduce | Impala | Pig | Oozie | Kafka | Flink
ETL, ELT and Orchestration: ETL | ELT | Apache Airflow | Astronomer | Matillion | Talend | dbt | Qlik Replicate | DataStage | SSIS | Sqoop | Tidal | REST APIs | CDC
Databases and Storage: PostgreSQL | MySQL | SQL Server | Oracle | MongoDB | Cassandra | Redis | Elasticsearch
Architecture and Modeling: data pipelines | batch processing | real-time processing | streaming | event-driven | data warehouse | data lake | lakehouse | data marts | Medallion Architecture | Bronze | Silver | Gold | Star Schema | Snowflake Schema | Data Vault 2.0 | dimensional modeling | OLTP | OLAP | data products
Governance, Quality and Security: data governance | data quality | metadata management | schema evolution | Great Expectations | GDPR | HIPAA | RBAC | KMS | Secrets Manager | encryption | security | compliance
BI and Delivery: Power BI | Tableau | SSRS | SAS | Erwin | Jira | Confluence | Agile | Scrum | technical leadership | code reviews | mentoring | stakeholder management
Technology Application Highlights:
Cloud-native ingestion and transformation using ADF, AWS Glue, Matillion, Cloud Composer, Dataflow, Airflow, dbt, Talend, Qlik Replicate, DataStage, and SSIS.
Real-time architecture using Kafka, Flink, Kinesis, Azure Event Hubs, Pub/Sub, Spark Structured Streaming, Snowpipe, CDC, and event-driven processing.
Lakehouse governance using Databricks, Delta Lake, Unity Catalog, Auto Loader, Delta Live Tables, Apache Iceberg, Medallion Architecture, metadata, lineage, and access controls.
Warehouse engineering using Snowflake, BigQuery, Redshift, Azure Synapse, Teradata, Oracle/Exadata, dimensional modeling, Data Vault 2.0, and semantic layers.
Platform automation using Terraform, CloudFormation, CI/CD, Jenkins, GitHub Actions, Azure DevOps, Docker, Kubernetes, PyTest, monitoring, logging, and alerting.
Professional Experience
Vanguard, PA Aug 2025 - Present
Senior Data Engineer
Responsibilities:
Build and optimize investment-data ELT pipelines with Python, PySpark, Spark SQL, Snowflake, and BigQuery for reporting, analytics, and downstream data products.
Tune Spark workloads by reviewing Spark UI metrics, partition sizes, shuffle behavior, data skew, caching, join strategies, and executor memory to keep scheduled pipelines within SLA targets.
Develop reusable Databricks notebooks and Delta Lake processing patterns for ingestion, transformation, validation, and curated lakehouse datasets.
Design Snowflake databases, schemas, tables, views, clustering strategies, and role-based access patterns for secure and maintainable analytical workloads.
Create Matillion and Spark ingestion frameworks for REST APIs, PostgreSQL, Oracle, MongoDB, files, Kafka topics, and cloud object storage, with audit columns, error handling, and restart capability.
Implement Spark Structured Streaming and Kafka pipelines for low-latency operational events and dependable downstream consumption.
Build GCP pipelines with BigQuery, Cloud Storage, Dataproc, Cloud Composer, Pub/Sub, and Dataflow for batch and near-real-time processing.
Develop and support Airflow DAGs in Cloud Composer, including dependencies, retries, backfills, alerting, and operational recovery.
Model analytical data with Star Schema and Data Vault techniques for BI dashboards, ad hoc analysis, and self-service reporting.
Package data services with Docker and deploy them to Kubernetes-based environments using repeatable configuration and release practices.
Maintain CI/CD pipelines in Jenkins, GitLab, GitHub Actions, and Azure DevOps for code review, automated testing, deployment, and version-controlled promotion across environments.
Monitor pipelines through CloudWatch and cloud-native logging, investigate failed jobs, document root causes, and coordinate production fixes with application and platform teams.
Apply RBAC, encryption, data quality checks, reconciliation controls, metadata, and lineage standards across AWS, Azure, GCP, Snowflake, and Databricks data assets.
Worked with investment-data owners and reporting teams to translate business rules into source-to-target mappings, reusable transformations, and clear acceptance criteria.
Added reconciliation checks, schema validation, exception logging, and rerun controls so failed records could be investigated without restarting an entire workflow.
Reviewed SQL and PySpark changes with other engineers, documented operating procedures, and supported releases through development, test, and production environments.
Improved Snowflake and BigQuery cost control by reviewing warehouse sizing, query history, partition pruning, clustering behavior, and recurring high-cost workloads.
Supported production incidents by tracing upstream dependencies, comparing source and target results, correcting data issues, and communicating recovery status to stakeholders.
Environment:
Python, PySpark, SQL, Apache Spark, Spark SQL, Databricks, Delta Lake, Snowflake, BigQuery, Cloud Storage, Dataproc, Cloud Composer, Pub/Sub, Dataflow, Kafka, Spark Structured Streaming, Airflow, REST APIs, relational databases, Docker, Kubernetes, Jenkins, GitLab CI/CD, CloudWatch, data modeling, RBAC, Agile/Scrum.

John Deere, Urbandale, Iowa May 2023 - Jul 2025
Senior Data Engineer
Responsibilities:
Built AWS data pipelines with S3, Glue, Lambda, EMR, Redshift, Athena, Talend, and Python to ingest and transform manufacturing, equipment, and supply-chain data.
Developed PySpark and Spark SQL jobs on EMR and Databricks, troubleshooting failed stages and tuning partitions, joins, storage formats, and cluster resources for stable daily processing.
Used Redshift Spectrum and optimized SQL to query large S3 datasets while keeping object storage as the primary data layer.
Integrated Snowflake with manufacturing and enterprise source systems and maintained secure databases, schemas, tables, stages, and governed data-sharing workflows.
Automated AWS Glue, Lambda, EC2, and S3 operations with Python and Boto3 and monitored workflows through CloudWatch alerts and logs.
Created Airflow workflows for ingestion, transformation, validation, dependency handling, and recovery, reducing manual coordination between teams.
Moved data from MySQL and other relational databases into HDFS and S3 using Sqoop and Glue, with reconciliation checks between source and target systems.
Supported streaming ingestion with Kinesis and Spark to make operational data available for near-real-time analysis.
Built dimensional models and Star Schema structures in Redshift and Snowflake for Tableau, QuickSight, and business reporting.
Used Glue crawlers, Pandas, and NumPy for metadata discovery, profiling, cleansing, and validation of incoming datasets.
Maintained CI/CD workflows with GitHub, Jenkins, CodePipeline, Systems Manager, and CloudTrail; tracked releases and defects in Jira and supported scheduled production jobs through Control-M.
Worked with analysts, operations teams, and database administrators to resolve data issues, tune SQL, and improve pipeline reliability.
Designed reusable ingestion patterns for CSV, JSON, Parquet, and relational sources, with consistent naming, metadata capture, validation, and error handling.
Optimized Redshift and Snowflake queries through execution-plan review, predicate pushdown, distribution and sort-key choices, and workload scheduling.
Maintained IAM roles, encryption settings, secrets, and least-privilege access for data pipelines operating across S3, EMR, Glue, Lambda, and Redshift.
Prepared deployment notes, support runbooks, data mappings, and troubleshooting guides that helped operations teams handle recurring failures and planned releases.
Partnered with database, reporting, QA, and platform teams during sprint planning, code review, testing, cutover, and post-release production support.
Environment:
AWS, S3, Glue, Lambda, EMR, EC2, Redshift, Redshift Spectrum, Athena, QuickSight, CloudWatch, Kinesis, Databricks, Snowflake, Python, Boto3, PySpark, Spark SQL, Hadoop, HDFS, MySQL, Sqoop, Airflow, Control-M, Tableau, Pandas, NumPy, GitHub, Jenkins, Jira.

Change Healthcare, TN Oct 2020 - Mar 2023
Data Engineer
Responsibilities:
Developed Azure Data Factory and Informatica PowerCenter pipelines for healthcare claims, operational, and reference data from relational databases and file-based sources.
Processed and transformed large datasets with Azure Databricks, PySpark, Spark SQL, and Delta Lake patterns for curated analytics layers.
Loaded governed data into Azure Synapse Analytics and optimized SQL, table design, distribution, and dimensional models for reporting workloads.
Managed Azure Data Lake Storage zones and file layouts to support dependable ingestion, retention, and downstream access.
Integrated Azure SQL Database and Cosmos DB sources into governed claims pipelines and applied validation and reconciliation checks before publishing curated data.
Designed Snowflake Schema, transactional, dimensional, and slowly changing dimension (SCD) models for consistent claims and finance reporting.
Protected sensitive healthcare data with Azure Key Vault, encryption, role-based access, and controlled Azure DevOps release processes.
Monitored ADF, Databricks, and Synapse workflows with Azure Monitor and Log Analytics, resolving failures and documenting production fixes in Jira.
Reviewed complex SQL and rebuilt maintainable transformation logic in Python where it improved testing, reuse, and supportability.
Produced Power BI dashboards and validated source-to-report results with business and finance teams before release.
Used R and SAS for statistical analysis, profiling, and investigation of trends and exceptions in healthcare datasets.
Built parameterized ADF pipelines and reusable Databricks transformations for incremental loads, dependency management, retry handling, and controlled reprocessing.
Created source-to-target mappings and reconciliation reports for claims data, resolving mismatched counts, invalid codes, duplicate records, and late-arriving files.
Tuned Synapse and SQL Server workloads through indexing, distribution choices, partitioning, query-plan analysis, and removal of unnecessary data scans.
Worked with security and compliance teams to apply access controls, encryption, secrets management, and audit-ready operating procedures for protected healthcare data.
Supported Agile delivery through estimation, peer review, test coordination, release validation, knowledge transfer, and production issue resolution.
Environment:
Microsoft Azure, Azure Data Factory, ADLS, Azure Databricks, Delta Lake, PySpark, Spark SQL, Azure Synapse Analytics, Azure SQL Database, Cosmos DB, Snowflake, Informatica PowerCenter, Python, SQL, R, SAS, Power BI, Key Vault, Azure Monitor, Log Analytics, Azure DevOps, Jira.
Schneider National, WI Jan 2019 - Sep 2020
Hadoop Developer
Responsibilities:
Developed Scala, Spark, PySpark, and Spark Streaming jobs for transportation and logistics data processing on Hadoop clusters.
Created internal and external Hive tables and wrote HiveQL for ingestion, validation, aggregation, and downstream reporting.
Moved data between relational databases, Netezza, Hive, and HDFS with Sqoop while validating row counts and daily load completeness.
Built DataFrames and RDD transformations with PySpark and Spark SQL and tuned queries used by reporting and operations teams.
Configured SSIS transformations, including Lookup, Derived Column, Data Conversion, Aggregate, and Conditional Split, for database ETL workflows.
Developed Pig and MapReduce jobs for data cleansing and large-volume batch processing on HDFS.
Designed HDFS-based ETL workflows covering acquisition, transformation, quality checks, storage, and scheduled delivery.
Supported Cassandra data processing and parallel metric calculations on the existing Hadoop platform.
Installed and configured Hive, Pig, Sqoop, Flume, and Oozie and maintained Oozie workflows for scheduled Hive and Pig jobs.
Monitored scheduled Hadoop and Oozie workloads, investigated failed jobs, corrected data and configuration issues, and restarted processing from safe recovery points.
Improved Hive and Spark performance by reviewing file sizes, partitions, joins, serialization formats, and query stages used in daily logistics reporting.
Validated data moving between Netezza, Hive, Cassandra, and HDFS through count checks, field-level comparisons, and exception reporting.
Maintained technical mappings, workflow documentation, installation notes, and production runbooks for Hadoop ecosystem components and scheduled jobs.
Worked with analysts and operations teams to clarify data requirements, investigate reporting differences, and deliver corrected datasets on schedule.
Environment:
Hadoop, HDFS, Apache Spark, Spark Streaming, Scala, Python, PySpark, Spark SQL, Hive, HiveQL, Cassandra, Netezza, Sqoop, Pig, MapReduce, Flume, Oozie, SSIS, relational databases.

Global Logic Technologies, Hyderabad, India Jun 2015 - Aug 2018
Data Analyst
Responsibilities:
Queried and joined data in MySQL and SQL Server, improving aggregation and stored-procedure logic used in recurring analysis and reporting.
Built SSIS packages with Lookup, Derived Column, Conditional Split, Aggregate, Pivot, Slowly Changing Dimension, Merge Join, and Union All transformations.
Developed stored procedures and datasets for drill-through, parameterized, tabular, matrix, linked, and cascading-parameter SSRS reports.
Created and published SSRS subscriptions and applied conditional formatting and validation rules to support consistent report delivery.
Analyzed large datasets in Excel with pivot tables, VLOOKUP, advanced formulas, and VBA automation to reduce repetitive manual work.
Performed data-quality checks for completeness, consistency, duplicates, and source-to-report accuracy before delivery.
Prepared clear business documentation in Microsoft Word and worked with stakeholders to explain findings, assumptions, and exceptions.
Participated in Agile planning and retrospectives and tracked requirements, issues, and delivery work in Jira.
Gathered reporting requirements from business users and converted them into SQL logic, dataset definitions, report layouts, and test scenarios.
Reviewed database tables, joins, indexes, and stored procedures to improve report response time and reduce repeated manual data preparation.
Created source-to-report validation queries and documented exceptions so analysts could trace reported values back to the underlying database records.
Coordinated user-acceptance testing, corrected report defects, and maintained release notes and support documentation for recurring business reports.
Helped standardize SSIS and SSRS development practices through reusable package patterns, naming conventions, peer review, and version-controlled changes.
Environment:
MySQL, SQL Server, SQL, T-SQL, SSIS, SSRS, stored procedures, dimensional reporting, Microsoft Excel, Pivot Tables, VLOOKUP, VBA, Microsoft Word, Jira, Agile/Scrum.

Education
Bachelor of Technology, Computer Science | 2011 - 2015
Keywords: continuous integration continuous deployment quality analyst machine learning user interface business intelligence sthree database active directory rlang information technology trade national procedural language Pennsylvania Tennessee Wisconsin

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7693
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: