| William Burston - Sr. Data Engineer |
| [email protected] |
| Location: Dallas, Texas, USA |
| Relocation: Yes |
| Visa: USA Citizen |
| Resume file: WILLIAM BURSTON_1786546666719.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
WILLIAM BURSTON
SENIOR DATA ENGINEER EMAIL - [email protected] PROFESSIONAL SUMMARY Senior Data Engineer with 7+ years of experience progressing from Python/software development into data engineering, big data, cloud data platforms, and enterprise analytics across financial services, healthcare, life sciences, telecom, and technology environments. Strong hands-on experience designing and supporting end-to-end ETL/ELT pipelines, data lakehouses, data warehouses, batch and streaming workflows, with a focus on reliable and scalable data delivery. Advanced experience with Python, SQL, PySpark, Apache Spark, Azure, AWS, Snowflake, Databricks, Azure Data Factory, Airflow, and Synapse, building data solutions across both cloud and hybrid environments. Experienced in building Bronze/Silver/Gold and Delta Lake architectures, integrating structured and semi-structured data, and developing curated datasets for analytics, reporting, regulatory, and operational use cases. Strong background in data quality, reconciliation, validation, performance tuning, monitoring, and production support, with ownership of pipelines from development and testing through deployment and troubleshooting. Hands-on experience with real-time data processing using Kafka, Azure Event Hubs, and Spark Structured Streaming, including windowing, watermarking, checkpointing, and stateful processing. Experienced in implementing CI/CD, Infrastructure as Code, cloud security, and automated deployments using Azure DevOps, Terraform, ARM templates, Git, Jenkins, GitHub Actions, Key Vault, IAM, RBAC, and managed identities. Comfortable working directly with analysts, data scientists, application developers, architects, compliance teams, and business stakeholders to translate requirements into practical data solutions and resolve production issues. Proven ability to take ownership of complex data initiatives, including AML data modernization, healthcare EHR/EMR integration, life-sciences analytics, telecom data platforms, cloud migration, and backend/API development, while mentoring team members and contributing in Agile environments. EDUCATION The University of Texas at Dallas (UT Dallas), Richardson, TX Master of Science in Computer Science | 2017 - 2019 Relevant Coursework: Data Mining, Database Systems, Cloud Computing, Big Data Analytics, Distributed Systems University of Texas at Arlington (UTA), Arlington, TX Bachelor of Science in Computer Science | 2013 - 2017 Focus: Programming, Data Structures & Algorithms, Database Systems, Software Engineering, Computer Networks SKILLS Category Skills Programming Python, SQL, PySpark, Scala, Bash, Pandas Data Engineering ETL/ELT, Data Integration, Data Pipelines, Data Transformation, Data Validation, Data Migration Big Data & Streaming Apache Spark, PySpark, Hadoop, Hive, Kafka, Spark Structured Streaming Azure Azure Data Factory, Databricks, ADLS Gen2, Synapse Analytics, Event Hubs, Key Vault, Azure Monitor, Entra ID AWS S3, Glue, Redshift, RDS, EC2, Lambda, IAM, Step Functions Data Warehousing Snowflake, Synapse, Redshift, Dimensional Modeling, Star Schema, Data Lakehouse, Delta Lake Databases PostgreSQL, SQL Server, Oracle, MySQL, MongoDB, Teradata ETL & Orchestration Airflow, Azure Data Factory, AWS Glue, dbt, Databricks Workflows, AutoSys Data Quality & Governance Great Expectations, dbt Tests, Data Profiling, Schema Validation, Data Lineage, Data Governance Performance Partitioning, Partition Pruning, Broadcast Joins, Caching, AQE, Query Optimization, Z-Ordering DevOps & Infrastructure Git, GitHub, Azure DevOps, Jenkins, GitHub Actions, Docker, Kubernetes, Terraform, CI/CD APIs & Backend REST APIs, Flask, Django, JSON, API Integration Monitoring & Support Azure Monitor, Dynatrace, Grafana, Prometheus, Splunk, ServiceNow, RCA, Production Support BI & Methodologies Power BI, Excel, KPI Reporting, Agile, Scrum, SDLC, Stakeholder Collaboration Bank of America - Senior Data Engineer | Jan 2025 - Present Own the design, development, and production support of enterprise data engineering solutions for the Global Financial Crimes (GFC) organization, supporting AML, KYC, CDD, transaction monitoring, regulatory reporting, and financial crime analytics. Designed and developed scalable ETL/ELT pipelines using Azure Data Factory, leveraging Copy Activities, Mapping Data Flows, Lookup, Stored Procedure, Linked Services, Integration Runtimes, and parameterized pipelines to ingest high-volume transaction, customer, account, and regulatory data from multiple banking systems. Built and optimized distributed data processing workloads using Azure Databricks, PySpark, Python, and Spark SQL, processing billions of records across ADLS Gen2, Delta Lake, and Azure Synapse Analytics while applying partitioning, broadcast joins, caching, Adaptive Query Execution (AQE), and partition pruning to improve pipeline performance. Developed a modern Azure Data Lakehouse architecture on ADLS Gen2 using Bronze, Silver, and Gold layers, establishing reliable raw-data ingestion, automated cleansing and standardization, and curated datasets for AML investigations, compliance analytics, risk reporting, and downstream consumption. Implemented Delta Lake pipelines with schema enforcement, schema evolution, MERGE operations, incremental processing, compaction, OPTIMIZE, Z-Ordering, and retention management to improve data reliability and query performance across large financial datasets. Supported the modernization of legacy AML data pipelines by performing source-to-target mapping, data profiling, transformation design, reconciliation, and migration of Teradata-based workloads into Azure-based data platforms while maintaining business and regulatory requirements. Developed real-time and near-real-time streaming pipelines using Azure Event Hubs, Kafka, Spark Structured Streaming, windowing, watermarking, and stateful aggregations to support transaction monitoring, suspicious activity detection, and time-sensitive financial crime analytics. Designed analytical data models and optimized Azure Synapse Analytics workloads, including dedicated SQL pools, external tables, PolyBase-based ingestion, SQL transformations, stored procedures, views, and dimensional models supporting Retail Banking, Risk, Compliance, and Financial Crimes reporting. Automated complex data workflows and dependencies using Azure Data Factory, Databricks Workflows, and AutoSys, implementing scheduling, dependency management, parameterization, restartability, error handling, and operational controls to ensure reliable delivery within defined SLAs. Established automated data quality and reconciliation frameworks using Python and SQL to validate record counts, schema changes, duplicate records, null patterns, transaction totals, source-to-target balances, and critical AML attributes before promoting datasets between environments. Implemented enterprise CI/CD and DevOps practices using Azure DevOps, Git, and YAML pipelines across DEV, UAT, QA, and Production environments; automated deployment of ADF pipelines, Databricks notebooks, SQL objects, and configuration changes while maintaining controlled release processes. Built reusable Infrastructure as Code (IaC) using Terraform and ARM templates to provision and manage Azure Data Factory, Databricks, ADLS Gen2, Synapse, Event Hubs, Key Vault, networking, and supporting cloud resources consistently across environments. Strengthened security and governance across financial data platforms using Azure Key Vault, Managed Identities, RBAC, Azure AD/Entra ID, encryption, secrets management, access controls, data classification, and audit logging, ensuring sensitive customer and transaction data remained protected and traceable. Took ownership of production operations and L2/L3 support, using Azure Monitor, Log Analytics, Dynatrace, Grafana, Prometheus, and Splunk to monitor pipeline health and data-processing workloads; investigated failures, performed root-cause analysis, implemented permanent fixes, and participated in post-incident reviews through ServiceNow. Work closely with AML analysts, compliance teams, data architects, product owners, and other engineers in an Agile/Scrum environment, translating regulatory and business requirements into technical solutions and contributing to code reviews, design discussions, documentation, and knowledge sharing while mentoring junior engineers and helping newer team members understand enterprise data engineering practices. AbbVie - Big Data Engineer | Nov 2023 Dec 2024 Designed and maintained scalable big data pipelines supporting life sciences research, analytics, and reporting, taking ownership of data ingestion, transformation, validation, and delivery across AWS, Snowflake, and Databricks environments. Built end-to-end ETL/ELT workflows using Python, SQL, Apache Airflow, and AWS Glue, integrating data from SQL Server, PostgreSQL, APIs, files, and other enterprise sources into Amazon S3 and Snowflake for centralized analytics. Developed distributed data processing pipelines using PySpark and Databricks to transform large research and operational datasets, applying partitioning, caching, optimized joins, and incremental processing to improve throughput and reduce processing time. Engineered AWS Glue jobs to ingest and transform data from Amazon S3, implemented reusable transformation logic, and loaded curated datasets into Snowflake while using IAM roles, AWS Lambda, and event-based triggers to automate and secure pipeline execution. Designed and optimized Snowflake databases, schemas, tables, views, stored procedures, and materialized views for analytics workloads; tuned complex SQL queries and transformations, reducing query execution time by approximately 50% for frequently used datasets. Built reliable Airflow DAGs with scheduling, dependencies, retries, failure handling, and monitoring to automate recurring data workflows; developed Python-based validation and cleansing processes using Pandas to handle missing values, duplicates, schema inconsistencies, and data-quality issues. Led the migration of on-premises SQL Server and PostgreSQL datasets to Amazon Redshift, using Informatica for extraction and transformation and validating approximately 120M production records through row-count, primary-key, and hash-based reconciliation checks. Developed data mapping and integration frameworks to bring together large datasets from multiple source systems, establishing consistent business rules and standardized transformations that improved downstream reporting accuracy and reduced manual data preparation. Built and maintained curated datasets powering Spotfire and Power BI dashboards used by research and business teams, reducing manual reporting effort by approximately 40% and improving visibility into key operational and research metrics. Contributed to backend data applications using Python and Flask, developing APIs and database integrations that allowed internal applications and analytics users to access validated datasets and data-processing services. Supported an intelligent document-processing workflow using AWS Textract and downstream extraction services, developing data-mapping templates, validating extracted fields, and creating Snowflake views consumed by analytics and Spotfire dashboards. Worked closely with data scientists, research teams, analysts, application developers, and business stakeholders in an Agile environment, participating in design discussions, code reviews, production troubleshooting, documentation, and continuous improvement while taking ownership of data pipeline reliability and delivery. HCA Healthcare - Data Engineer | Sep 2022 - Oct 2023 Designed and supported core ETL/ELT pipelines, databases, and data warehouse solutions for healthcare data, integrating EHR/EMR records, clinical encounters, patient, provider, claims, laboratory, pharmacy, and operational data from multiple hospital and enterprise source systems. Built scalable data ingestion and transformation workflows using Python, SQL, Apache Airflow, Azure Data Factory, and dbt, automating the movement of high-volume healthcare data into Azure Data Lake Storage Gen2, Azure Synapse Analytics, and Databricks. Developed PySpark and Spark SQL transformation jobs in Databricks to process large clinical datasets, applying partitioning, optimized joins, caching, and incremental processing to improve pipeline performance and support downstream analytics. Established a Bronze/Silver/Gold data architecture using ADLS Gen2 and Delta Lake, landing raw EHR/EMR data in the Bronze layer, standardizing and validating it in Silver, and publishing curated healthcare datasets and dimensional models in Gold for reporting and analytics. Built and maintained Azure Synapse Analytics data warehouse objects, including tables, views, stored procedures, external tables, and dimensional models, supporting clinical reporting, operational analytics, and healthcare performance dashboards. Developed near-real-time ingestion pipelines using Apache Kafka and Spark Structured Streaming for selected clinical and operational events, implementing checkpointing, windowing, watermarking, and error-handling mechanisms to provide reliable downstream data delivery. Implemented healthcare-focused data quality and reconciliation processes using Python, SQL, dbt tests, and Great Expectations to validate schemas, record counts, null values, duplicates, referential integrity, and source-to-target balances across EHR/EMR integrations. Designed secure data pipelines for HIPAA-sensitive PHI, using Azure Entra ID, RBAC, Azure Key Vault, Managed Identities, encryption, private endpoints, and access-controlled storage to protect patient and clinical information while maintaining auditability and appropriate data access. Automated pipeline scheduling, dependencies, retries, monitoring, and failure recovery through Airflow and Azure Data Factory, taking ownership of production issues, performing root-cause analysis, and resolving ETL failures to maintain reliable data delivery. Implemented CI/CD and Infrastructure as Code practices using Azure DevOps, Git, Terraform, and Docker, supporting repeatable deployments of data pipelines and cloud resources across development, testing, and production environments. Partnered with clinical analysts, data scientists, application teams, and business stakeholders to translate EHR/EMR and healthcare reporting requirements into reliable data models and pipelines, while participating in Agile ceremonies, code reviews, technical documentation, and production support. AT&T - Data Engineer | Data Analyst (Remote) | July 2020 - Aug 2022 Progressed from data analysis into hands-on data engineering, working with SQL, Python, AWS, ETL, and reporting to turn large telecom datasets into reliable data products and actionable business insights. Analyzed high-volume telecom customer, billing, network, usage, and operational datasets using advanced SQL, Python, and R to identify trends, investigate root causes, and support data-driven decisions across business and operations teams. Designed and maintained ETL pipelines using Python, PySpark, Apache Spark, AWS Glue, Sqoop, and Kafka, integrating data from relational databases, files, APIs, and streaming sources into Amazon S3, Snowflake, and Amazon RDS. Built an AWS-based data platform using Amazon S3, EC2, Lambda, RDS, and Snowflake, organizing raw and processed datasets into structured data layers that supported both operational reporting and analytics workloads. Developed automated data cleansing, validation, and transformation processes for multi-source datasets, implementing checks for duplicates, missing values, schema inconsistencies, data types, and source-to-target reconciliation before data was made available for reporting. Created and maintained Power BI dashboards and KPI reports covering customer activity, service performance, operational metrics, and business trends, while automating recurring Excel reports to reduce manual reporting effort and improve reporting consistency. Designed and optimized Snowflake and Hive data models, SQL queries, tables, and schemas for analytical workloads; also worked with Oracle and MongoDB to extract and analyze structured and semi-structured data from different source systems. Automated data workflows and cloud infrastructure using AWS Step Functions, Terraform, and Lambda, while implementing CI/CD practices with Jenkins, GitHub Actions, and Git to improve deployment consistency and maintain version-controlled data engineering code. Worked closely with analysts, business stakeholders, and engineering teams in an Agile/Scrum environment, gathering reporting requirements, translating them into technical solutions, tracking work through JIRA, monitoring production workloads with New Relic, and taking ownership of data issues from investigation through resolution. DoorDash - Python Developer | Software Engineer | May 2019 - June 2020 Developed and maintained backend services using Python, Django, Flask, and REST APIs, supporting internal and customer-facing applications with a focus on clean, maintainable, and reliable code. Built reusable Python automation scripts and backend utilities to streamline operational workflows, data processing, validation, and routine engineering tasks, reducing manual effort for development and support teams. Designed and integrated RESTful APIs for exchanging order, delivery, restaurant, customer, and operational data between internal services, implementing request validation, error handling, logging, and authentication controls. Worked extensively with SQL and PostgreSQL, writing complex queries, joins, stored procedures, and data-access logic to support application features, operational reporting, and backend workflows; investigated slow queries and improved database performance through indexing and query optimization. Developed data-processing components using Python, Pandas, and SQL to clean, transform, validate, and prepare high-volume operational datasets used by engineering and analytics teams. Integrated third-party and internal services through REST/JSON APIs, handling asynchronous workflows, retries, timeouts, exception handling, and structured logging to make backend integrations more resilient. Used Git, GitHub, Docker, Jenkins, and Linux as part of the development lifecycle, participating in code reviews, unit testing with PyTest, CI/CD deployments, debugging, and production issue resolution while following Agile/Scrum practices. Worked closely with senior engineers, product managers, QA, and data teams to understand requirements and deliver backend features, taking ownership of assigned services from development and testing through deployment and production support while building a strong foundation in Python software engineering, APIs, SQL, distributed systems, and cloud-based applications. Keywords: continuous integration continuous deployment quality analyst business intelligence sthree active directory rlang information technology Idaho Texas Utah |