| Sai Teja - Sr Data engineer |
| [email protected] |
| Location: Austin, Texas, USA |
| Relocation: Yes |
| Visa: H1B |
| Resume file: Sai_Data Engineer_1786569672256.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
SAITEJA
Sr. Data Engineer [email protected] PROFESSIONAL SUMMARY Senior Data Engineer with 8+ years of experience designing and delivering scalable cloud-native data platforms, real-time streaming solutions, and AI-powered analytics across telecommunications, banking, financial services, insurance, and higher education domains. Extensive expertise in building high-performance ETL/ELT pipelines using Python, SQL, PySpark, Apache Spark, Databricks, Delta Lake, and dbt, processing terabytes to petabytes of structured and unstructured data. Strong hands-on experience with Azure Cloud, including Azure Data Factory, Azure Databricks, Azure Synapse Analytics, Azure Data Lake Storage Gen2, Microsoft Fabric, Azure Functions, Azure Event Hubs, Azure AI Search, Azure Key Vault, Azure Monitor, and Azure DevOps for enterprise-scale data engineering solutions. Experienced in developing AI-powered data applications using Azure OpenAI Service, LangChain, Retrieval-Augmented Generation (RAG), Azure AI Foundry, Prompt Flow, Vector Embeddings, MLflow, GitHub Copilot, and Hugging Face Transformers to enable intelligent search, semantic analytics, and AI-assisted decision-making. Proficient in designing real-time streaming architectures using Apache Kafka, Spark Structured Streaming, and Azure Event Hubs for high-volume event processing, telemetry analytics, fraud detection, and operational monitoring. Strong expertise in modern data warehousing and data modeling using Snowflake, Azure Synapse Analytics, Amazon Redshift, SQL Server, Delta Lake, Data Vault 2.0, and dimensional modeling to support enterprise reporting and advanced analytics. Hands-on experience implementing cloud infrastructure automation and DevOps using Terraform, GitHub Actions, Azure DevOps, ARM Templates, Docker, Jenkins, Git, and CI/CD pipelines, following Infrastructure as Code (IaC) best practices. Skilled in developing secure enterprise integrations using FastAPI, REST APIs, Postman, Apache Airflow, Logic Apps, and Azure Functions, enabling scalable workflow orchestration and seamless application integration. Expertise in data governance, metadata management, and regulatory compliance using Microsoft Purview, Unity Catalog, Collibra, Apache Atlas, Azure Key Vault, IAM, and RBAC, ensuring enterprise security, lineage, and governance standards. Strong analytical and database optimization skills with experience in SQL Server, Kusto Query Language (KQL), Cosmos DB, Power BI, Tableau, and advanced SQL tuning, improving reporting performance and operational insights. Experienced working in Agile/Scrum environments, collaborating with architects, product owners, business stakeholders, and cross-functional engineering teams using JIRA and Confluence, while mentoring junior engineers and driving engineering best practices. Proven ability to architect scalable, secure, and AI-enabled data ecosystems by leveraging Azure, AWS, Databricks, Snowflake, Apache Spark, Kafka, and modern DataOps/MLOps practices, delivering high-performance enterprise data solutions that support business intelligence, advanced analytics, and digital transformation initiatives. TECHNICAL SKILLS Programming Languages: Python, SQL, PySpark, Spark SQL Big Data Technologies: Apache Spark, Databricks, Delta Lake, Delta Live Tables, Apache Kafka, Spark Structured Streaming Cloud Platform: AWS, Azure Data Warehousing: Snowflake, Azure Synapse Analytics, Amazon Redshift, SQL Server ETL / ELT Tools: Azure Data Factory, Apache Airflow, dbt (Data Build Tool), Databricks Workflows Databases: SQL Server, Cosmos DB, Snowflake, Amazon Redshift AI / Generative AI: Azure OpenAI, LangChain, Retrieval-Augmented Generation (RAG), Azure AI Foundry, Prompt Flow, Vector Embeddings, Hugging Face Transformers, MLflow, GitHub Copilot Data Modeling: Dimensional Modeling, Star Schema, Snowflake Schema, Data Vault 2.0, Data Mart Design Streaming Technologies: Apache Kafka, Spark Structured Streaming, Azure Event Hubs, Azure Stream Analytics API & Integration: FastAPI, REST APIs, Postman, Azure Functions, Logic Apps Infrastructure as Code (IaC): Terraform, ARM Templates DevOps / CI-CD: Azure DevOps, GitHub Actions, Jenkins, Git, Docker Data Governance & Security: Microsoft Purview, Unity Catalog, Collibra, Apache Atlas, Azure Key Vault, AWS IAM, RBAC Business Intelligence: Power BI, Tableau Monitoring & Logging: Azure Monitor, Azure Log Analytics, Amazon CloudWatch, Kusto Query Language (KQL) Version Control: Git, GitHub Methodologies: Agile, Scrum, SDLC, CI/CD, DataOps, MLOps PROFESSIONAL EXPERIENCE Senior Data Engineer Client- EQUINIX California, USA Feb 2025 Present Designed high-performance distributed data pipelines using PySpark and Databricks, processing petabyte-scale network infrastructure and telemetry data with improved throughput and reliability. Developed enterprise AI-powered data solutions using Azure OpenAI Service, LangChain, and Retrieval-Augmented Generation (RAG) to enable intelligent document search and knowledge retrieval across structured and unstructured datasets. Architected enterprise-scale data platforms using Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse Analytics, and Azure Key Vault, enabling secure, scalable, and high-performance cloud data solutions. Developed robust ELT frameworks using SQL and Delta Lake, implementing incremental data loading, Change Data Capture (CDC), and ACID-compliant storage for enterprise reporting. Designed and optimized scalable PySpark and SQL-based data pipelines to process high-volume structured and semi-structured data, reducing end-to-end data processing time by 40% while improving data quality and reliability. Built real-time streaming architectures using Apache Kafka and Spark Structured Streaming to ingest and process network events, customer activity, and infrastructure logs. Automated cloud infrastructure and CI/CD deployments using Azure DevOps, ARM Templates, Microsoft Fabric, Azure Monitor, and Azure Log Analytics, improving deployment efficiency and operational visibility. Designed scalable analytical data models using Snowflake and dbt, enabling standardized transformations and self-service analytics across multiple business domains. Built intelligent data processing pipelines using Azure AI Foundry, Azure AI Search, Prompt Flow, and Vector Embeddings, enabling semantic search and AI-assisted analytics for business users. Implemented workflow orchestration using Apache Airflow integrated with Docker, automating end-to-end pipeline execution, monitoring, and recovery. Developed enterprise REST APIs using FastAPI and Postman, enabling secure integration between data platforms and internal business applications. Implemented infrastructure automation using Terraform and GitHub Actions, reducing manual deployment efforts and ensuring consistent Infrastructure as Code (IaC) practices. Established enterprise data governance using Microsoft Purview and Unity Catalog, improving metadata management, data lineage, security, and regulatory compliance. Optimized enterprise database performance using SQL Server and Kusto Query Language (KQL) to accelerate operational reporting and infrastructure monitoring. Integrated GitHub Copilot, MLflow, Hugging Face Transformers, and Python into data engineering workflows to accelerate code generation, experiment tracking, and AI model deployment. Led Agile delivery by collaborating with architects, product owners, and cross-functional engineering teams using JIRA and Confluence, mentoring junior engineers and driving best practices in data engineering. Senior Data Engineer Client- JP MORGAN CHASE New York, USA Sep 2023 Jan 2025 Designed enterprise-scale PySpark data pipelines on Databricks to process billions of financial transactions, improving data processing performance and scalability. Developed complex Python applications and optimized SQL queries to support regulatory reporting, fraud analytics, and customer insights. Implemented Apache Kafka streaming pipelines integrated with Spark Structured Streaming for real-time payment and transaction processing. Architected and deployed cloud-native data platforms using AWS S3, AWS Glue, AWS Lambda, and IAM, enabling secure, scalable, and automated enterprise data processing. Built scalable ELT workflows using Azure Data Factory integrated with Delta Lake to support incremental data ingestion and ACID-compliant storage. Implemented metadata-driven ETL/ELT frameworks using Azure Data Factory, Databricks, and Snowflake, enabling automated ingestion, transformation, and monitoring of enterprise data with minimal manual intervention. Designed enterprise data models using Snowflake and implemented Data Vault 2.0 architecture for historical tracking and audit compliance. Automated workflow scheduling and dependency management using Apache Airflow, integrating with dbt (Data Build Tool) for modular data transformations. Developed REST-based data integration services using FastAPI and managed API lifecycle through Postman for seamless application interoperability. Implemented DevOps best practices using GitHub Actions and Terraform, automating infrastructure provisioning and CI/CD deployments. Built high-performance analytical solutions using Amazon Redshift, Amazon EMR, Amazon Athena, CloudWatch, and Step Functions to optimize data warehousing, monitoring, and workflow orchestration. Established enterprise data governance by implementing Collibra and Apache Atlas for metadata management, data lineage, and compliance reporting. Collaborated with cross-functional Agile teams using JIRA and Confluence, conducting code reviews, sprint planning, and production support while mentoring junior data engineers. Data Engineer Trine University Jan 2023 Aug 2023 Engineered secure data integration solutions using Azure Data Factory and Azure Data Lake Storage Gen2 to consolidate student, faculty, and academic datasets. Built scalable analytics pipelines with Azure Databricks and Delta Live Tables to automate data transformation and improve processing efficiency. Designed and maintained enterprise data warehouses using Azure Synapse Analytics and Microsoft Fabric to support institutional reporting and analytics. Implemented real-time event processing using Azure Event Hubs and Azure Stream Analytics for monitoring student engagement and application events. Collaborated with solution architects, business stakeholders, and cross-functional engineering teams to design cloud-native data solutions, ensuring secure, high-performance, and scalable data platforms that supported enterprise analytics and reporting. Developed serverless data processing workflows with Azure Functions and Logic Apps to automate data validation and notification processes. Configured metadata management and governance using Microsoft Purview and Azure Key Vault to ensure data security, lineage, and compliance. Optimized database performance through Cosmos DB and Azure Monitor, improving query response times and proactively identifying operational issues. Managed infrastructure deployment and environment provisioning using Terraform and Docker, enabling consistent, repeatable, and scalable cloud deployments. Data Engineer TCS (PRUDENTIAL FINANCIAL) Jan 2019 June 2022 Developed scalable ETL pipelines using Python and SQL to ingest, transform, and process insurance policy and claims data for enterprise analytics. Built distributed data processing jobs with PySpark and Apache Spark to handle large-scale datasets and improve batch processing performance. Designed cloud-based data integration solutions using AWS S3 and AWS Glue to automate secure data ingestion and transformation workflows. Created dimensional data models and optimized warehouse performance using Snowflake and Amazon Redshift for reporting and analytics. Orchestrated end-to-end data pipelines using Apache Airflow and Apache Kafka to support reliable workflow scheduling and real-time data streaming. Developed interactive business intelligence dashboards using Power BI and Tableau to visualize operational KPIs and executive metrics. Implemented automated deployments and version control using Git and Jenkins, ensuring consistent CI/CD practices across development environments. Performed data validation and pipeline optimization using Databricks and Delta Lake to improve data quality, reliability, and query performance. Data Engineer L&T (CITI BANK) Apr 2018 Dec 2018 Designed, developed, and maintained scalable ETL pipelines using SQL and Python to ingest, transform, and load banking data from multiple source systems into enterprise data warehouses. Developed complex SQL queries, stored procedures, views, and performance tuning techniques to support financial reporting, regulatory compliance, and business analytics. Built automated data validation and reconciliation processes to ensure data accuracy, completeness, and consistency across banking applications. Integrated data from core banking, customer, transaction, and payment systems using batch processing and scheduled ETL workflows. Collaborated with business analysts, data architects, and application teams to gather requirements and deliver data solutions aligned with banking business needs. Implemented data quality checks, exception handling, logging, and monitoring to improve pipeline reliability and reduce data processing failures. Supported reporting teams by developing optimized datasets and database objects for dashboards using Power BI and Tableau, enabling faster business decision-making. Participated in production deployments, performance optimization, troubleshooting, and root cause analysis while following Agile methodology and version control best practices. CERTIFICATIONS AWS Certified Data Engineer Associate | Amazon Web Services (AWS) | 2024 Google Professional Data Engineer | Google Cloud (GCP) | 2023 Databricks Certified Associate Developer for Apache Spark | Databricks | 2023 EDUCATION Master of Science in Business Analytics: Trine University, Detroit, Michigan. Bachelor of Engineering in Computer Science: JNTUH, India Keywords: continuous integration continuous deployment artificial intelligence business intelligence sthree database |