| Navya Sri - Data Engineer / AI/ML |
| [email protected] |
| Location: Aurora, Illinois, USA |
| Relocation: yes |
| Visa: H1B |
|
Navya Sri. B
Data Engineer | : (309) 332-5836 | : [email protected] Senior Data Engineer with 8+ years of experience engineering scalable cloud-native data platforms across AWS, Azure, and GCP for enterprise analytics, AI/ML, and advanced analytics workloads. Expert in Python, SQL, Apache Spark, PySpark, Scala, ETL/ELT, distributed processing, and modern data lakehouse, data warehouse, and data platform architectures. Hands-on expertise delivering high-throughput batch and real-time streaming pipelines using Kafka, Spark Structured Streaming, Pub/Sub, Dataflow, Databricks, AWS EMR, Azure Databricks, Delta Lake, and CDC patterns. Proven ability to build reliable ingestion, transformation, cleansing, enrichment, validation, and orchestration frameworks for multi-terabyte datasets. Strong background in Databricks Lakehouse, Delta Lake, Unity Catalog, metadata management, data lineage, data quality, governance, security, RBAC, IAM, encryption, and performance optimization. Experienced with Spark tuning, partitioning, caching, predicate pushdown, broadcast joins, AQE, Z-Ordering, schema evolution, fault tolerance, and operational troubleshooting. Advanced AI/ML data platform experience supporting feature stores, feature engineering, MLOps, model training and deployment pipelines, MLflow, Vertex AI, Azure Machine Learning, GenAI/RAG workloads, and collaboration with Data Scientists, ML Engineers, Architects, Product teams, and DevOps. Strong cloud warehouse expertise across BigQuery, Snowflake, Redshift, and Synapse, with production CI/CD and DataOps practices. Experienced in secure, production-grade data engineering with Git, Jenkins, GitHub Actions, Docker, Kubernetes, Terraform, Airflow, and Cloud Composer. Delivers high-performance, governed platforms with strong data integrity, scalability, reliability, cost efficiency, monitoring, and automated deployment across enterprise environments. Proven expertise in designing modern data lakehouse architectures, implementing scalable ETL/ELT frameworks, optimizing distributed data processing workloads, and delivering production-grade data platforms that power enterprise analytics and AI solutions. Experience working with metadata management, data governance practices, data lineage, and data quality frameworks across enterprise data platforms. Strong experience working with cloud data warehouses including BigQuery, Snowflake, Amazon Redshift, and Azure Synapse Analytics, designing dimensional data models that support large-scale analytics workloads. Developed real-time and batch data pipelines using distributed processing frameworks (Apache Spark/Kafka), with experience in streaming data processing and scalable data architectures. Collaborates closely with data scientists, analysts, product teams, and platform engineers to deliver reliable, scalable, and high-performing data platforms that enable data-driven decision making and machine learning innovation. Experience working with financial data platforms and supporting enterprise reporting, reconciliation, and data workflows in financial services environments. Experienced in implementing DevOps best practices using Git, Jenkins, Docker, Kubernetes, and CI/CD pipelines to support reliable and automated deployment of data engineering solutions. Familiar with cybersecurity best practices including secure data access, IAM-based controls, encryption, governance standards, compliance requirements, and secure cloud data platform design within enterprise financial environments. Cloud & Data Platforms: AWS (S3, Glue, EMR, Redshift, Lambda, Athena, Kinesis, IAM, Lake Formation), Azure (Azure Databricks, ADLS Gen2, Synapse, Azure Data Factory, Event Hubs, Azure ML, Key Vault, Entra ID), GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Cloud Composer, Vertex AI, Cloud Storage) Big Data & Streaming: Apache Spark, PySpark, Spark SQL, Spark Structured Streaming, Scala, Apache Kafka, Kafka Connect, Hadoop, Hive, Apache Beam, CDC, Event-Driven Architecture Data Engineering & Architecture: ETL/ELT, Data Lake, Lakehouse, Data Warehouse, Delta Lake, Delta Live Tables, Unity Catalog, Medallion Architecture, Data Modeling, Dimensional Modeling, Data Integration, Data Quality, Data Validation, Metadata Management, Data Lineage, Data Governance, Schema Evolution Programming: Python, Scala, Java, SQL, Spark SQL Warehousing & Databases: Snowflake, BigQuery, Amazon Redshift, Azure Synapse, PostgreSQL, MongoDB, Redis, MySQL, Oracle AI/ML & MLOps: Feature Store, Feature Engineering, MLOps, MLflow, Vertex AI, Azure Machine Learning, Model Training, Model Deployment Pipelines, Model Serving, GenAI, LLMOps, RAG, LangChain, LangGraph, Vector Embeddings, Vector Search, LLM Evaluation Orchestration & DataOps: Apache Airflow, Cloud Composer, Azure Data Factory, Databricks Workflows, Jenkins, GitHub Actions, Git, CI/CD, Terraform, Docker, Kubernetes, Cloud Build Performance & Reliability: Spark Performance Tuning, Partitioning, Caching, Broadcast Joins, Predicate Pushdown, Adaptive Query Execution, Z-Ordering, File Compaction, Cluster Optimization, Monitoring, Fault Tolerance, Cost Optimization Security & Governance: Unity Catalog, RBAC, IAM, Lake Formation, Azure Key Vault, Microsoft Entra ID, Data Encryption, Secure Data Access, Auditability, Compliance APIs & Microservices: REST APIs, FastAPI, Spring Boot Operating Systems: Linux / Unix Financial Data & Systems: Financial transaction processing, reporting, reconciliation, GL, AP, AR Financial Data & Systems: Financial data processing, reporting, and reconciliation workflows Experience working with transaction-based datasets and financial analytics Familiarity with financial operations concepts (GL, AP, AR, reconciliation) PROFESSIONAL EXPERIENCE ________________________________________ FIS (Fidelity National Information Services) Senior Data Engineer Data Governance & SQL USA | Jul 2023 Present Company Overview: FIS (Fidelity National Information Services) is a global financial technology provider delivering banking infrastructure, payment processing platforms, and financial analytics solutions used by financial institutions and fintech companies worldwide. Project Goal: Designed and implemented a scalable cloud-based enterprise data platform capable of processing high-volume financial transaction datasets, enabling real-time analytics, fraud detection, and enterprise reporting across distributed cloud environments. Key Contributions Designed and implemented scalable cloud-native data platforms using Python, SQL, PySpark, Scala, Apache Spark, Azure Databricks, AWS EMR, and GCP services to process multi-terabyte financial datasets for analytics, AI/ML, fraud detection, and enterprise reporting. Built batch and real-time ingestion architectures using Apache Kafka, Kafka Connect, Spark Structured Streaming, Google Pub/Sub, Dataflow, CDC, watermarking, checkpointing, schema evolution, and fault-tolerant processing for high-volume transaction events. Architected Databricks Lakehouse solutions using Delta Lake, Delta Live Tables, Unity Catalog, and Medallion Architecture across Bronze, Silver, and Gold layers; implemented metadata management, lineage, access controls, data quality, and governed data products. Developed production ETL/ELT pipelines for ingestion, transformation, cleansing, enrichment, validation, deduplication, and incremental loading from relational databases, APIs, files, and streaming sources into BigQuery, Snowflake, Redshift, and Synapse. Implemented CDC and incremental ingestion using Delta Lake MERGE/UPSERT patterns, Spark Structured Streaming, Kafka, checkpoints, watermarks, and Databricks Workflows to maintain reliable near-real-time datasets. Optimized Spark and Databricks workloads through partition strategy, broadcast joins, caching, predicate pushdown, Adaptive Query Execution, Z-Ordering, file compaction, cluster tuning, and workload/resource optimization. Built and maintained AI/ML data platform capabilities including feature engineering, Feature Store integrations, MLOps workflows, MLflow tracking, model training datasets, model deployment pipelines, and production model-serving data flows supporting Data Science and ML teams. Integrated Vertex AI, Azure Machine Learning, MLflow, GenAI/RAG pipelines, vector search, LangChain, and LangGraph with governed enterprise data platforms to enable machine learning and advanced analytics workloads. Implemented data quality and validation frameworks using automated profiling, reconciliation, completeness, accuracy, consistency, schema checks, and exception monitoring to protect data integrity across critical pipelines. Established security and governance controls using Unity Catalog, RBAC, IAM, AWS Lake Formation, Azure Key Vault, Microsoft Entra ID, encryption, auditability, lineage, and compliance-aligned access patterns. Orchestrated multi-stage pipelines with Apache Airflow, Cloud Composer, Azure Data Factory, and Databricks Workflows; automated dependency management, retries, monitoring, scheduling, and operational recovery. Built CI/CD and DataOps automation using Git, Jenkins, GitHub Actions, Terraform, Docker, Kubernetes, and Cloud Build for repeatable deployment of data pipelines, Databricks workloads, and ML platform components. Troubleshot production failures across Spark jobs, Kafka streams, CDC processes, data quality checks, schema changes, cluster resources, and orchestration workflows while improving reliability, scalability, observability, and cost efficiency. Partnered with Data Scientists, ML Engineers, Data Architects, DevOps, Product Owners, analysts, and business stakeholders to translate requirements into secure, high-performance data products and AI/ML-ready datasets. JMPC Data Engineer USA | Jul 2022 Jun 2023 Company Overview: JPMorgan Chase & Co. is one of the largest global financial institutions providing investment banking, digital banking platforms, and financial analytics services to millions of customers worldwide. Project Goal: Developed scalable cloud-based data infrastructure supporting enterprise analytics and reporting by modernizing legacy data pipelines and migrating large-scale financial datasets to cloud-native data platforms. Key Contributions Developed scalable cloud ETL/ELT pipelines using Python, PySpark, SQL, Apache Spark, Kafka, and Airflow to ingest and transform large-scale financial datasets across cloud data platforms. Implemented real-time streaming pipelines with Kafka and Pub/Sub, applying Spark Structured Streaming concepts for high-throughput, low-latency processing and resilient event ingestion. Designed lake and warehouse architectures with dimensional data models supporting enterprise reporting, analytics, data science, and downstream machine learning workloads. Built automated data validation, profiling, reconciliation, and quality controls using Python and SQL; resolved data integrity issues and improved trust in analytical datasets. Implemented CDC, incremental ingestion, reusable ingestion frameworks, and multi-stage orchestration patterns to improve pipeline scalability, reliability, and development velocity. Optimized Spark transformations and cloud warehouse queries through partitioning, query tuning, resource optimization, and efficient distributed processing strategies. Supported AI/ML enablement through feature engineering, ML-ready datasets, MLOps integration, feature store workflows, model training/deployment data pipelines, and collaboration with Data Scientists and ML Engineers. Applied security and governance practices including IAM, RBAC, secure cloud access, encryption, monitoring, auditability, and enterprise data governance standards. Automated deployments with Git, Jenkins, Docker, Kubernetes, and CI/CD practices while collaborating with platform, security, architecture, and product teams. Cognizant Data Engineer India | Jul 2017 Oct 2021 Company Overview: Cognizant is a global IT consulting and digital transformation services company providing enterprise cloud, analytics, and data engineering solutions for global organizations. Project Goal: Developed scalable enterprise data platforms supporting operational analytics, reporting systems, and data science initiatives through distributed data processing and cloud-based data architectures. Key Contributions Developed enterprise batch and streaming data pipelines using Apache Spark, PySpark, SQL, Hadoop, Kafka, and Airflow for large-scale operational analytics and reporting. Designed cloud data lake and warehouse architectures for structured and semi-structured datasets, supporting scalable ETL/ELT and downstream BI and data science workloads. Implemented distributed Spark transformations, reusable ingestion frameworks, and performance tuning strategies to improve processing throughput, scalability, and reliability. Integrated data from APIs, relational databases, flat files, and event streams; applied cleansing, enrichment, validation, and reconciliation controls to maintain data quality. Built workflow orchestration with Apache Airflow and automated monitoring for multi-stage ingestion and transformation processes. Supported machine learning and advanced analytics teams with curated datasets, feature engineering pipelines, governed data access, and production-ready data services. Applied security, governance, metadata, lineage, and operational best practices across enterprise data platforms and collaborated with cross-functional engineering teams. EDUCATION ________________________________________ Master of Science Computer Science 2023 Bachelor of Technology Computer Science- 2017 Keywords: continuous integration continuous deployment artificial intelligence machine learning business intelligence sthree information technology Arkansas Colorado Idaho |