Home

SAI VINUTHNA TATA - Data Scientist
[email protected]
Location: Dallas, Texas, USA
Relocation: YES
Visa: GC
SAI VINUTHNA TATA
Sr. Data Scientist(Gen AI/Agentic AI)
Email: [email protected] | Contact: +1(469) 759-9088 | LinkedIn




PROFESSIONAL SUMMARY:
Senior Data Scientist (Gen AI/AgenticAI) with 12+ years of experience designing and delivering scalable Data Science, Machine Learning, Artificial Intelligence, Predictive Analytics, Statistical Modeling, and Cloud solutions across Healthcare, Financial Services, Government, and Retail domains using Microsoft Azure, AWS, and Google Cloud Platform (GCP).
Extensive expertise in developing production-grade Machine Learning and Generative AI solutions using OpenAI GPT-4o, Azure OpenAI, Amazon Bedrock, LangChain, LangGraph, LlamaIndex, Model Context Protocol (MCP), AI Agents, and Multi-Agent Systems for intelligent automation, knowledge management, document intelligence, and conversational AI.
Experienced building customer-facing AI applications and analytics solutions that improve user experience, digital interactions, and workflow efficiency for healthcare providers, physicians, clinical teams, and business users through AI-driven insights, intelligent automation, and data-driven decision support capabilities across enterprise platforms.
Strong experience architecting Retrieval-Augmented Generation (RAG) solutions using Hybrid Search, Semantic Search, OpenAI Embeddings, Sentence Transformers, Pinecone, FAISS, Azure AI Search, and document intelligence frameworks to deliver accurate, context-aware, explainable AI insights and production-ready enterprise AI applications.
Proven expertise in building end-to-end Data Science, Machine Learning, and AI solutions, including data ingestion, feature engineering, statistical modeling, model development, evaluation, deployment, monitoring, optimization, and performance tuning using Python, Apache Spark, PySpark, Databricks, Scikit-learn, TensorFlow, PyTorch, and MLflow.
Hands-on experience developing secure cloud-native ML services, AI microservices, and scalable APIs using Python, FastAPI, Docker, and Kubernetes, enabling production AI applications, predictive analytics, and intelligent business workflows.
Experienced in implementing MLOps and LLMOps practices using Azure Machine Learning, MLflow, Docker, Kubernetes, Azure DevOps, GitHub Actions, and Jenkins, with expertise in CI/CD, model deployment, model versioning, model monitoring, model governance, prompt evaluation, observability, and continuous model performance monitoring.
Strong background in Machine Learning, Statistical Analysis, Deep Learning, Natural Language Processing (NLP), Predictive Analytics, Recommendation Systems, Time-Series Forecasting, Fraud Detection, Anomaly Detection, Healthcare Analytics, and Explainable AI (SHAP), delivering intelligent, data-driven business solutions.
Experienced designing scalable Data Engineering platforms using Apache Spark, PySpark, Apache Kafka, Azure Data Factory, Databricks, Delta Lake, Snowflake, SQL Server, and ETL/ELT pipelines supporting large-scale Machine Learning, Data Science, AI, analytics, feature engineering, and distributed data processing workloads.
Experienced in developing personalization solutions, customer segmentation models, recommendation systems, behavioral analytics, and data-driven insights to improve customer experience and business decision-making.
Proficient in building executive dashboards and analytics solutions using Power BI, Tableau, SQL, and Python, transforming complex business data into actionable business insights through interactive visualizations and KPI reporting.
Collaborative engineering professional with extensive experience working in Agile Scrum environments, partnering with Product Owners, Solution Architects, Data Engineers, Software Developers, Clinical Experts, Business Stakeholders, and executive leadership to translate complex analytical findings into actionable recommendations for technical and non-technical stakeholders.

TECH STACK:

Programming, Data Analysis & APIs Python, Pandas, NumPy, SciPy, Java, SQL, JavaScript, TypeScript, HTML5, CSS3, Shell Scripting, FastAPI, Flask, Django, Spring Boot, React.js, jQuery, Bootstrap, AJAX, REST APIs, OpenAPI
Statistical Analysis & Data Science EDA, Statistical Modeling, Hypothesis Testing, A/B Testing, Business Analytics, Customer Analytics, Marketing Analytics, Personalization Modeling, Predictive Analytics, Feature Engineering, Regression Analysis, Classification, Clustering, Recommendation Systems, Forecasting, Anomaly Detection, Model Evaluation
Generative AI & Agentic AI OpenAI GPT-4o, Azure OpenAI, Anthropic Claude, Google Gemini, Amazon Bedrock, LangChain, LangGraph, LlamaIndex, Large Language Models (LLMs), Model Context Protocol (MCP), MCP Tools Integration, Agentic AI, Multi-Agent Systems, AI Agents, AI Copilots, Prompt Engineering, Function Calling, Structured Outputs, Agent Memory, Tool Calling, Planning & Reasoning
RAG, NLP & Document Intelligence Retrieval-Augmented Generation (RAG), Hybrid Search, Semantic Search, OpenAI Embeddings, Sentence Transformers, Pinecone, FAISS, ChromaDB, Hugging Face Transformers, BERT, spaCy, NLTK, TF-IDF, OCR, Azure AI Search, Azure AI Document Intelligence, HL7/FHIR
Machine Learning & Deep Learning Scikit-learn, TensorFlow, PyTorch, XGBoost, Random Forest, Gradient Boosting, Decision Trees, Logistic Regression, Linear Regression, Support Vector Machines (SVM), K-Means Clustering, Principal Component Analysis (PCA), Spark MLlib
MLOps, LLMOps & DevOps MLflow, Azure Machine Learning, Amazon SageMaker, Docker, Kubernetes, Azure Kubernetes Service (AKS), Azure DevOps, GitHub Actions, Jenkins, CI/CD, Model Versioning, Model Monitoring, Model Governance, Prompt Evaluation, LLM Evaluation, RAGAS, Langfuse, Git, GitHub, Linux
Data Engineering & Big Data Apache Spark, PySpark, Apache Airflow/DAGs, Spark MLlib, Databricks, Azure Databricks, Apache Airflow, Azure Data Factory, Apache Kafka, Spark Streaming, Delta Lake, BigQuery, ETL/ELT Pipelines, Data Processing, Feature Engineering Pipelines, Data Warehousing
Cloud Platforms & Services Microsoft Azure, Azure OpenAI, Azure AI Search, Azure Machine Learning, Azure Blob Storage, Azure Data Lake Storage, Azure Functions, Azure API Management, Azure Key Vault, Azure Monitor, Application Insights, Amazon SageMaker, AWS Glue, AWS Lambda, Amazon S3, Amazon Bedrock, Amazon Redshift, Google Cloud Platform (GCP), Google AI Platform, BigQuery, Cloud Storage, Cloud Dataproc
Databases & Analytics SQL, SQL Server, PostgreSQL, MySQL, Oracle, Snowflake, MongoDB, Azure SQL Database, Azure Cosmos DB, Power BI, Tableau, Matplotlib, Microsoft Excel
Daya Quality, Governance, Security, Monitoring & Methodologies Data Validation Frameworks, Data Quality Checks, Data Profiling, Data Cleansing, Data Drift Detection, Schema Validation, Data Governance, Privacy Controls, PII Protection OAuth2, JWT, RBAC, HIPAA, PII Masking, Responsible AI, Azure AI Content Safety, SHAP, OpenTelemetry, Prometheus, Grafana, Postman, Agile Scrum, SDLC, Microservices Architecture, Design Patterns


WORK EXPERIENCE:

Client: Quest Diagnostics, Dallas, Texas
Senior Data Scientist (Generative AI & Agentic AI) Aug 2023 Present
Architected scalable Generative AI and Data Science platforms using LLMs, Agentic AI, RAG, and cloud services, enabling healthcare analytics, intelligent automation, and data-driven decision support across enterprise clinical workflows.
Designed scalable autonomous multi-agent AI systems using LangGraph, LangChain, and LlamaIndex on Azure Cloud Platform for autonomous reasoning, planning, memory management, dynamic tool orchestration, and intelligent task execution.
Built AI Copilots using GPT-4o, Azure OpenAI, and Azure AI Studio for physicians, providers, and laboratory users, enabling conversational healthcare assistance, intelligent search, contextual question answering, and improved digital user experience.
Developed enterprise RAG frameworks integrating EHRs, laboratory reports, clinical documentation, and medical guidelines using Azure AI Search, hybrid retrieval, reranking, and semantic search for accurate clinical knowledge retrieval.
Engineered scalable ETL/ELT pipelines using Python, SQL, Pyspark, AI Document Intelligence, OCR, semantic chunking, Azure Blob Storage, and metadata extraction to extract, transform, validate, and process structured and unstructured datasets.
Designed scalable embedding and vector search pipelines using OpenAI Embeddings, Azure AI Search, Pinecone, and FAISS for high-performance semantic retrieval, clinical knowledge discovery, and low-latency vector search optimization.
Implemented advanced Prompt Engineering strategies including few-shot learning, dynamic prompting, structured outputs, function calling, and response grounding to improve LLM accuracy, consistency, explainability, reliability and robustness.
Developed secure AI orchestration services using Python, FastAPI, REST APIs, Azure Functions, and microservices, enabling seamless enterprise AI integration with healthcare applications, distributed AI services, and intelligent automation.
Developed distributed data engineering workflows using PySpark, SQL, Databricks, Delta Lake, and Data Factory for feature engineering, large-scale processing, analytics, and production machine learning pipeline development.
Developed and optimized supervised and unsupervised machine learning models using Scikit-learn, XGBoost, TensorFlow, and PyTorch for prediction, anomaly detection, forecasting, feature engineering, and healthcare business analytics.
Designed enterprise LLMOps and MLOps pipelines using MLflow, Azure Machine Learning, Docker, AKS, Azure DevOps, GitHub Actions, and Jenkins for model deployment, performance monitoring, and lifecycle management.
Performed exploratory data analysis (EDA), statistical analysis, feature validation, and data quality assessments using Python, SQL, and PySpark to identify patterns, anomalies, & actionable insights from enterprise datasets with improved accuracy.
Implemented comprehensive model evaluation and monitoring frameworks using RAGAS, Langfuse, ML evaluation metrics, and feedback analytics to improve model performance, reliability, accuracy, and production AI quality.
Integrated Azure OpenAI, Azure AI Search, Azure ML, Key Vault, Azure API Management, AKS, and Azure Functions with secure enterprise authentication, RBAC controls, and governed cloud security frameworks for scalable AI deployments.
Implemented Model Context Protocol (MCP) architectures connecting LLMs with enterprise APIs, healthcare systems, external knowledge repositories, and internal business applications across Azure Cloud Platform services using secure integrations.
Designed AI and ML monitoring frameworks using Azure Monitor, OpenTelemetry, Prometheus, and Grafana to track model performance, data quality, drift indicators, latency, reliability, & operational metrics for continuous production optimization.
Collaborated with business stakeholders, data engineers, analysts, architects, and security teams to translate analytical findings into actionable insights while delivering governed AI and machine learning solutions across enterprise environments.
Implemented Responsible AI practices including guardrails, PII masking, prompt injection protection, audit logging, explainability, governance frameworks, and Azure AI Content Safety for secure enterprise healthcare AI deployments.

Environment: Python, SQL, Scikit-learn, TensorFlow, PyTorch, XGBoost, OpenAI GPT-4o, Azure OpenAI, LangChain, LangGraph, LlamaIndex, Model Context Protocol (MCP), Agentic AI, Multi-Agent Systems, Retrieval-Augmented Generation (RAG), Prompt Engineering, OpenAI Embeddings, Azure AI Search, Pinecone, FAISS, Azure AI Document Intelligence, MLflow, Azure Machine Learning, Azure Databricks, Apache Spark, PySpark, Azure Data Factory, Delta Lake, Azure Blob Storage, FastAPI, REST APIs, Docker, Kubernetes, Azure Kubernetes Service (AKS), Azure DevOps, GitHub Actions, Jenkins, Azure Monitor, Application Insights, OpenTelemetry, Prometheus, Grafana, Power BI, HL7/FHIR, HIPAA, Agile Scrum.

Client: BHG Financial, Syracuse, NY
Applied AI / Machine Learning Engineer Feb 2021 July 2023
Architected scalable Applied AI and Machine Learning solutions using AWS Cloud, Python, and Scikit-learn for credit risk assessment, underwriting analytics, fraud detection, lending decision automation and predictive financial analytics.
Engineered end-to-end Machine Learning pipelines using Apache Spark, PySpark, Databricks, Amazon SageMaker, and MLflow for data ingestion, feature engineering, model training, evaluation, deployment, monitoring, and lifecycle management.
Developed predictive Machine Learning models using XGBoost, Random Forest, Gradient Boosting, and Logistic Regression for classification, regression, default prediction, customer risk profiling, creditworthiness assessment, and lending risk analysis.
Designed personalization and recommendation engines using collaborative filtering, clustering, behavioral analytics, and customer segmentation to deliver targeted recommendations and improve customer engagement.
Implemented enterprise Natural Language Processing (NLP) solutions using spaCy, NLTK, Hugging Face Transformers, and AWS Comprehend to analyze loan documents, underwriting reports, and customer communications.
Built intelligent document processing and analytical pipelines using AWS Textract, OCR, metadata extraction, and validation rules to automate loan applications, identity verification, bank statements, tax documents, and regulatory compliance workflows.
Implemented Explainable AI (XAI) frameworks using SHAP, feature importance analysis, statistical validation, and model interpretability techniques to improve regulatory compliance, model transparency, and stakeholder confidence.
Developed scalable Machine Learning inference services using Python, FastAPI, Flask, Docker, and Amazon API Gateway for secure real-time predictions, model integration, and enterprise lending applications across distributed environments.
Engineered distributed data pipelines using Apache Spark, AWS Glue, Amazon S3, Snowflake, and SQL Server to process transactional data, customer profiles, payment history, lending datasets, and analytical workloads.
Designed Machine Learning lifecycle workflows using Amazon SageMaker, MLflow, Docker, Kubernetes, Jenkins, and GitHub Actions to automate model deployment, versioning, monitoring, CI/CD pipelines, and production lifecycle management.
Evaluated early Generative AI capabilities using OpenAI APIs and Amazon Bedrock for document summarization, financial knowledge retrieval, contextual search, AI-assisted analytical insights, and intelligent decision support.
Developed interactive Power BI and Tableau dashboards using SQL to visualize lending performance, portfolio health, credit risk metrics, operational KPIs, predictive analytics, executive decision-making insights, and business performance trends.
Optimized production Machine Learning models through feature engineering, hyperparameter tuning, cross-validation, and statistical evaluation, improving prediction accuracy, operational efficiency, and model reliability.
Developed secure cloud-native ML applications using Amazon SageMaker, AWS Lambda, Amazon ECS, AWS IAM, and AWS Secrets Manager to support scalable model deployment, secure API integration, and enterprise production workloads.
Collaborated with data engineers, software developers, risk analysts, compliance teams, product owners, and business stakeholders using Agile Scrum to deliver scalable, production-ready Applied AI and Machine Learning solutions.
Environment: Python, SQL, Scikit-learn, XGBoost, Gradient Boosting, Logistic Regression, Statistical Modeling, Apache Spark, PySpark, Databricks, MLflow, Amazon SageMaker, AWS Glue, Amazon S3, Snowflake, Power BI, Tableau, spaCy, NLTK, Hugging Face Transformers, AWS Textract, AWS Comprehend, OpenAI API, Amazon Bedrock, FastAPI, Flask, Docker, Kubernetes, Amazon ECS, AWS Lambda, Amazon API Gateway, AWS IAM, AWS Secrets Manager, GitHub Actions, Jenkins, REST APIs, SHAP, Agile Scrum.

Client: State of New York, Albany, NY
Machine Learning Engineer Jun 2019 Jan 2021
Architected scalable Machine Learning and Data Science solutions on Google Cloud Platform (GCP) using Python, statistical modeling, and predictive analytics for revenue forecasting, fraud detection, budget optimization, and enterprise financial planning.
Engineered end-to-end ML pipelines using Scikit-learn, TensorFlow, MLflow, Google AI Platform, and Cloud Storage covering data preparation, feature engineering, model training, evaluation, deployment, monitoring, and lifecycle management.
Developed supervised and unsupervised Machine Learning models using Random Forest, XGBoost, Decision Trees, and clustering techniques for classification, regression, anomaly detection, forecasting, and financial risk analysis.
Built & implemented scalable ETL/ELT workflows pipelines for Machine Learning models using Apache Spark, PySpark, Cloud Dataproc, & BigQuery to extract, transform, validate, & process large-scale financial datasets for ML model development.
Performed exploratory data analysis (EDA), statistical analysis, feature engineering, and data validation using Python, SQL, and PySpark to identify patterns, anomalies, data quality issues, and actionable business insights.
Developed real-time analytical pipelines using Apache Kafka, Spark Streaming, and Cloud Pub/Sub to process financial transactions, operational events, and streaming datasets supporting forecasting and real-time fraud detection.
Developed NLP-based feature extraction pipelines using spaCy, NLTK, and BERT for document classification, named entity recognition, semantic analysis, and structured transformation of government reports to support predictive analytics.
Operationalized production Machine Learning models using Docker, Google Kubernetes Engine (GKE), MLflow, and Cloud Build enabling scalable deployment, automated workflows, model versioning, and continuous performance monitoring.
Engineered ML dataset preparation workflows using BigQuery, Cloud SQL, and Snowflake for large-scale data integration, transformation, dataset preparation, and preprocessing, improving query performance, training efficiency, and model readiness.
Developed interactive Power BI and Tableau dashboards visualizing financial KPIs, revenue trends, expenditure forecasts, fraud indicators, ML model performance, predictive analytics, and executive reporting insights for business stakeholders.
Optimized production ML models through feature selection, hyperparameter tuning, cross-validation, statistical evaluation, and model monitoring improving prediction accuracy, reliability, and operational effectiveness.
Implemented data quality and governance practices including validation rules, anomaly detection, access controls, audit logging, and secure data management to maintain reliable enterprise datasets with improved data integrity.
Collaborated with data scientists, data engineers, analysts, project managers, and business stakeholders using Agile Scrum methodologies to deliver scalable machine learning solutions supporting statewide financial operations.
Environment: Python, SQL, Scikit-learn, TensorFlow, Statistical Modeling, Random Forest, XGBoost, Decision Trees, MLflow, Apache Spark, PySpark, Apache Kafka, Spark Streaming, spaCy, NLTK, BERT, Google Cloud Platform (GCP), Google AI Platform, BigQuery, Cloud Storage, Cloud Dataproc, Cloud Pub/Sub, Cloud SQL, Docker, Google Kubernetes Engine (GKE), Cloud Build, Snowflake, Power BI, Tableau, Git, Agile Scrum.
Client: TJ Maxx, Framingham, MA
Data Scientist Oct 2016 May 2019
Developed predictive ML models using Python, Scikit-learn, Random Forest, Gradient Boosting, and Logistic Regression for demand forecasting, inventory optimization, pricing analytics, customer behavior modeling, and retail performance improvement
Performed exploratory data analysis (EDA), feature engineering, statistical modeling, hypothesis testing, and data visualization using Pandas, NumPy, and SciPy to uncover customer purchasing patterns, sales trends, and actionable business insights.
Designed customer segmentation models using K-Means Clustering, Hierarchical Clustering, PCA, and behavioral analytics to improve targeted marketing campaigns, personalized promotions, customer retention, and merchandising effectiveness.
Built recommendation systems using collaborative filtering, association rule mining, and market basket analysis to generate product recommendations, cross-selling opportunities, customer affinity insights, and promotional optimization strategies.
Developed time-series forecasting models using ARIMA, Holt-Winters, and statistical forecasting techniques to predict product demand, seasonal purchasing behavior, inventory replenishment planning, and store-level sales forecasting.
Evaluated predictive models using feature selection, hyperparameter tuning, cross-validation, statistical validation, and performance metrics to improve model accuracy, robustness, reliability, and business decision-making.
Performed A/B testing, statistical experimentation, and regression analysis to evaluate pricing strategies, promotional campaigns, customer engagement, and merchandising effectiveness using quantitative business metrics.
Developed interactive Power BI and Tableau dashboards visualizing merchandising KPIs, sales performance, customer purchasing behavior, inventory trends, demand forecasts, and executive decision-making supporting strategic retail initiatives.
Collaborated with merchandising teams, pricing analysts, marketing managers, supply chain stakeholders, and executive leadership using Agile Scrum to deliver data-driven insights supporting enterprise retail strategy and operational excellence.
Environment: Python, SQL, Pandas, NumPy, SciPy, Scikit-learn, Statistical Modeling, Random Forest, Gradient Boosting, Decision Trees, Logistic Regression, K-Means Clustering, Hierarchical Clustering, Principal Component Analysis (PCA), Association Rule Mining, Market Basket Analysis, ARIMA, Holt-Winters, Matplotlib, Power BI, Tableau, PostgreSQL, SQL Server, Git, Agile Scrum.

Client: Biocon, Hyderabad, India.
Python Developer Feb 2015- Aug 2016
Developed Python-based enterprise applications using Django, Flask, HTML5, CSS3, JavaScript, and Bootstrap to automate pharmaceutical workflows, operational reporting, business processes, improving enterprise operational efficiency.
Built scalable backend modules using Python, Django, SQL, and REST APIs, implementing reusable business logic, authentication, authorization, session management, role-based access control, and application integration across healthcare systems.
Developed SQL queries and optimized MySQL, PostgreSQL, and SQL Server databases through stored procedures, indexing, query tuning, and schema optimization to improve transaction processing, reporting performance, and data integrity.
Developed automated data processing, data validation, and reporting utilities using Python, Pandas, SQL, Matplotlib, and Microsoft Excel, improving reporting accuracy, efficiency, business analytics, and decision-making across departments.
Designed interactive reporting dashboards and business interfaces using HTML5, CSS3, JavaScript, jQuery, and Bootstrap, enabling business users to visualize operational metrics, analyze key performance indicators, and support operational reporting.
Performed data validation, exception handling, logging, application security, and performance optimization to ensure reliable data processing, accurate reporting, software stability, enterprise application reliability, and production system performance.
Collaborated with business analysts, QA engineers, project managers, and cross-functional teams using Git, Jenkins, and Agile Scrum to deliver scalable, maintainable, production-ready business applications supporting operational and reporting requirements.

Environment: Python 2.7, SQL, Pandas, Matplotlib, Microsoft Excel, Django, Flask, REST APIs, MySQL, PostgreSQL, SQL Server, HTML5, CSS3, JavaScript (ES5), jQuery, Bootstrap, AJAX, Git, Jenkins, Linux, Apache HTTP Server, Agile Scrum.

Education: Bachelors in computer science, Guru Nanak Institute of Technology, India
Keywords: continuous integration continuous deployment quality analyst artificial intelligence machine learning javascript business intelligence sthree database Massachusetts New York

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7700
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: