Home

Kalyan - Data Scientist with 11+ years Experience
[email protected]
Location: Charlotte, North Carolina, USA
Relocation: yes
Visa:
Resume file: Swapnil Chandrakant Resume_1785187778731.docx
Please check the file(s) for viruses. Files are checked manually and then made available for download.
Swapnil Chandrakant Banduke
Lead Data Scientist | AI Engineer | Machine Learning Engineer
+1 (972)-332-1877 | [email protected] | LinkedIn: www.linkedin.com/in/swapnil-chandrakant
PROFESSIONAL SUMMARY
Senior Data Scientist with over 11 years of experience designing, developing, and deploying enterprise-scale machine learning, predictive analytics, and AI solutions across Banking, Healthcare, Retail, and Digital Marketing domains, delivering data-driven products that support strategic business decisions and operational excellence.
Demonstrated expertise in building end-to-end machine learning solutions by translating complex business problems into scalable analytical models using Python, SQL, PySpark, and cloud-native technologies while collaborating closely with cross-functional stakeholders throughout the project lifecycle.
Extensive experience developing supervised, unsupervised, and ensemble learning models to solve business challenges involving fraud detection, customer segmentation, demand forecasting, recommendation systems, credit risk assessment, pricing optimization, and healthcare analytics.
Strong foundation in statistical modeling, probability theory, hypothesis testing, feature engineering, feature selection, dimensionality reduction, model evaluation, and experimental design to improve prediction accuracy and business outcomes across large-scale enterprise applications.
Hands-on experience working with machine learning algorithms including Logistic Regression, Random Forest, Decision Trees, Gradient Boosting, XGBoost, LightGBM, Support Vector Machines, K-Means Clustering, Principal Component Analysis, and time-series forecasting models.
Designed scalable feature engineering pipelines that process structured, semi-structured, and unstructured datasets from multiple enterprise data sources while ensuring data quality, consistency, governance, and reproducibility throughout the analytical lifecycle.
Built production-ready Natural Language Processing solutions using transformer-based models, text embeddings, semantic search, Named Entity Recognition, sentiment analysis, document classification, and information extraction techniques for enterprise business applications.
Developed enterprise AI applications integrating Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), vector databases, and prompt engineering techniques to improve knowledge discovery, intelligent search, and document understanding for business users.
Leveraged Python libraries including Pandas, NumPy, Scikit-learn, TensorFlow, PyTorch, Matplotlib, Seaborn, and SciPy to develop scalable analytical solutions, automate workflows, and build reusable machine learning components.
Built scalable ETL and ELT workflows using Apache Airflow, Databricks, Snowflake, and SQL-based transformation frameworks to support enterprise reporting, advanced analytics, and machine learning initiatives.
Implemented cloud-native machine learning solutions utilizing AWS services including S3, EC2, Lambda, SageMaker, CloudWatch, IAM, and RDS to automate model training, deployment, monitoring, and operational scalability.
Worked extensively with Azure cloud services for enterprise analytics by leveraging Azure Machine Learning, Azure Data Factory, Azure Storage, Azure SQL Database, and Azure DevOps to support production-ready AI and data engineering solutions.
Implemented MLOps best practices by automating model training, validation, versioning, deployment, and monitoring using MLflow, GitHub Actions, CI/CD pipelines, and model performance tracking frameworks.
Built interactive dashboards and executive reporting solutions using Tableau and Power BI, transforming complex analytical insights into intuitive visualizations that supported executive decision-making and operational performance monitoring.
Collaborated with data engineers, software developers, business analysts, product owners, and domain experts to deliver scalable analytical products that aligned with business objectives, regulatory standards, and operational requirements.
Experienced in working with relational databases, cloud data warehouses, and distributed storage platforms by writing optimized SQL queries, stored procedures, complex joins, and performance-tuned analytical workloads supporting enterprise reporting and AI applications.
Built enterprise forecasting solutions using statistical and machine learning approaches to improve inventory planning, healthcare resource allocation, financial risk assessment, customer demand prediction, and operational planning.
Strong understanding of Responsible AI principles, model governance, explainable machine learning, bias detection, and ethical AI practices to ensure enterprise-grade, transparent, and reliable analytical solutions.
Proven ability to communicate technical concepts effectively to both technical and non-technical stakeholders by translating analytical findings into actionable business recommendations that drive measurable organizational value.
Passionate about building intelligent, scalable, and production-ready AI solutions by combining strong analytical thinking, modern machine learning technologies, cloud computing, and software engineering best practices to solve complex enterprise business challenges.
TECHNICAL SKILLS
Programming Languages: Python, SQL, PySpark, R, Shell Scripting
Machine Learning: Supervised Learning, Unsupervised Learning, Ensemble Learning, Classification, Regression, Clustering, Recommendation Systems, Time-Series Forecasting, Anomaly Detection, Feature Engineering, Feature Selection, Hyperparameter Tuning, Model Evaluation, Cross Validation
Artificial Intelligence: Machine Learning, Deep Learning, Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Prompt Engineering, AI Model Deployment, Explainable AI (XAI), Responsible AI
Natural Language Processing: NLTK, spaCy, Hugging Face Transformers, BERT, Sentence Transformers, Text Classification, Named Entity Recognition (NER), Sentiment Analysis, Topic Modeling, Semantic Search, Text Embeddings
Deep Learning Frameworks: TensorFlow, Keras, PyTorch
Machine Learning Libraries: Scikit-learn, XGBoost, LightGBM, CatBoost, Pandas, NumPy, SciPy, Matplotlib, Seaborn
Big Data & Data Engineering: Apache Spark, PySpark, Databricks, Apache Airflow, Apache Kafka, Snowflake, ETL, ELT, Data Warehousing, Data Modeling, Data Quality, Data Pipelines
Generative AI & Vector: OpenAI, Azure OpenAI, LangChain, FAISS, Pinecone, Vector Embeddings, Semantic Retrieval, Knowledge Retrieval
Cloud Platforms: Amazon Web Services (AWS), Microsoft Azure
AWS Services: Amazon S3, EC2, SageMaker, Lambda, IAM, RDS, CloudWatch, VPC
Azure: Azure Machine Learning, Azure Data Factory, Azure SQL Database, Azure Blob Storage, Azure DevOps, Azure Key Vault
MLOps & DevOps: MLflow, Docker, Kubernetes, GitHub Actions, Git, CI/CD Pipelines, Model Monitoring, Model Versioning
Databases: Snowflake, SQL Server, PostgreSQL, MySQL, Oracle
Business Intelligence & Visualization: Tableau, Microsoft Power BI, Advanced Microsoft Excel
Development Methodologies: Agile, Scrum, SDLC, Test-Driven Development (TDD), Continuous Integration, Continuous Deployment

Ally Financial, Charlotte, NC Jan 2023 Present
Lead AI Data Scientist
Project: Enterprise AI-Powered Fraud Detection & Credit Risk Intelligence Platform
Architected and deployed enterprise-grade machine learning solutions for fraud detection, credit risk assessment, customer behavior analytics, and financial forecasting by leveraging Python, SQL, PySpark, XGBoost, and Scikit-learn, enabling business teams to make faster and more informed lending decisions.
Developed Retrieval-Augmented Generation (RAG) solutions using Azure OpenAI, LangChain, vector embeddings, and enterprise knowledge repositories to assist business users in retrieving financial policies, operational documentation, and regulatory guidelines through conversational AI.
Engineered scalable feature engineering pipelines by integrating customer transactions, credit bureau information, digital banking activity, and third-party financial datasets using PySpark, Snowflake, and Apache Airflow, significantly improving model performance and data reliability.
Designed advanced fraud detection models utilizing Gradient Boosting, Random Forest, and anomaly detection techniques to identify suspicious financial activities while minimizing false positives across high-volume transaction processing systems.
Developed predictive credit risk models by combining historical customer profiles, repayment behavior, banking transactions, and alternative financial indicators to improve underwriting decisions and automate risk scoring across lending portfolios.
Built customer lifetime value, propensity, and churn prediction models using supervised machine learning algorithms, enabling personalized financial product recommendations and improving customer retention through data-driven engagement strategies.
Automated end-to-end machine learning workflows by implementing reusable feature engineering frameworks, model training pipelines, hyperparameter optimization, and validation processes using MLflow, GitHub Actions, and CI/CD practices to accelerate production deployments.
Leveraged transformer-based language models to automate document classification, customer inquiry categorization, and intelligent content summarization, reducing manual review efforts while improving response accuracy across customer support operations.
Collaborated with data engineers to build distributed data processing pipelines using Apache Spark, PySpark, Snowflake, and cloud-native storage services capable of processing millions of customer records efficiently for analytical and machine learning workloads.
Designed explainable machine learning solutions using SHAP, LIME, and feature importance analysis to improve model transparency and support regulatory compliance for financial risk management and lending decisions.
Optimized SQL queries, stored procedures, and analytical workloads across enterprise data warehouses to improve data retrieval performance, enabling faster feature generation and reducing model training cycle times.
Implemented model monitoring frameworks to continuously evaluate prediction accuracy, model drift, feature stability, and production performance while establishing automated retraining strategies for evolving financial datasets.
Built scalable REST-based model inference services using Docker containers and Kubernetes orchestration, enabling secure deployment of predictive models across cloud-hosted enterprise banking applications.
Developed customer segmentation frameworks by combining behavioural analytics, transaction history, demographic attributes, and digital engagement metrics, enabling targeted marketing campaigns and personalized financial product offerings.
Performed extensive exploratory data analysis and statistical modeling to identify hidden trends, data quality issues, and predictive variables influencing customer acquisition, credit performance, and fraud patterns across multiple banking products.
Applied advanced feature selection, dimensionality reduction, ensemble learning, and hyperparameter optimization techniques to maximize predictive performance while reducing model complexity and improving production scalability.
Partnered with business stakeholders, risk analysts, compliance teams, and product owners to translate complex regulatory and operational requirements into scalable AI-driven solutions aligned with organizational objectives.
Established enterprise MLOps best practices by implementing model versioning, automated deployment pipelines, reproducible experimentation, artifact management, and governance standards to improve collaboration across data science and engineering teams.
Developed executive dashboards and analytical reporting solutions using Power BI, providing senior leadership with real-time visibility into fraud trends, portfolio performance, customer behavior, and model effectiveness.
Mentored junior data scientists by conducting technical design reviews, code reviews, model validation sessions, and knowledge-sharing workshops while promoting software engineering best practices and reusable analytical frameworks.
Actively participated in Agile ceremonies including sprint planning, backlog refinement, estimation sessions, technical demonstrations, and retrospective meetings while collaborating with cross-functional teams to deliver high-quality analytical solutions within project timelines.
Continuously evaluated emerging AI technologies, foundation models, and modern machine learning frameworks to identify opportunities for enhancing enterprise analytics capabilities while ensuring secure, scalable, and responsible adoption of AI across financial applications.
Environment: Python, SQL, PySpark, Pandas, NumPy, SciPy, Scikit-learn, TensorFlow, PyTorch, XGBoost, LightGBM, CatBoost, Hugging Face Transformers, OpenAI, Azure OpenAI, LangChain, Prompt Engineering, Retrieval-Augmented Generation (RAG), FAISS, Pinecone, Apache Spark, Databricks, Snowflake, Apache Airflow, Apache Kafka, MLflow, Feature Store, Model Registry, Docker, Kubernetes, Git, GitHub Actions, Jenkins, CI/CD, Azure Machine Learning, Azure DevOps, Azure Blob Storage, Azure Data Factory, Power BI, Tableau, SHAP, LIME, REST APIs, FastAPI, Flask, Linux, Jira, Confluence, Agile, Scrum.
New Mexico Department of Health | Santa Fe, NM Jan 2021 Dec 2022
Senior Machine Learning Engineer
Project: Population Health Analytics & Disease Surveillance Platform
Designed and implemented predictive analytics solutions using Python, SQL, Spark, and Scikit-learn to analyze statewide public health datasets, enabling healthcare leadership to identify disease trends, monitor population health, and support evidence-based policy decisions.
Built scalable machine learning models for disease risk prediction by integrating electronic health records, laboratory reports, demographic information, and social determinants of health, improving early identification of high-risk patient populations.
Developed statistical forecasting models using regression analysis, ensemble learning, and time-series techniques to predict disease outbreaks, hospital resource utilization, and vaccination demand, supporting proactive healthcare planning across multiple public health programs.
Engineered robust data ingestion and transformation pipelines using PySpark, Apache Airflow, and Azure Data Factory to consolidate structured and semi-structured healthcare data from multiple clinical, laboratory, and surveillance systems into centralized analytical platforms.
Performed comprehensive exploratory data analysis on millions of healthcare records to uncover hidden trends, eliminate data quality issues, identify missing values, and generate meaningful insights that supported epidemiological investigations and operational reporting.
Applied advanced feature engineering, dimensionality reduction, and feature selection techniques to improve predictive model performance while maintaining model interpretability and compliance with healthcare reporting standards.
Developed Natural Language Processing solutions to extract meaningful clinical information from physician notes, laboratory reports, discharge summaries, and public health documentation using spaCy, transformer-based language models, and semantic text processing techniques.
Automated healthcare reporting processes by developing reusable analytical workflows and interactive Power BI dashboards, enabling executive leadership to monitor disease prevalence, vaccination progress, mortality rates, and healthcare service utilization in near real time.
Designed patient risk stratification models using classification algorithms, probability scoring, and behavioral analytics to assist healthcare providers in prioritizing preventive interventions and improving patient care outcomes.
Built scalable ETL and ELT workflows using Azure SQL Database, Snowflake, Apache Spark, and SQL to support enterprise reporting, machine learning model training, and longitudinal healthcare analytics across multiple business units.
Collaborated with epidemiologists, healthcare analysts, clinicians, and public health officials to transform complex analytical requirements into production-ready machine learning solutions that addressed critical public health initiatives and operational challenges.
Evaluated multiple supervised and ensemble learning algorithms including Logistic Regression, Random Forest, Gradient Boosting, XGBoost, and LightGBM to identify optimal predictive models for healthcare classification and forecasting problems.
Developed explainable AI solutions using SHAP and feature importance analysis, enabling clinicians and healthcare administrators to understand model predictions while increasing trust and transparency in AI-assisted decision-making.
Implemented automated model validation, performance monitoring, and retraining strategies using MLflow and Azure Machine Learning to ensure production models remained accurate, stable, and aligned with evolving healthcare datasets.
Integrated cloud-native analytical services including Azure Machine Learning, Azure Blob Storage, Azure Data Factory, and Azure DevOps to streamline machine learning lifecycle management and accelerate deployment of enterprise healthcare solutions.
Partnered with data engineers and software development teams to deploy containerized machine learning services using Docker and Kubernetes, enabling secure and scalable model inference across enterprise healthcare applications.
Participated in Agile development processes by contributing to sprint planning, backlog refinement, technical design discussions, code reviews, and cross-functional collaboration while ensuring timely delivery of high-quality analytical solutions.
Established data governance, security, and privacy best practices by implementing data validation, access controls, audit logging, and HIPAA-compliant analytical workflows to safeguard sensitive healthcare information throughout the machine learning lifecycle.
Continuously researched emerging machine learning techniques, healthcare AI innovations, and statistical methodologies to modernize analytical capabilities, improve predictive accuracy, and support data-driven public health decision-making across statewide healthcare initiatives.
Environment: Python, SQL, PySpark, Pandas, NumPy, Scikit-learn, TensorFlow, Keras, PyTorch, XGBoost, LightGBM, NLP, spaCy, NLTK, Hugging Face Transformers, BERT, Apache Spark, Databricks, Snowflake, Apache Airflow, Azure Machine Learning, Azure Data Factory, Azure SQL Database, Azure Blob Storage, Azure Key Vault, Azure DevOps, MLflow, Docker, Kubernetes, Git, GitHub, Power BI, Tableau, PostgreSQL, SQL Server, REST APIs, Feature Engineering, Statistical Modeling, SHAP, LIME, Healthcare Analytics, Linux, Jira, Confluence, Agile, Scrum.
The Home Depot | Atlanta, GA Jul 2017 Dec 2020
Machine Learning Engineer
Project: Retail Demand Forecasting & Customer Intelligence Platform
Collaborated with merchandising, supply chain, and digital commerce teams to develop machine learning solutions that improved demand forecasting, customer analytics, and inventory planning across enterprise retail operations.
Built predictive demand forecasting models using Python, PySpark, XGBoost, and statistical time-series techniques to analyze historical sales trends, seasonal demand, and promotional activities for accurate inventory planning.
Developed customer segmentation models by analyzing transactional, demographic, and behavioral data to identify purchasing patterns and support targeted marketing campaigns across multiple retail channels.
Designed recommendation models using collaborative filtering and machine learning algorithms to improve personalized product suggestions, cross-selling opportunities, and overall customer shopping experience.
Engineered scalable data pipelines using Apache Spark, PySpark, SQL, and Snowflake to process large retail datasets and deliver high-quality features for machine learning applications.
Built pricing analytics models by evaluating customer purchasing behavior, competitor pricing, and promotional performance to support data-driven pricing strategies and revenue optimization.
Performed exploratory data analysis and feature engineering on sales, inventory, and customer datasets to identify key business drivers and improve predictive model performance.
Developed inventory optimization models by combining demand forecasts, supplier lead times, and warehouse capacity data to reduce stock shortages while improving inventory utilization.
Leveraged ensemble learning techniques including Random Forest, Gradient Boosting, XGBoost, and LightGBM to improve prediction accuracy for customer propensity, sales forecasting, and inventory planning.
Automated ETL workflows using Apache Airflow, SQL, and Databricks to streamline data ingestion, transformation, and preparation processes supporting enterprise reporting and machine learning initiatives.
Built interactive Tableau and Power BI dashboards that enabled business users to monitor sales performance, inventory trends, campaign effectiveness, and operational KPIs through real-time visual analytics.
Implemented A/B testing and statistical analysis to evaluate marketing campaigns, pricing strategies, and customer engagement initiatives, enabling data-driven business decisions.
Worked closely with business stakeholders, data engineers, and product teams to translate retail requirements into scalable analytical solutions that supported merchandising and digital transformation initiatives.
Participated in deploying machine learning models using Docker and cloud-based infrastructure while monitoring model performance and continuously improving prediction accuracy in production environments.
Contributed to Agile development by participating in sprint planning, code reviews, technical discussions, and cross-functional collaboration to deliver high-quality analytics solutions within project timelines.
Environment: Python, SQL, PySpark, Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, Random Forest, Apache Spark, Databricks, Snowflake, Apache Airflow, Tableau, Power BI, SQL Server, Oracle, MySQL, AWS (S3, EC2, RDS), Docker, Git, Jenkins, Feature Engineering, Time-Series Forecasting, Prophet, ARIMA, Customer Segmentation, Recommendation Systems, A/B Testing, Statistical Analysis, ETL, Data Warehousing, Linux, Jira, Agile, Scrum.
Quotient | Mountain View, CA Nov 2015 Jun 2017
Data Scientist
Digital Marketing Analytics & Customer Segmentation Platform
Supported digital advertising and consumer marketing initiatives by analysing large-scale campaign, customer, and transactional datasets to identify user behavior patterns and uncover opportunities for improving campaign effectiveness.
Converted raw business data into meaningful analytical datasets using Python, SQL, and data cleansing techniques, providing reliable inputs for reporting, predictive modeling, and marketing intelligence initiatives.
Investigated customer purchasing trends through exploratory data analysis and statistical profiling, helping marketing teams better understand audience preferences and optimize campaign targeting strategies.
Assisted in developing predictive models that estimated customer response, conversion probability, and campaign performance using Logistic Regression, Decision Trees, Random Forest, and other supervised learning techniques.
Produced reusable analytical datasets by integrating information from multiple relational databases, improving data consistency and reducing the effort required for recurring business analysis.
Evaluated campaign success through statistical analysis, hypothesis testing, and A/B experimentation, enabling business teams to make evidence-based decisions for future marketing investments.
Created automated Python utilities for validating incoming data, identifying anomalies, handling missing values, and standardizing datasets before analytical processing, improving overall data quality across projects.
Prepared business-focused dashboards and executive reports using Tableau, enabling marketing leaders to monitor customer acquisition, campaign performance, click-through rates, and conversion metrics through interactive visualizations.
Worked closely with campaign managers and business analysts to translate analytical findings into actionable recommendations that improved audience targeting, promotional effectiveness, and customer engagement.
Improved SQL query performance by optimizing joins, indexing strategies, and aggregation logic, reducing report execution times and supporting faster analytical decision-making.
Contributed to model validation activities by comparing algorithm performance, evaluating prediction accuracy, and documenting analytical findings to support model selection and business adoption.
Participated in Agile development activities including sprint planning, backlog discussions, knowledge-sharing sessions, and peer reviews while continuously expanding expertise in machine learning, statistical analysis, and data visualization.
Environment: Python, SQL, Pandas, NumPy, SciPy, Scikit-learn, Logistic Regression, Decision Trees, Random Forest, Gradient Boosting, K-Means Clustering, PCA, Matplotlib, Seaborn, Tableau, Power BI, Microsoft Excel, SQL Server, Oracle, MySQL, ETL, Data Modeling, Data Cleansing, Feature Engineering, Statistical Analysis, Hypothesis Testing, A/B Testing, Jupyter Notebook, Git, Linux, Jira, Agile, Scrum.

Future Group | Bengaluru, India Oct 2013 Aug 2015
Business Intelligence Analyst
Project: Retail Sales Analytics & Customer Insights Platform
Collaborated with merchandising, sales, and category management teams to analyze retail sales, inventory, and customer transaction data, enabling business users to make informed merchandising and operational decisions.
Collected, cleansed, transformed, and validated large datasets from multiple operational systems using SQL, Excel, and Python to ensure data consistency and accuracy for business reporting and analytical initiatives.
Developed SQL queries, stored procedures, views, and analytical reports to support daily sales analysis, inventory tracking, and executive management reporting across multiple retail business units.
Designed interactive Tableau dashboards and management reports that provided real-time visibility into sales performance, product movement, inventory levels, and key business KPIs.
Worked closely with ETL developers and database administrators to validate data quality, reconcile discrepancies, and maintain reliable analytical datasets for enterprise reporting solutions.
Supported inventory optimization initiatives by analyzing product movement, stock availability, replenishment cycles, and warehouse performance, helping reduce inventory imbalances across retail locations.
Participated in data warehouse testing, report validation, and user acceptance testing (UAT) to ensure analytical solutions met business requirements before production deployment.
Collaborated with business analysts and project stakeholders to understand reporting requirements, translate business needs into technical specifications, and deliver analytical solutions aligned with organizational objectives.
Maintained detailed documentation for SQL scripts, reporting procedures, business rules, and analytical workflows to support knowledge sharing and ongoing application maintenance.
Contributed to Agile development activities by participating in sprint planning, requirement discussions, testing cycles, and cross-functional team meetings while continuously expanding technical expertise in data analytics and business intelligence.
Environment: Python, SQL, PL/SQL, Pandas, NumPy, Microsoft Excel, Tableau, Power BI, Oracle Database, SQL Server, MySQL, ETL, Data Warehousing, Data Modeling, Data Cleansing, Data Validation, Stored Procedures, Views, Statistical Analysis, Regression Analysis, Forecasting, Customer Segmentation, Sales Analytics, Inventory Analytics, Retail Analytics, KPI Reporting, Git, Jira, Agile, Scrum.



Certifications:

Databricks Certified Data Engineer Professional
Snowflake SnowPro Core Certification
AWS Certified Machine Learning Engineer Associate(MLA-C01)
Azure AI Engineer Associate (AI-102)
Keywords: continuous integration continuous deployment artificial intelligence business intelligence sthree rlang procedural language mtv mountain view bay area California Georgia New Mexico North Carolina

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7587
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: