| nikhila - Senior Generative AI/ML Engineer |
| [email protected] |
| Location: York, New York, USA |
| Relocation: YES |
| Visa: H1B |
| Resume file: Nikhila_GenAI_ML_Engineer_1787924751780.docx Please check the file(s) for viruses. Files are checked manually and then made available for download. |
|
Senior Generative AI/ML Engineer
NIKHILA J Phone: 17323999405 104 Senior Generative AI/ML Engineer Email: [email protected] PROFESSIONAL SUMMARY Senior Gen AI / ML Engineer with 6+ years delivering production-grade Generative AI, agentic, and machine learning systems across banking, healthcare, insurance, and financial services with measurable business impact. Specialized in Agentic AI, Generative AI, LLMs, and RAG architectures, designing multi-agent workflows on LangGraph, CrewAI, LangChain, LlamaIndex, AutoGen, and DSPy with planner, tool-caller, and critic patterns. Expert in LLM fine-tuning techniques including LoRA, QLoRA, PEFT, Prefix-Tuning, DPO, PPO, RLHF, RLAIF, and SFT, with hands-on knowledge distillation, pruning, and INT8/FP8 quantization for cost-efficient inference. Strong inference optimization and LLMOps practitioner using Triton Inference Server, TensorRT-LLM, CUDA graphs, ONNX Runtime, KServe, and NVIDIA GPU Operator to maximize GPU throughput and minimize latency. Proven record of shipping enterprise RAG systems on AWS Bedrock, Azure AI Foundry, and Vertex AI using Pinecone, Weaviate, FAISS, Milvus, and OpenSearch with hybrid BM25 retrieval, reranking, and citation pipelines. Deep expertise in classical ML, deep learning, NLP, and computer vision with PyTorch, TensorFlow, Hugging Face Transformers, scikit-learn, XGBoost, BERT, BioBERT, spaCy, OpenCV, YOLO, and Tesseract. End-to-end MLOps and DevOps practitioner using MLflow, Kubeflow, Docker, Kubernetes (AKS, EKS, GKE), OpenShift, Helm, ArgoCD, Terraform, GitHub Actions, Jenkins, Azure DevOps, and AWS CodePipeline. Strong data engineering foundation for AI workloads across Snowflake, Databricks, BigQuery, PySpark, Apache Kafka, Apache Airflow, Apache Beam, dbt, Delta Lake, and Apache Iceberg for batch and streaming. Designed and operated cloud-native AI microservices on Azure, AWS, and GCP with multi-region failover, blue-green and canary deployments, GPU autoscaling, and 99.99% availability SLAs. Recognized for cross-functional leadership and mentorship, partnering with product, risk, legal, and compliance to enforce responsible AI guardrails, HIPAA, SOC2, SOX, PCI DSS, and SR 11-7 governance. TECHNICAL SKILLS Programming: Python, SQL, R, Java, JavaScript, TypeScript, Scala, C++, Bash, Go Cloud & Data Platforms: AWS: Bedrock, SageMaker, Lambda, S3, Glue, Lake Formation, ECS Fargate, EKS, API Gateway, Step Functions, EventBridge, OpenSearch, Comprehend, Kinesis, EMR, Redshift, Athena, CloudWatch, IAM, KMS, CDK, CodePipeline Azure: AI Foundry, OpenAI, Azure ML, Functions, Synapse, Cognitive Search, AI Document Intelligence, AKS, Data Factory, Monitor, App Insights, Key Vault, Purview, DevOps GCP: Vertex AI, BigQuery, Cloud Storage, Compute Engine, Dataflow (Beam), Dataproc, Cloud Composer, Cloud Monitoring, Cloud Logging, Pub/Sub, GKE Data Engineering & Streaming: Snowflake, Databricks, PySpark, Apache Spark, Apache Kafka, Spark Structured Streaming, Apache Airflow, MWAA, Apache Beam, dbt, Delta Lake, Apache Iceberg, Hadoop, Hive, Informatica PowerCenter, SSIS Machine Learning & Deep Learning: PyTorch, TensorFlow, Keras, scikit-learn, XGBoost, LightGBM, CatBoost, TabNet, Prophet, ARIMA, Random Forest, SVM, Gradient Boosting, A/B Testing, champion-challenger LLMs & Generative AI: GPT-4o, GPT-5, Claude Opus, Claude Sonnet, Claude Haiku, Gemini 1.5/2, Llama 3/3.1, Mistral, Mixtral, PaLM, Hugging Face Transformers, RAG, Prompt Engineering (few-shot, chain-of-thought, ReAct, self-consistency), Function Calling, Agentic AI, MCP Agentic & Multi-Agent Frameworks: LangChain, LangGraph, CrewAI, LlamaIndex, AutoGen, Semantic Kernel, DSPy Inference Optimization & LLMOps: Triton Inference Server, TensorRT-LLM, CUDA graphs, dynamic batching, response caching, GPU memory optimization, KServe, NVIDIA GPU Operator, model ensembles. Container Platforms: Kubernetes, OpenShift, Helm, KServe, NVIDIA GPU Operator. NLP: BERT, BioBERT, RoBERTa, Sentence Transformers, spaCy, NLTK, gensim, entity extraction, sentiment, topic modeling, summarization, query understanding. Computer Vision: OpenCV, YOLO, CNNs, GANs, Tesseract, PaddleOCR, Azure AI Document Intelligence, image quality checks, layout parsing Feature Stores, Search & Vector DBs: Databricks Feature Store, Elasticsearch, Azure Cognitive Search, Weaviate (VPC), FAISS/Pinecone. MLOps & Deployment: MLflow (tracking/registry), Docker, Kubernetes (AKS/EKS/GKE), FastAPI/Flask microservices, CI/CD (Jenkins, Azure DevOps), canary releases & rollback, model/feature drift monitoring Workflow & Orchestration: Apache Airflow (incl. MWAA), Cloud Composer, Databricks Jobs/Workflows, EventBridge schedules Api: FastAPI, Flask, REST APIs, schema validation, OpenAPI/Swagger, authentication/authorization, API integration, observability hooks. Databases & Warehouses: PostgreSQL, MongoDB, MySQL, Redis, DynamoDB, Oracle, SQL Server, Cassandra, HBase, Neo4j, BigQuery, Snowflake Observability & Monitoring: Prometheus, Grafana, CloudWatch, Azure Monitor/App Insights, GCP Cloud Monitoring/Logging Visualization & Apps: React, TypeScript, Tableau, Power BI, Looker Studio (Data Studio), Streamlit internal experiment/telemetry dashboards, Prometheus, Grafana, CloudWatch, Azure Monitor, GCP Monitoring, alerting/SLAs/SLOs, dashboards, ML telemetry, custom model metrics, API monitoring, GPU performance monitoring, KPI tracking. Dimensionality Reduction & Topic Modeling: PCA, UMAP, t-SNE, LSA, LDA, NMF, SHAP, LIME, Captum, fairness and parity monitoring Version Control & IaC: Git, GitLab, Jenkins, Azure DevOps, Terraform, CI/CD best practices PROFESSIONAL EXPERIENCE Bank of America, Remote Sep 2025 present Senior Generative AI & ML Engineer Responsibilities: Designed multi-agent systems with LangGraph, CrewAI, and DSPy orchestrating Claude Opus, Claude Sonnet, and GPT-4o for regulatory document analysis, automating 70% of compliance review workflows across 12 lines of business. Deployed enterprise agentic RAG pipelines on AWS Bedrock combining Pinecone vector search with hybrid BM25 retrieval and Cohere reranking, raising answer relevance from 0.71 to 0.92 across 4M internal policy documents. Built production fine-tuning workflows applying LoRA, QLoRA, and DPO on Llama 3.1 70B for SR 11-7 model risk validation, cutting inference costs 58% while preserving accuracy within 1.2 points of GPT. Engineered evaluation harnesses with Ragas, LangSmith, and custom LLM-as-judge metrics, scoring 14 production agents weekly on faithfulness, groundedness, and hallucination rate across 8,000 synthetic test cases. Containerized agent services with Docker on Amazon EKS behind API Gateway, scaling to 2.3M daily inference calls at p95 latency under 800ms across three AWS regions with active-active failover. Established prompt versioning and rollback strategies using MLflow and Bedrock Prompt Management, shrinking incident recovery time from 45 minutes to under 6 minutes. Partnered with risk, legal, and InfoSec to codify responsible AI guardrails using Amazon Comprehend PII detection and content safety filters, blocking 99.4% of policy violations before model invocation. Implemented Model Context Protocol servers exposing 32 internal banking tools to LLM agents, enabling secure tool calling for account, transaction, and risk lookups under fine-grained IAM controls. Optimized Triton Inference Server pipelines with TensorRT-LLM and FP8 quantization on EKS GPU nodes, raising large-context summarization throughput 2.8x and freeing 18 A100 instance hours daily. Drove cost optimization across Bedrock and SageMaker workloads through prompt compression, response caching, and right-sizing of GPU inference endpoints, saving $1.8M annualized. Constructed CUDA-graph-accelerated inference workflows for batch agent execution, cutting per-call GPU memory footprint by 34% and enabling 4x concurrent agent sessions on the same hardware. Led migration of legacy XGBoost credit-risk scoring models to a unified MLflow registry on EKS, halving model retraining cycle time from 21 days to 9 days. Standardized LLMOps tooling across the AI Center of Excellence with Helm, ArgoCD, KServe, and LangFuse, accelerating model promotion velocity by 38% for 11 partner teams. Configured multi-region active-active failover for agent services on EKS with Route 53 and health-based routing, achieving 99.99% availability and 90-second recovery during DR drills. Hardened defense-in-depth security across AWS using VPC endpoints, KMS encryption, IAM least privilege, and Secrets Manager, satisfying SOC2 and internal model governance audits across 24 services. Mentored 2 mid-level engineers on agentic patterns, prompt engineering, and evaluation methodology, lifting team velocity by 28% measured through sprint completion rates. Authored internal LLMOps standards covering prompt management, evaluation gates, drift monitoring, and red-teaming, adopted by 9 engineering teams across the AI Center of Excellence. Pioneered red-teaming and adversarial evaluation workflows using jailbreak corpora and prompt-injection test suites, hardening 14 production agents against 312 attack patterns. Integrated DSPy-based prompt programs with automatic prompt optimization, lifting end-task accuracy by 11 points across regulatory Q&A and compliance summarization tasks. Defined SLOs and golden-signal dashboards for agent latency, cost-per-call, tool-call success rate, and groundedness, tied to PagerDuty escalations for 14 production services. Tech Stack: Python, FastAPI, PyTorch, LangGraph, CrewAI, DSPy, Hugging Face, Claude, GPT-4o, Llama 3.1, Triton, TensorRT-LLM, ONNX, MLflow, AWS Bedrock, SageMaker, Lambda, S3, Step Functions, EKS, API Gateway, OpenSearch, Pinecone, KMS, Comprehend, Docker, Helm, ArgoCD, KServe, Terraform, GitHub Actions, Ragas, LangSmith, LangFuse, CloudWatch, Datadog, OpenTelemetry. CareFirst BlueCross BlueShield, Remote Jan 2024 Aug 2025 Gen AI/ML Engineer Responsibilities: Designed and shipped a RAG clinical assistant using LangChain, OpenAI GPT-4, and FAISS over 1.8M de-identified member records, helping care managers reduce chart review time by 42%. Operationalized HIPAA-compliant data pipelines on AWS using Glue, Lambda, and S3 to ingest FHIR claims and authorization data at 95GB daily volume, with PII redaction reaching 99.7% precision. Fine-tuned BERT and BioBERT classifiers for prior-authorization triage on 220K labeled cases, lifting precision from 0.78 to 0.89 and trimming the manual review queue by 31%. Delivered scikit-learn and XGBoost risk-stratification models behind FastAPI services on Amazon ECS Fargate, serving 600K predictions per day with 99.95% uptime. Spearheaded MLOps workflows using MLflow, GitHub Actions, and Terraform, cutting model promotion lead time from 14 days to 3 days across 18 model registries. Developed an A/B testing framework for prompt variants and retrieval strategies, running 47 experiments that informed rollout decisions for 9 production GenAI features. Instrumented an observability stack with CloudWatch, Datadog, OpenTelemetry, Prometheus, and Grafana to monitor model drift, latency, GPU saturation, and token usage, surfacing 23 degradation events before customer impact. Composed a multi-step agent on AWS Bedrock and Claude 3 to summarize clinical notes, extract ICD-10 codes, and draft member outreach messages, reducing coder workload by 36%. Connected Pinecone- and Weaviate-backed semantic search across 600K provider records with hybrid lexical retrieval and BGE reranking, raising top-3 match rate from 71% to 94%. Crafted prompt engineering patterns including few-shot exemplars, chain-of-thought, and self-consistency for clinical summarization, lifting ROUGE-L scores by 19 points. Tuned GPT-4 inference cost through prompt compression, batching, and routing low-risk queries to Llama 3 hosted on SageMaker with TensorRT-LLM, lowering per-query spend by 47%. Launched Triton Inference Server with INT8 quantization for clinical summarization workloads, doubling GPU throughput while keeping p95 latency under 2.1 seconds at peak hours. Activated blue-green and canary deployment patterns on EKS with Helm and ArgoCD, enabling safe rollouts of 9 GenAI features with automatic rollback on drift, latency, or bias regression. Wired clinical OCR and entity extraction using Tesseract, PaddleOCR, and Hugging Face Transformers, standardizing PHI redaction and ontology mapping for chart abstraction. Secured workloads end-to-end using VPC endpoints, KMS encryption, IAM least privilege, and AWS Secrets Manager, satisfying HIPAA, HITRUST, and internal governance audits across 24 services. Produced fairness and bias dashboards with Evidently AI and Grafana, monitoring demographic parity within +/- 2.5% across protected classes for prior-auth and risk-stratification models. Coached 8 junior engineers and clinical analysts on LangChain, vector databases, and evaluation methodology, accelerating internal adoption of GenAI tooling across 5 product squads. Tech Stack: Python, FastAPI, PyTorch, TensorFlow, LangChain, Hugging Face, GPT-4, Claude 3, Llama 3, BERT, BioBERT, scikit-learn, XGBoost, Triton, TensorRT-LLM, ONNX, MLflow, AWS Bedrock, SageMaker, Glue, Lambda, S3, ECS Fargate, EKS, OpenSearch, Pinecone, Weaviate, FAISS, Tesseract, PaddleOCR, Docker, Helm, ArgoCD, Terraform, GitHub Actions, CloudWatch, Datadog, OpenTelemetry, Prometheus, Grafana, Evidently AI. JP Morgan Chase Jan 2023 Dec 2023 ML Engineer Responsibilities: Productionized ML models on AWS SageMaker for transaction fraud detection across 4TB of daily payment data, raising precision from 0.82 to 0.91 and saving an estimated $4.2M in chargebacks. Rolled out XGBoost, LightGBM, and PyTorch sequence models behind real-time scoring APIs on Amazon ECS with 99.97% availability and sub-100ms p95 latency. Set up MLOps pipelines using MLflow, SageMaker Pipelines, and GitHub Actions for automated retraining, validation, and canary deployments across 22 model variants. Prototyped early RAG workflows with LangChain, OpenAI GPT-3.5, and Pinecone for internal policy Q&A, accelerating analyst self-service by 35% in pilot rollout. Trained BERT and RoBERTa classifiers on Hugging Face for transaction memo categorization across 11M records, lifting macro F1 from 0.74 to 0.88. Created streaming feature pipelines with Apache Kafka and PySpark Structured Streaming for trade surveillance, sustaining 35K events per second with sub-200ms enrichment latency. Packaged model services with Docker on Amazon EKS using Helm charts, enabling rolling updates, GPU scheduling, and zero-downtime deploys across 4 trading desks. Refined dimensionality reduction pipelines using PCA, UMAP, and t-SNE for high-dimensional fraud feature spaces, surfacing 6 anomaly clusters that informed new rule-based detectors. Tracked model performance with Evidently AI and CloudWatch dashboards, catching 17 drift events that triggered timely retraining and avoided two production incidents. Shipped model retraining triggers using AWS EventBridge and Step Functions tied to drift thresholds, automating 12 retraining cycles quarterly without manual intervention. Reviewed model registry, lineage, and approval workflows in MLflow integrated with ServiceNow change management, satisfying SOX, SR 11-7, and internal model risk governance. Tutored 4 junior ML engineers on Python, MLOps, and feature engineering, accelerating their ramp-up by 40% measured through PR throughput and reviewer feedback. Compressed inference workloads with ONNX export and TorchScript compilation for sequence models, cutting per-request CPU latency by 32% and freeing budget for 3 additional models on shared infrastructure. Tech Stack: Python, PySpark, SQL, FastAPI, PyTorch, Hugging Face, BERT, RoBERTa, XGBoost, LightGBM, scikit-learn, LangChain, OpenAI GPT-3.5, Pinecone, Apache Kafka, Spark Structured Streaming, Snowflake, AWS SageMaker, Lambda, S3, ECS, EKS, EventBridge, Step Functions, MLflow, Docker, Helm, Terraform, GitHub Actions, CloudWatch, Evidently AI, Great Expectations. CGI, IND Jan 2022 Aug 2022 Data Scientist Responsibilities: Modeled customer segmentation and propensity-to-purchase using scikit-learn, XGBoost, and pandas on 5M customer records, lifting campaign conversion rates by 22%. Curated NLP classifiers with Hugging Face Transformers including BERT and RoBERTa for support ticket triage across 180K records, raising routing accuracy from 0.71 to 0.86. Surfaced features and ran exploratory analysis in Python and SQL on Oracle and SQL Server, generating 9 actionable insights that informed pricing and retention strategy. Generated Tableau and Power BI dashboards for finance and operations stakeholders, replacing 22 manual Excel reports and saving roughly 18 analyst hours per week. Applied dimensionality reduction with PCA, t-SNE, and UMAP on 18M-row datasets to visualize customer clusters and inform segmentation logic for marketing campaigns. Drafted LDA and NMF topic-modeling pipelines with gensim and NLTK over 240K support tickets, surfacing 12 latent themes that drove a 19% reduction in repeat call volume. Automated hyperparameter tuning with Optuna and Hyperopt across XGBoost and LightGBM classifiers, lifting macro F1 by an average of 6.4 points across 9 use cases. Documented modeling methodology, data lineage, and validation results in Confluence, accelerating onboarding for three new analysts and supporting transition to the production team. Exposed classification and segmentation models through lightweight Flask APIs to internal teams, instrumented with logging and request metrics for downstream observability. Validated model fairness across customer cohorts using SHAP and LIME, catching 4 protected-class disparities and prompting feature redesigns prior to production rollout. Tech Stack: Python, SQL, pandas, NumPy, scikit-learn, XGBoost, LightGBM, Hugging Face Transformers, BERT, RoBERTa, spaCy, NLTK, gensim, Optuna, Hyperopt, Oracle, SQL Server, Tableau, Power BI, Confluence, Git. LTI Mindtree, IND Jan 2020 Dec 2021 Data Scientist Responsibilities: Programmed Python data science pipelines using pandas, NumPy, and scikit-learn to deliver customer churn and credit risk models on 3.5M records, lifting recall from 0.68 to 0.81. Shaped BERT-based NLP classifiers with TensorFlow and Hugging Face for product review sentiment analysis on 800K reviews, reaching weighted F1 of 0.87. Forecasted loan repayment volumes with Prophet and ARIMA time-series models, cutting forecast error from 14% to 7% and informing capital allocation decisions. Wrote A/B test analysis utilities in Python and SQL on AWS Redshift and authored Jupyter-based reproducible analysis templates, helping product managers evaluate 28 experiments quarterly. Served models in production with TensorFlow Serving and Docker on AWS EC2 behind an Application Load Balancer, sustaining 1,200 inference requests per second with p99 under 240ms. Streamed Kafka events through PySpark Structured Streaming for fraud surveillance, shrinking end-to-end detection latency from 12 seconds to 1.8 seconds. Accelerated deep learning training with PyTorch Distributed Data Parallel across 4-GPU nodes for retail demand forecasting, lifting MAPE accuracy by 14 points over baseline ARIMA. Scaled recommendation engines using collaborative filtering, matrix factorization, and FAISS-based nearest neighbor search across 4M product catalog entries, improving click-through rate by 17%. Scripted Airflow DAGs for model retraining and validation, automating 22 production workflows and reducing weekly engineering toil by 28 hours. Monitored Kubernetes-deployed services with Prometheus and Grafana, alerting on accuracy drift, latency SLO breaches, and feature distribution shifts across 16 models. Extracted named entities and topic clusters from customer call transcripts with spaCy, NLTK, and LDA, surfacing 60% more actionable insights for marketing teams. Embedded SHAP-based explainability layers into credit-risk and churn models, lifting stakeholder trust and accelerating model approval by risk committees. Migrated legacy on-prem analytics to AWS using EKS and Terraform, establishing CI/CD pipelines with Jenkins and shrinking model deployment time from several hours to under 20 minutes. Tech Stack: Python, Java, SQL, pandas, NumPy, scikit-learn, XGBoost, LightGBM, TensorFlow, PyTorch, PyTorch DDP, TensorFlow Serving, Hugging Face Transformers, BERT, spaCy, NLTK, Prophet, ARIMA, FAISS, SHAP, Apache Kafka, Spark Structured Streaming, Apache Airflow, AWS Redshift, EC2, S3, EKS, MLflow, Docker, Kubernetes, Jenkins, Terraform, Prometheus, Grafana. CERTIFICATIONS 1. AWS Certified Machine Learning Specialty 2. AWS Certified Solutions Architect Associate 3. Microsoft Certified Azure AI Engineer Associate 4. Databricks Certified Machine Learning Professional 5. Google Cloud Professional Machine Learning Engineer EDUCATION Webster University Master s in Data analytics 2022 August to 2024 May BIET, Hyderabad Bachelor s in technology 2016 June to 2020 May Keywords: cplusplus continuous integration continuous deployment artificial intelligence machine learning business intelligence sthree rlang golang Delaware |