| Position: AI/ML Inference Engineer at Charlotte, North Carolina, USA |
| Email: [email protected] |
|
http://bit.ly/4ey8w48 https://jobs.nvoids.com/job_details.jsp?id=3407757&uid=42a11e0b412f4a8d9ad287000f8b31b3 Need local to - Charlotte, NC only TITLE- AI/ML Inference Engineer Location- Charlotte, NC 3 Open Roles | Hybrid 3 Days Onsite Need local Duration: 3-6 Months Interview- Onsite Note Final interview will be In-person AI/ML Inference Engineer Major Financial Services Organization About the Role A leading financial services institution is building a greenfield vLLM inference platform from the ground up. These three roles sit at the intersection of large language model serving, GPU infrastructure, and enterprise MLOps delivering the low-latency, high-throughput AI backbone powering the bank's next generation of intelligent products. You will have direct ownership over inference stack performance on NVIDIA H200 GPU clusters in a regulated financial environment. What You'll Do Design, deploy, and optimize vLLM inference infrastructure on NVIDIA H200 GPU clusters using TensorRTLLM, Triton Inference Server, and SGLang Implement advanced inference optimizations: continuous batching, speculative decoding, KV cache/prefix caching, and quantization (FP8, AWQ, GPTQ) Architect and manage GPU orchestration on Kubernetes using KServe, OpenShift AI, Helm, and Run:AI Leverage CUDA, NCCL, and MIG to maximize throughput and manage multi-GPU tensor parallelism Build ML observability pipelines using Arize AI, Prometheus, and Grafana Manage cloud-native infrastructure on GCP with Terraform Conduct performance benchmarking and capacity planning for growing inference workloads Required Qualifications 5+ years in ML engineering, MLOps, or AI infrastructure Hands-on LLM inference serving experience (vLLM, TensorRTLLM, Triton, or equivalent) NVIDIA H200 or comparable high-performance GPU cluster experience strongly preferred candidates with H200 experience will typically have working knowledge across the full inference stack Kubernetes-based ML serving (KServe, OpenShift AI, Helm) CUDA optimization, NCCL, and MIG partitioning experience GCP and Terraform (or equivalent IaC) Financial services or regulated industry background required Preferred Qualifications Arize AI or comparable ML observability platforms Claude/Anthropic tooling or agentic AI workflows (Claude Cowork) Open-source inference framework contributions or published benchmarking results Immediate Start Industry: Financial Services / Regulated Enterprise Thank you, Ashish Kumar | Technical Recruiter Verve IT Consulting Inc O: 917-259-0969 EXT- 119 Keywords: artificial intelligence machine learning information technology card North Carolina Position: AI/ML Inference Engineer [email protected] http://bit.ly/4ey8w48 https://jobs.nvoids.com/job_details.jsp?id=3407757&uid=42a11e0b412f4a8d9ad287000f8b31b3 |
| [email protected] View All |
| 11:52 PM 28-May-26 |