Home

Position: AI/ML Inference Engineer at Charlotte, North Carolina, USA
Email: [email protected]
http://bit.ly/4ey8w48
https://jobs.nvoids.com/job_details.jsp?id=3407757&uid=42a11e0b412f4a8d9ad287000f8b31b3

Need local to - Charlotte, NC only

TITLE-
AI/ML Inference Engineer

Location- Charlotte, NC 3 Open Roles | Hybrid 3 Days Onsite Need local

Duration: 3-6 Months

Interview- Onsite

Note Final interview will be In-person

AI/ML Inference Engineer Major Financial Services Organization

About the Role
A leading financial services institution is building a greenfield vLLM inference platform from the ground up. These three roles sit at the intersection of large language model serving, GPU infrastructure, and enterprise MLOps delivering the low-latency, high-throughput AI backbone powering the bank's next generation of intelligent products.
You will have direct ownership over inference stack performance on NVIDIA H200 GPU clusters in a regulated financial environment.

What You'll Do
Design, deploy, and optimize vLLM inference infrastructure on NVIDIA H200 GPU clusters using TensorRTLLM, Triton Inference Server, and SGLang
Implement advanced inference optimizations: continuous batching, speculative decoding, KV cache/prefix caching, and quantization (FP8, AWQ, GPTQ)
Architect and manage GPU orchestration on Kubernetes using KServe, OpenShift AI, Helm, and Run:AI
Leverage CUDA, NCCL, and MIG to maximize throughput and manage multi-GPU tensor parallelism
Build ML observability pipelines using Arize AI, Prometheus, and Grafana
Manage cloud-native infrastructure on GCP with Terraform
Conduct performance benchmarking and capacity planning for growing inference workloads

Required Qualifications
5+ years in ML engineering, MLOps, or AI infrastructure
Hands-on LLM inference serving experience (vLLM, TensorRTLLM, Triton, or equivalent)
NVIDIA H200 or comparable high-performance GPU cluster experience strongly preferred candidates with H200 experience will typically have working knowledge across the full inference stack
Kubernetes-based ML serving (KServe, OpenShift AI, Helm)
CUDA optimization, NCCL, and MIG partitioning experience
GCP and Terraform (or equivalent IaC)
Financial services or regulated industry background required

Preferred Qualifications
Arize AI or comparable ML observability platforms
Claude/Anthropic tooling or agentic AI workflows (Claude Cowork)
Open-source inference framework contributions or published benchmarking results

Immediate Start Industry: Financial Services / Regulated Enterprise

Thank you,

Ashish Kumar | Technical Recruiter

Verve IT Consulting Inc

O: 917-259-0969 EXT- 119

Keywords: artificial intelligence machine learning information technology card North Carolina
Position: AI/ML Inference Engineer
[email protected]
http://bit.ly/4ey8w48
https://jobs.nvoids.com/job_details.jsp?id=3407757&uid=42a11e0b412f4a8d9ad287000f8b31b3
[email protected]
View All
11:52 PM 28-May-26


To remove this job post send "job_kill 3407757" as subject from [email protected] to [email protected]. Do not write anything extra in the subject line as this is a automatic system which will not work otherwise.

Pages not loading, taking too much time to load, server timeout or unavailable, or any other issues please contact admin at [email protected]


Time Taken: 8

Location: Charlotte, North Carolina