MlOps Engineer Job
Employer: Sandeep Kanchi
SpiderID: 14218242
Location: Santa Clara, California
Posted: 8/3/2026
Wage:
Priority Review Date: 11/1/2026
Job Code / NOC / SOC:
Category: General
Job Description:
Apply here - https://tinyurl.com/52zawa2b
Must-Have Skills:
5-7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
Strong proficiency in Python and software engineering best practices.
Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and Continuous/Dynamic Batching.
Demonstrated hands-on experience building, deploying, troubleshooting, and optimizing production-grade LLM and GenAI solutions.
Experience deploying, scaling, and monitoring production-grade GenAI/LLM applications.
Exposure to AI Observability, Governance, and Responsible AI practices.
Good-to-Have Skills:
Hands-on experience with LLM Fine-Tuning using PEFT, SFT, CPT, LoRA, and QLoRA techniques.
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT.
Knowledge of distributed training and multi-GPU environments.
Experience with Agentic AI frameworks such as LangGraph, AutoGen, or CrewAI.
Understanding of simulation platforms, digital twins, modeling & simulation workflows, or scientific computing.
Must-Have Skills:
5-7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
Strong proficiency in Python and software engineering best practices.
Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and Continuous/Dynamic Batching.
Demonstrated hands-on experience building, deploying, troubleshooting, and optimizing production-grade LLM and GenAI solutions.
Experience deploying, scaling, and monitoring production-grade GenAI/LLM applications.
Exposure to AI Observability, Governance, and Responsible AI practices.
Good-to-Have Skills:
Hands-on experience with LLM Fine-Tuning using PEFT, SFT, CPT, LoRA, and QLoRA techniques.
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT.
Knowledge of distributed training and multi-GPU environments.
Experience with Agentic AI frameworks such as LangGraph, AutoGen, or CrewAI.
Understanding of simulation platforms, digital twins, modeling & simulation workflows, or scientific computing.
Contact Information:
| Contact Name: Sandeep Kanchi | Type: Recruiter |
| Company: | |
| Web Site: https://tinyurl.com/52zawa2b | |