Build and operate production machine learning and generative AI infrastructure. Responsibilities include automating training and deployment pipelines, managing Kubernetes-based model serving, tracking models and data versions, monitoring drift and performance, optimizing inference with ONNX and TensorRT, supporting LLMOps and vector databases, and implementing security controls. Requires strong Python, cloud, container, ML framework, SQL, distributed systems, and GPU expertise, plus a mandatory ML certification.
This is a remote position.
AI / ML Ops Engineer
Job Details
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 4 to 8 years
- Relevant Experience Required: 3+ years of dedicated MLOps or DevOps experience deploying and managing machine learning models in production
- Mandatory Certification: AWS Certified Machine Learning - Specialty, Google Cloud Certified Professional Machine Learning Engineer, or Databricks Certified Machine Learning Professional
Job Summary
We are seeking an experienced AI / ML Ops Engineer to bridge the gap between data science research and cloud infrastructure execution. The ideal candidate will build automated pipelines to train, test, deploy, and monitor machine learning models and generative AI systems securely and at scale within enterprise container environments.
Key Responsibilities
- Design and automate end-to-end ML pipelines (Continuous Training and Continuous Deployment) using orchestration engines like Kubeflow, MLflow, or AWS SageMaker Pipelines.
- Orchestrate containerized model deployments, configuring low-latency inference endpoints, auto-scaling GPU/CPU clusters, and model serving runtimes on Kubernetes (KServe, Triton Inference Server).
- Implement robust model tracking and data versioning foundations, managing feature stores (e.g., Feast), model registries, and version controls for massive datasets using DVC.
- Build automated AI performance and data monitoring gates, tracking model accuracy decay, data drift indicators, concept drift parameters, and system processing latencies in real time.
- Optimize inference execution environments, leveraging model compilation engines (e.g., ONNX, TensorRT) and quantization strategies to shorten response times and minimize cloud compute costs.
- Integrate generative AI and LLM operational frameworks (LLMOps), configuring semantic caching layers, vector database scaling parameters (e.g., Pinecone, Milvus), and prompt validation pipelines.
- Govern machine learning access controls and security profiles, configuring strict data separation barriers, model access tokens, and encryption protocols to safeguard sensitive inference logs.
Requirements
- 4 to 8 years of core software engineering, DevOps, or data engineering experience, with 3+ dedicated years actively building and maintaining MLOps automation infrastructures.
- Strong technical mastery of Python programming, container orchestration (Docker, Kubernetes), ML frameworks (PyTorch, TensorFlow, Hugging Face), and advanced SQL.
- Deep structural understanding of distributed system mechanics, GPU resource management limits, model deployment patterns (Shadow, Canary, A/B), and cloud provider API governance.
- Mandatory certification: AWS Machine Learning Specialty, Google Cloud Professional ML Engineer, or Databricks ML Professional.
Preferred Qualifications
- Prior experience implementing RAG (Retrieval-Augmented Generation) pipelines or fine-tuning open-source LLM layers in production.
- Familiarity with infrastructure-as-code scripting tools like Terraform to provision ML cluster topologies.
Similar Jobs
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads the architecture, development, and productionization of AI/ML and GenAI solutions for healthcare operations. Responsibilities include building clinical and claims data pipelines, developing LLM and RAG applications, implementing MLOps and LLMOps, ensuring HIPAA-compliant governance, translating business needs into technical roadmaps, communicating with stakeholders, and mentoring engineers.
Top Skills:
AirflowAWSAzureAzure MlBigQueryCi/CdDaskDatabricksDockerFaissGCPGitGoHugging FaceJavaKafkaKubernetesLangchainLlamaindexMlflowNeo4JNumpyPandasPgvectorPineconePrefectPytestPythonPyTorchSagemakerScalaSnowflakeSparkSQLTensorFlowTerraformTransformersVertex Ai
Artificial Intelligence • Productivity • Software • Automation
Leads Zapier’s India entity as statutory Board Director and General Manager. Owns local operations, governance, legal, tax, audit, transfer pricing, regulatory compliance, risk management, vendor relationships, and people decisions. Serves as the bridge between India-based teams and global leadership, translating strategy into execution and establishing operating rhythms for a fully remote workforce. Partners with Deloitte, HR, finance, auditors, regulators, and global executives while overseeing hiring, performance, compensation, employee relations, and compliance with Indian labor requirements.
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Leads industrial engineering for semiconductor assembly and test manufacturing, including capacity and resource planning, labor optimization, operational excellence, cost reduction, manufacturing analytics, and factory systems. Develops capacity models, staffing standards, dashboards, and automation initiatives while supporting budgets and capital planning. Builds and mentors an industrial engineering team and collaborates globally to standardize manufacturing practices and improve efficiency, utilization, cycle time, and cost performance.
Top Skills:
Advanced AnalyticsAIManufacturing SimulationMesExcelPower BISQLTableau
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.



.jpeg)