The AI Platform Engineer will design and optimize AI/ML infrastructure, automate pipelines, and manage Kubernetes for effective machine learning model deployment and collaboration.
AHEAD builds platforms for digital business. By weaving together advances in cloud infrastructure, automation and analytics, and software delivery, we help enterprises deliver on the promise of digital transformation.
At AHEAD, we prioritize creating a culture of belonging, where all perspectives and voices are represented, valued, respected, and heard. We create spaces to empower everyone to speak up, make change, and drive the culture at AHEAD.
We are an equal opportunity employer, and do not discriminate based on an individual's race, national origin, color, gender, gender identity, gender expression, sexual orientation, religion, age, disability, marital status, or any other protected characteristic under applicable law, whether actual or perceived.
We embrace all candidates that will contribute to the diversification and enrichment of ideas and perspectives at AHEAD.
We are seeking an experienced AI Platform Engineer to design, deploy, and optimize AI/ML infrastructure, AI workflows, and automated pipelines. This role focuses on building scalable environments for training and deploying machine learning models, leveraging modern orchestration, automation, and GPU acceleration technologies. You will collaborate with data scientists and platform engineers to drive efficient resource utilization and scalable operations across cloud and hybrid environments.
Key Responsibilities
- Kubernetes for AI/ML: Architect and manage Kubernetes clusters tailored to AI/ML workloads.
- GPU Orchestration: Implement Run:ai and operators for GPU resource orchestration and workload scheduling.
- Automation & Pipelines: Develop and maintain Python-based automation scripts and ML pipelines; automate infrastructure provisioning with Terraform and configuration management with Ansible.
- Notebooks & Collaboration: Create and manage Jupyter Notebooks for experimentation and collaboration.
- NVIDIA Integration: Integrate and optimize NVIDIA Enterprise Suite components (CUDA, NeMo Framework, Triton, TensorRT, GPU drivers) for accelerated computing.
- MLOps Practices: Establish and maintain MLOps best practices for model lifecycle management, CI/CD, and monitoring (e.g., MLflow, Kubeflow).
- Collaboration: Work closely with data scientists and platform engineers to ensure efficient resource utilization and scalability across environments.
Required Skills & Experience
- Strong proficiency in Python and experience with ML frameworks (TensorFlow, PyTorch).
- Hands-on experience with Kubernetes and container orchestration.
- Familiarity with Run:ai or similar GPU scheduling platforms.
- Expertise in Terraform and Ansible for infrastructure automation.
- Experience with Jupyter Notebooks for ML development.
- Knowledge of NVIDIA Enterprise Suite (CUDA, NeMo Framework, Triton, GPU drivers).
- Solid understanding of MLOps principles and tools (e.g., MLflow, Kubeflow).
- Background in deploying and scaling AI workloads in cloud or hybrid environments.
Qualifications
- 4+ years in platform architecture or solutions architecture, with 2+ years focused on AI/ML workloads.
- Experience with high-performance computing (HPC) environments.
- Familiarity with distributed training and model optimization techniques.
- Certification in Kubernetes or cloud platforms (AWS, Azure, GCP).
Why AHEAD:
Through our daily work and internal groups like Moving Women AHEAD and RISE AHEAD, we value and benefit from diversity of people, ideas, experience, and everything in between.
We fuel growth by stacking our office with top-notch technologies in a multi-million-dollar lab, by encouraging cross department training and development, sponsoring certifications and credentials for continued learning.
India Employment Benefits include:
Comprehensive health insurance coverage for employees, with options to extend coverage to dependents
Paid time off and company holidays, along with additional leave benefits as per policy
Flexible work arrangements, supporting work-life balance
Learning and development opportunities to support continuous growth and upskilling
Employee wellness initiatives and programs focused on physical and mental well-being
Retirement and statutory benefits in line with India regulations
Inclusive and people-first culture, with a strong focus on collaboration and ownership
AHEAD Gurugram, Haryana, IND Office
One Horizon Center Golf Course Road, DLF Phase V Sector 43, Gurugram, India, 122002
Similar Jobs
Cloud • Information Technology • Security • Software • Cybersecurity
Design, scale, and maintain secure, highly available AWS infrastructure and platform capabilities for production AI/ML workloads. Build Terraform-based IaC, manage GitLab CI/CD pipelines, implement centralized observability with Prometheus/Grafana and alerting, drive DORA metrics, enable autoscaling/self-healing Kubernetes deployments, enforce platform governance and security best practices, and mentor engineering teams.
Top Skills:
AlertmanagerAws EcsAws EksAws IamAws LambdaAws S3Aws VpcBashCluster AutoscalerGitlab Ci/CdGitlab RunnersGoGrafanaHelmHorizontal Pod Autoscaler (Hpa)KarpenterKubernetes (Eks)OpsgeniePagerdutyPrometheusPythonTerraform
Automotive
Design, build, integrate, and operate enterprise-scale AI security and DevSecOps platforms. Embed security into developer workflows and CI/CD, automate security processes, support vulnerability remediation, enable secure AI development, and drive adoption of emerging security technologies and developer enablement.
Top Skills:
AWSAzureAzure DevopsCi/CdContainer SecurityGCPGithub ActionsGithub EnterpriseGitlab Ci/CdInfrastructure As CodeKubernetesPowershellPythonSecrets Detection/ManagementShell ScriptingSoftware Composition Analysis (Sca)Software Supply Chain SecurityStatic Application Security Testing (Sast)TektonTerraform
Artificial Intelligence • Information Technology • Professional Services • Consulting
Design and build realistic cloud and backend environments to evaluate AI-driven tasks. Create distributed-system scenarios covering deployment, troubleshooting, scaling, security, disaster recovery, observability, and validation tests. Produce reference implementations and faulty configurations, document architectures and procedures, and collaborate with technical leads to refine specifications.
Top Skills:
AuthenticationAuthorizationC++Ci/CdCloud InfrastructureDatabasesDevOpsDurable StorageGoIamInfrastructure AutomationJavaJavaScriptObservabilityPythonQueuesRustTypescript
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.



