NVIDIA
Teams at NVIDIA
Recently posted jobs
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve the reliability, availability, and performance of large-scale GeForce NOW services. Participate in incident triage and on-call rotations, build automation and tooling, enhance observability (metrics/logs/traces), drive SLO/SRI practices, run postmortems, and design/operate Kubernetes-based services across cloud and datacenter environments.
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and operate a Kubernetes-based platform for global network infrastructure. Own cluster lifecycle, provisioning, upgrades, GitOps delivery, observability, capacity, and recovery. Develop automation, provide production support and on-call incident response for network services, and drive issues from detection through verified resolution.
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead design, operation, and reliability of large-scale AI/HPC clusters. Manage day-to-day operations, incident response, automation, performance tuning, scheduler and storage optimization, and collaborate with researchers and global teams to improve GPU-accelerated infrastructure.
