logo

JobNob

Your Career. Our Passion.

Senior MLOps Engineer


Quantiphi


Location

Trivandrum | India


Job description

Job Summary We are seeking an experienced Platform Engineer with expertise in MLOps and handling distributed systems, particularly Kubernetes and Slurm, along with a strong background in managing Multi-GPU, Multi-Node Deep Learning job scheduling. Proficiency in Linux (Ubuntu) systems, the ability to create intricate shell scripts, good proficiency in working with configuration management tools and sufficient understanding of deep learning workflow.

Location: Mumbai, Bangalore, Trivandrum

What we Offer ● The opportunity to solve tough business problems with innovative AI solutions. ● A fast-paced environment where you can make a difference. ● The ability to work and grow with an incredible and fun team. ● Highly competitive compensation (and other perks, of course).

Scope of Responsibilities ● Design, deploy, and maintain distributed systems using Kubernetes and Slurm for optimal resource utilization and workload management. ● Lead the configuration and optimization of Multi-GPU, Multi-Node Deep Learning job scheduling, ensuring efficient computation and data processing. ● Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions. ● Experience in working with On-prem NVIDIA GPU servers. ● Develop and maintain complex shell scripts for various system automation tasks, enhancing efficiency and reducing manual intervention. ● Monitor system performance, identify bottlenecks, and implement necessary adjustments to ensure high availability and reliability. ● Troubleshoot and resolve technical issues related to the distributed system, job scheduling, and deep learning processes. ● Stay updated with industry trends and emerging technologies in distributed systems, deep learning, and automation.

Job Requirements Skill Set Needed: ● Strong communication and collaboration skills to work effectively within a cross-functional team. ● Python. Hands-on experience in MLOps. ● Good to have at least one ML framework understanding - PyTorch / TF. ● Experience in shell scripting. ● Good understanding of logical networks. ● Understanding of CV, NLP. Cloud native stack. ● Proven experience in designing, deploying, and managing distributed systems, with a focus on Kubernetes and Slurm. ● Sufficient understanding of AI Model Training and Deployment and Strong background in Multi-GPU, Multi-Node Deep Learning job scheduling and resource management. ● Proficiency in Linux systems, particularly Ubuntu, and the ability to navigate and troubleshoot related issues. ● Extensive experience creating complex shell scripts for automation and system orchestration. ● Familiarity with continuous integration and deployment (CI/CD) processes. ● Excellent problem-solving skills and the ability to diagnose and resolve technical issues promptly.

Good to have: ● Previously working on NVIDIA Ecosystem or well aware of NVIDIA Ecosystem. ● Good at Slurm, Kubernetes, Linux, and AI Deployment tools.


Job tags



Salary

All rights reserved