Agilerobotsse
Munich, DE · On-site · Full-time
<div class="TranslationField-module__fieldItem___g4pRX"> <p><strong>The AI Teams at Agile Robots</strong> are looking for an <strong>ML Platform Engineer (m/f/d)</strong>, who will build and operate the distributed training, deployment, and experimentation infrastructure that research, data, and robotics teams depend on to move models from prototype to production.</p> </div> <div class="TranslationField-module__fieldItem___g4pRX"> <h3 class="TokenisedTypography-module__EzGiEkH0__v6-3-8 TranslationField-module__fieldItemName___Q0F_7">Your Responsibilities</h3> <ul> <li><strong>Training Infrastructure:</strong> Design and scale distributed training workflows for large models using tools such as PyTorch Distributed, DeepSpeed, and cluster schedulers like SLURM or Kubernetes.</li> <li><strong>ML Platform:</strong> Build and maintain containerised ML environments that support reproducible experimentation and benchmarking.</li> <li><strong>CI/CD Pipelines:</strong> Develop and maintain CI/CD pipelines for machine learning systems to enable reliable testing, training, and deployment of models.</li> <li><strong>Lifecycle Management:</strong> Implement experiment tracking, model versioning, and reproducibility workflows using tools such as ClearML or Weights & Biases.</li> <li><strong>Observability:</strong> Set up monitoring systems such as Prometheus and Grafana to track model performance and system health and detect drift in production.</li> <li><strong>Cross-Team Collaboration:</strong> Work with research, data, and robotics teams to connect new models to robust production systems.</li> </ul> </div> <div class="TranslationField-module__fieldItem___g4pRX"> <h3 class="TokenisedTypography-module__EzGiEkH0__v6-3-8 TranslationField-module__fieldItemName___Q0F_7">Essential Skills</h3> <ul> <li><strong>Background and Experience:</strong> Degree in Computer Science, Software Engineering, or a related field, with professional experience building and operating ML or software infrastructure in production.</li> <li><strong>Distributed Training:</strong> Experience designing and operating distributed training systems on Kubernetes and Docker, using PyTorch Distributed, DeepSpeed, and schedulers such as SLURM.</li> <li><strong>CI/CD for ML:</strong> Experience building CI/CD pipelines that support reliable model testing, training, and deployment.</li> <li><strong>Cloud Infrastructure:</strong> Experience operating ML workloads on cloud infrastructure, preferably AWS.</li&g
Agilerobotsse
Posted via Arbeitnow
Apply Now takes you to Rozgoo, where auto-apply can submit your application for this role. Updated 9 days ago.
Apply Now