Distributed Training Engineer - Remote Work | REF#298260
At BairesDev, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley.
Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
As a Distributed Training Engineer, you will own the scaling and optimization of training workloads across multi-GPU and multi-node environments. You will provide the necessary systems fluency to handle massive datasets and models, ensuring peak performance through mixed-precision and distributed coordination.
What You'll Do
Apply now!
Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
As a Distributed Training Engineer, you will own the scaling and optimization of training workloads across multi-GPU and multi-node environments. You will provide the necessary systems fluency to handle massive datasets and models, ensuring peak performance through mixed-precision and distributed coordination.
What You'll Do
- Implement and optimize distributed training strategies using PyTorch DDP, FSDP, and DeepSpeed.
- Profile training performance and GPU utilization using Nvidia ecosystem tools like Nsight.
- Design mixed-precision and memory-optimization workflows to maximize training throughput and reduce memory footprint.
- Solve complex coordination and synchronization issues across multi-node clusters using CUDA and Triton.
- Scale massive models through parameter sharding, pipeline parallelism, and efficient data sharding techniques.
- Partner with machine learning research teams to align infrastructure capabilities with evolving model scaling requirements.
- 4+ years of experience in Machine Learning Engineering or Distributed Systems.
- Proven expertise in distributed training frameworks like PyTorch DDP, FSDP, or DeepSpeed.
- Hands-on experience with Nvidia GPU profiling tools such as Nsight.
- Strong proficiency in CUDA programming and performance optimization for multi-node clusters.
- Experience with inference optimization tools like Triton or TensorRT.
- Advanced proficiency in English.
- 100% remote work (from anywhere).
- Excellent compensation in USD or your local currency if preferred
- Hardware and software setup for you to work from home.
- Flexible hours: create your own schedule.
- Paid parental leaves, vacations, and national holidays.
- Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
- Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities.
Apply now!
Empleos Recomendados
Executive Assistant (Junior)
Publicado hace 7 horas
Líder de Soporte y Desarrollo
Publicado hace 7 horas
Ingeniero/a de senior de GRC
Publicado hace 8 horas
Business Development Representative (Junior)
Publicado hace 8 horas
Marketing Specialist (Junior)
Publicado hace 8 horas

