Experience: 5+ Years
Job Summary :
We are looking for an experienced Senior Software Engineer or Technical Lead to design and build scalable machine learning platforms for training, optimizing, evaluating, packaging, and deploying large-scale AI models.
The role requires strong expertise in distributed systems, GPU computing, ML infrastructure, and production-grade software engineering. You will work closely with applied scientists, platform engineers, hardware teams, compiler teams, and infrastructure specialists.
Key Skills & Expertise :
-
- Distributed ML systems
- GPU computing
- ML infrastructure
- Production-grade software engineering
- Distributed ML training
- Multi-node GPU clusters
- Data, tensor, pipeline, and model parallelism
- GPU utilization and memory efficiency
- Communication performance and fault recovery
- Model optimization
- Quantization
- Pruning
- Knowledge distillation
- Compression
- Evaluation and artifact-management workflows
- CI/CD
- Regression testing
- Observability
- GPU workload profiling
- System performance optimization
- Production operations
- System design and architecture
- C++, Java, C#, Python
- CUDA kernels or ML/low-level kernels
- Containers
- Kubernetes
- AWS infrastructure
- Technical leadership
Job Responsibilities :
- Design and develop distributed ML platform services and reusable libraries
- Build scalable training capabilities for large language and multimodal models
- Support data, tensor, pipeline, and model parallelism across multi-node GPU clusters
- Improve training throughput, GPU utilization, memory efficiency, communication performance, and fault recovery
- Define stable APIs and platform interfaces for model onboarding, training, evaluation, and deployment
- Integrate model optimization techniques such as quantization, pruning, knowledge distillation, and compression
- Build evaluation and artifact-management workflows to measure model quality and system performance
- Develop automated validation, CI/CD, regression testing, observability, and release pipelines
- Profile GPU workloads and resolve end-to-end system performance bottlenecks
- Build production-operational mechanisms, including metrics, alarms, dashboards, runbooks, and root-cause corrective actions
- Collaborate with model, compiler, runtime, hardware, security, infrastructure, and product teams
- Prepare technical designs, evaluate architecture trade-offs, and drive engineering decisions
- Mentor engineers and improve code reviews, design reviews, testing, and development practices
Required Qualifications :
- Bachelor’s degree in Computer Science, Engineering, or a related technical field
- 5+ years of professional software development experience
- Strong programming skills in C++, Java, C#, Python, or a similar language
- Experience programming with at least one modern language such as Java, C++, or C# including object-oriented design, or experience with CUDA kernels or ML/low-level kernels
- Experience with containers, Kubernetes, AWS infrastructure, CI/CD, observability, and production operations
- Experience designing distributed systems, high-performance computing platforms, or scalable backend systems
- Strong understanding of system design, concurrency, reliability, scalability, and performance optimization
- Experience as a technical lead, architect, mentor, or engineering team lead
- Experience across the software development lifecycle, including coding, testing, source control, build systems, deployment, and production support
Preferred Qualifications :
- Experience with PyTorch, TensorFlow, JAX, NeMo, or Megatron-LM
- Experience building distributed ML training, inference, evaluation, or data platforms
- Knowledge of CUDA, GPU kernels, GPU profiling, and performance optimization
- Experience with containers, Kubernetes, cloud infrastructure, CI/CD, and observability tools
- Knowledge of model compression, quantization, pruning, distillation, compilation, or edge AI deployment
- Experience with large language models, multimodal models, and distributed GPU training
- Experience designing extensible platform APIs and reusable software frameworks
- Experience working with applied science, hardware, compiler, runtime, or product engineering teams
Key Skills :
- Distributed ML Systems
- GPU Computing
- ML Infrastructure
- PyTorch
- CUDA
- Kubernetes
- CI/CD
- Model Optimization
- Performance Engineering
- System Architecture
- Technical Leadership
What We’re Looking For :
- Experience: 5+ years of professional software development experience.
- Technical Expertise: Strong expertise in distributed systems, GPU computing, ML infrastructure, and production-grade software engineering.
- Programming Skills: Strong programming skills in C++, Java, C#, Python, or a similar language.
- ML Platform Expertise: Experience designing distributed systems, high-performance computing platforms, or scalable backend systems.
- Performance: Strong understanding of system design, concurrency, reliability, scalability, and performance optimization.
- Technical Leadership: Experience as a technical lead, architect, mentor, or engineering team lead.
- Collaboration: Ability to work closely with applied scientists, platform engineers, hardware teams, compiler teams, and infrastructure specialists.
Application :
Please send your CV to hr@nyxses.com
Website: www.nyxses.com