Posted time June 7, 2026 Location Bangalore Job type Full-time

Experience: 5+ Years

Job Summary :

We are looking for an experienced Senior Software Engineer or Technical Lead to design and build scalable machine learning platforms for training, optimizing, evaluating, packaging, and deploying large-scale AI models.

The role requires strong expertise in distributed systems, GPU computing, ML infrastructure, and production-grade software engineering. You will work closely with applied scientists, platform engineers, hardware teams, compiler teams, and infrastructure specialists.

Key Skills & Expertise :

    • Distributed ML systems
  • GPU computing
  • ML infrastructure
  • Production-grade software engineering
  • Distributed ML training
  • Multi-node GPU clusters
  • Data, tensor, pipeline, and model parallelism
  • GPU utilization and memory efficiency
  • Communication performance and fault recovery
  • Model optimization
  • Quantization
  • Pruning
  • Knowledge distillation
  • Compression
  • Evaluation and artifact-management workflows
  • CI/CD
  • Regression testing
  • Observability
  • GPU workload profiling
  • System performance optimization
  • Production operations
  • System design and architecture
  • C++, Java, C#, Python
  • CUDA kernels or ML/low-level kernels
  • Containers
  • Kubernetes
  • AWS infrastructure
  • Technical leadership

Job Responsibilities :

  • Design and develop distributed ML platform services and reusable libraries
  • Build scalable training capabilities for large language and multimodal models
  • Support data, tensor, pipeline, and model parallelism across multi-node GPU clusters
  • Improve training throughput, GPU utilization, memory efficiency, communication performance, and fault recovery
  • Define stable APIs and platform interfaces for model onboarding, training, evaluation, and deployment
  • Integrate model optimization techniques such as quantization, pruning, knowledge distillation, and compression
  • Build evaluation and artifact-management workflows to measure model quality and system performance
  • Develop automated validation, CI/CD, regression testing, observability, and release pipelines
  • Profile GPU workloads and resolve end-to-end system performance bottlenecks
  • Build production-operational mechanisms, including metrics, alarms, dashboards, runbooks, and root-cause corrective actions
  • Collaborate with model, compiler, runtime, hardware, security, infrastructure, and product teams
  • Prepare technical designs, evaluate architecture trade-offs, and drive engineering decisions
  • Mentor engineers and improve code reviews, design reviews, testing, and development practices

Required Qualifications :

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field
  • 5+ years of professional software development experience
  • Strong programming skills in C++, Java, C#, Python, or a similar language
  • Experience programming with at least one modern language such as Java, C++, or C# including object-oriented design, or experience with CUDA kernels or ML/low-level kernels
  • Experience with containers, Kubernetes, AWS infrastructure, CI/CD, observability, and production operations
  • Experience designing distributed systems, high-performance computing platforms, or scalable backend systems
  • Strong understanding of system design, concurrency, reliability, scalability, and performance optimization
  • Experience as a technical lead, architect, mentor, or engineering team lead
  • Experience across the software development lifecycle, including coding, testing, source control, build systems, deployment, and production support

Preferred Qualifications :

  • Experience with PyTorch, TensorFlow, JAX, NeMo, or Megatron-LM
  • Experience building distributed ML training, inference, evaluation, or data platforms
  • Knowledge of CUDA, GPU kernels, GPU profiling, and performance optimization
  • Experience with containers, Kubernetes, cloud infrastructure, CI/CD, and observability tools
  • Knowledge of model compression, quantization, pruning, distillation, compilation, or edge AI deployment
  • Experience with large language models, multimodal models, and distributed GPU training
  • Experience designing extensible platform APIs and reusable software frameworks
  • Experience working with applied science, hardware, compiler, runtime, or product engineering teams

Key Skills :

  • Distributed ML Systems
  • GPU Computing
  • ML Infrastructure
  • PyTorch
  • CUDA
  • Kubernetes
  • CI/CD
  • Model Optimization
  • Performance Engineering
  • System Architecture
  • Technical Leadership

What We’re Looking For :

  • Experience: 5+ years of professional software development experience.
  • Technical Expertise: Strong expertise in distributed systems, GPU computing, ML infrastructure, and production-grade software engineering.
  • Programming Skills: Strong programming skills in C++, Java, C#, Python, or a similar language.
  • ML Platform Expertise: Experience designing distributed systems, high-performance computing platforms, or scalable backend systems.
  • Performance: Strong understanding of system design, concurrency, reliability, scalability, and performance optimization.
  • Technical Leadership: Experience as a technical lead, architect, mentor, or engineering team lead.
  • Collaboration: Ability to work closely with applied scientists, platform engineers, hardware teams, compiler teams, and infrastructure specialists.

Application :

Please send your CV to hr@nyxses.com

Website: www.nyxses.com