
Remote opportunity at
NebiusSenior ML Engineer (Token Factory)
Nebius seeks a Senior ML Engineer (Token Factory) for a full-time, remote position to join Nebius Cloud, which operates one of the largest GPU clouds globally.…
Career Tools
About This Role
Nebius seeks a Senior ML Engineer (Token Factory) for a full-time, remote position to join Nebius Cloud, which operates one of the largest GPU clouds globally. The employer provides an AI cloud platform enabling developers and enterprises to handle everything from data and model training to production deployment without building internal infrastructure. Headquartered in Amsterdam with a Nasdaq listing (NBIS), Nebius maintains a workforce of over 1,500 people and R&D…
Job Description
Nebius seeks a Senior ML Engineer (Token Factory) for a full-time, remote position to join Nebius Cloud, which operates one of the largest GPU clouds globally. The employer provides an AI cloud platform enabling developers and enterprises to handle everything from data and model training to production deployment without building internal infrastructure. Headquartered in Amsterdam with a Nasdaq listing (NBIS), Nebius maintains a workforce of over 1,500 people and R&D hubs spanning Europe, the UK, North America, and Israel.
In this capacity, the engineer will work within the Token Factory initiative, focusing on high-performance inference and fine-tuning platforms designed to maximize throughput, minimize latency, and optimize cost-per-token across tens of thousands of GPUs. Key areas of focus include identifying LLM inference bottlenecks for speedups, supporting architectures like GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, and GLM-5, implementing speculative decoding, and productionizing low-precision pipelines using FP8, NVFP4, and MXFP4.
This opportunity is well-suited for machine learning professionals with a strong background in transformer architecture, GPU performance profiling, and large neural network training. The company offers a collaborative international environment characterized by constant growth, flexibility, ownership, and impactful AI projects.
Responsibilities
- Identify LLM inference bottlenecks to drive production speedups across various LLM architectures at scale
- Implement novel speculative decoding architectures and optimize components of dense, mixture-of-experts, autoregressive, and parallel LLM designs
- Contribute to open-source inference engines
- Design and productionize low-precision training and inference pipelines including FP8, NVFP4, and MXFP4 with measurable gains in throughput and cost-efficiency
Requirements
- Profound understanding of machine learning theoretical foundations and transformer architecture
- Experience profiling GPU workloads utilizing Nsight, PyTorch profiler, or similar tools
- Comprehension of GPU memory hierarchy and compute-memory tradeoffs
- Familiarity with concepts such as MHA, RoPE, KV-cache, Flash Attention, and quantization
- Knowledge of performance aspects related to large neural network training including sharding strategies, custom kernels, and hardware features
- Strong software engineering skills with primary usage of Python
- Deep experience utilizing modern deep learning frameworks
- Proficiency in software engineering approaches such as CI/CD, version control, and unit testing
- Strong communication and leadership abilities
- Authorization to work in the country of application with proof of eligibility required upon hire
Qualifications
- Experience working with and contributing to open-source inference engines such as vLLM, SGLang, and TensorRT-LLM
- Experience with kernel languages or DSLs including Triton, Cute, CUTLASS, and CUDA
- Track record of building and delivering products within a dynamic, startup-like environment
- Experience developing large distributed systems or high-load web services
- Open-source projects showcasing engineering capabilities
- Excellent command of the English language with superior writing, articulation, and communication skills
Core Skills
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Frequently Asked Questions
Answers are based only on the employer’s listing; where it doesn’t say, neither do we.
What is the location and remote status for this role?
The position is remote.
What is the employment type?
This is a full-time employment opportunity.
What is the salary for the Senior ML Engineer position?
The salary is not stated in the job posting.
What are the main technical focus areas for the Token Factory team?
The team focuses on high-performance inference and fine-tuning platforms, maximizing throughput, minimizing latency, optimizing cost-per-token across tens of thousands of GPUs, inference optimization, and low-precision training and inference pipelines.
Sample Interview Questions
AI-generated questions tailored to this specific role — a preview of the full practice set.