В приложении удобнее
Установить
Soniox

Software Engineer, LLM Inference

в Soniox

📍 Любляна (Словения)
Гибрид
📍 Любляна (Словения)
Помощь с переездом
Специализация
Data Scientist & Machine Learning
Уровень
Senior
Английский
B2 — Upper-Intermediate

Технологии/инструменты

CUDA
PyTorch
Triton

We’ll help you with the relocation process and paperwork from any country to Slovenia.

About the role

Soniox is pushing the boundaries of real-time AI, and we’re looking for an engineer to help us run large language models with exceptional speed, efficiency, and reliability at production scale.

In this role, you’ll work deep in the LLM inference stack, from vLLM scheduling and KV-cache management to CUDA kernels, distributed execution, and GPU profiling, optimizing every part of the path from request to generated token.

In this role, you will

  • Build and optimize our vLLM-based inference stack for low latency, high throughput, and maximum GPU utilization.
  • Optimize continuous batching, scheduling, prefill/decode, prefix caching, and KV-cache allocation and reuse.
  • Optimize or implement CUDA and Triton kernels for attention, GEMMs, sampling, normalization, and other critical model operations.
  • Evaluate and integrate technologies such as FlashAttention, FlashInfer, CUDA Graphs, torch.compile, speculative decoding, and quantization.
  • Optimize distributed inference using tensor, data, and expert parallelism, NCCL, NVLink/NVSwitch, and InfiniBand.
  • Work closely with researchers to bring new dense and MoE model architectures into production quickly and efficiently.

You might thrive in this role if you

  • Have deep hands-on experience with LLM inference, ideally working inside vLLM, SGLang, TensorRT-LLM, or similar systems, not just deploying them.
  • Understand TTFT, inter-token latency, throughput, continuous batching, PagedAttention, KV caching, and prefill vs. decode performance.
  • Are comfortable profiling GPUs and reasoning about compute, memory bandwidth, kernel launches, synchronization, and communication bottlenecks.
  • Have experience with CUDA, Triton, PyTorch, NCCL, and modern NVIDIA GPU architectures.
  • Understand Transformer internals including MHA/GQA, RoPE, KV cache, quantization, and MoE.
  • Have experience optimizing distributed, performance-critical systems in production.
  • Care deeply about performance, simplicity, and reliability, and take ownership from profiling through production deployment.

Why Soniox

You’ll help build one of the most technically advanced voice AI platforms in the world, and push LLM inference performance at every layer of the stack.

You’ll work directly with a world-class team of engineers and researchers on hard, measurable problems spanning models, GPU kernels, distributed systems, and production infrastructure.

You'll have a voice in how our technology evolves, how our company grows, and how AI transforms human communication.

Soniox

О компании Soniox

Сфера
Технологии, информационные средства и Интернет
Размер
11 - 50

Мы верим, что речь является самым естественным способом общения между людьми. Большинство существующих решений хорошо работают только для части языков, рынков и сценариев. Soniox был создан, чтобы это изменить.

Мы разрабатываем фундаментальный Voice AI, который помогает разработчикам, компаниям и обычным пользователям понимать и использовать речь независимо от языка, акцента или окружающих условий.

Сегодня Soniox предоставляет единую платформу для speech-to-text, text-to-speech и перевода речи в реальном времени. Мы используем только собственные модели, поддерживаем более 60 языков и созданы для глобальных продуктов, где особенно важны точность и минимальная задержка.

Мы фокусируемся на самых сложных задачах голосового ИИ: многоязычной речи, переключении языков внутри одного разговора, нескольких собеседниках, именах, числах, профессиональной лексике и сложном аудио. Все это работает как для распознавания речи в реальном времени, так и для ее генерации.

Soniox уже используют Perplexity, Samsung, LG, LiveKit, Krisp, Fireflies.ai, Wispr Flow, Truecaller, Vapi, Retell и тысячи других компаний и разработчиков. На нашей технологии создают голосовых агентов, продукты для встреч и контакт-центров, медицинские решения, диктовку, перевод и другие приложения, где голос становится основным способом взаимодействия с технологиями.

Soniox быстро растет, и вместе с продуктом растет наша команда. Мы ищем сильных специалистов в своих областях, которые хотят решать сложные технические задачи и вместе с нами строить голосовые технологии мирового уровня.