AI Inference Engineer, QVAC at Tether
Confirmed open
Optimise C++ inference engines including llama.cpp and ggml for fast private AI on edge devices from Hanoi.
- Location
- Hanoi, Vietnam
- Employment
- fulltime_permanent
- Work arrangement
- remote
- Technologies
- JavaScript, C++
Role overview
The QVAC inference engineer works on the C++ runtime that runs local AI models without relying on cloud infrastructure. The role improves model loading, memory use and execution across edge hardware. The posting lists Hanoi, Vietnam.
Responsibilities
- Deploy machine-learning models to edge devices with llama.cpp, ggml and relevant GPU frameworks.
- Improve inference speed, memory use and runtime stability across hardware platforms.
- Work with researchers to turn models into production implementations.
- Integrate on-device AI features into existing products.
Qualifications
- Strong C++ programming and practical llama.cpp or ggml inference-engine experience.
- Experience with CUDA, Vulkan, Metal or OpenCL GPU frameworks.
- Understanding of deep-learning models including transformers, LLMs and diffusion systems.
- Computer science, AI or machine-learning degree and evidence of AI research and development.
- Model fine-tuning, production inference, distributed systems or JavaScript are listed as bonus experience.
About Tether
Tether develops digital-token products, including USDT, alongside projects in peer-to-peer communications and AI. The employer describes a globally distributed remote team.
How to apply
Review the original posting and apply through the employer’s careers page.
View source and applyApply directlySource checked: 2026-09-29
Show what you can do
Build a free profile around your CV and real work. Share its link when you are ready. Applications for this role still go directly to the employer.
Build a free profile