← All jobs

AI Inference Engineer, QVAC at Tether

Confirmed open

Optimise C++ inference engines including llama.cpp and ggml for fast private AI on edge devices from Hanoi.

Tether logoTether
Location
Hanoi, Vietnam
Employment
fulltime_permanent
Work arrangement
remote
Technologies
JavaScript, C++

Role overview

The QVAC inference engineer works on the C++ runtime that runs local AI models without relying on cloud infrastructure. The role improves model loading, memory use and execution across edge hardware. The posting lists Hanoi, Vietnam.

Responsibilities

  • Deploy machine-learning models to edge devices with llama.cpp, ggml and relevant GPU frameworks.
  • Improve inference speed, memory use and runtime stability across hardware platforms.
  • Work with researchers to turn models into production implementations.
  • Integrate on-device AI features into existing products.

Qualifications

  • Strong C++ programming and practical llama.cpp or ggml inference-engine experience.
  • Experience with CUDA, Vulkan, Metal or OpenCL GPU frameworks.
  • Understanding of deep-learning models including transformers, LLMs and diffusion systems.
  • Computer science, AI or machine-learning degree and evidence of AI research and development.
  • Model fine-tuning, production inference, distributed systems or JavaScript are listed as bonus experience.

About Tether

Tether develops digital-token products, including USDT, alongside projects in peer-to-peer communications and AI. The employer describes a globally distributed remote team.

How to apply

Review the original posting and apply through the employer’s careers page.

View source and applyApply directly

Source checked: 2026-09-29

Show what you can do

Build a free profile around your CV and real work. Share its link when you are ready. Applications for this role still go directly to the employer.

Build a free profile

Related jobs