Skip to content
K3YHOLD

By K3YHOLD · Pronounced “human mini”

HueMann Mini

A voice model built to think on its feet.

HueMann Mini is the realtime voice model we’re building for maintenance coordination: speech in, speech out, built to listen while it talks, think a request through and call the tools a repair needs — without leaving the caller in silence. It’s in development, and when testing is done we plan to open-source it.

In development

01 · Built to

Hear you out. Think it through. Get it booked.

  • Listen while it talks

    Full-duplex by design: built for the pauses, the “mm-hm”s and the moment a caller cuts in — the way people actually talk on the phone.

  • Think without dead air

    Built to reason through a request mid-call while the conversation keeps moving, instead of leaving the caller in silence.

  • Call the tools a repair needs

    Built to look things up, check a schedule and book the visit, inside rules you set. In our harness, anything that commits money waits for a person.

  • More than one language

    We're finishing multilingual support. Tenants, owners and vendors don't all speak the same language, so the model shouldn't have to.

02 · One job

Built for maintenance coordination.

Most voice models are built to talk about anything. HueMann Mini is being built for the calls a repair actually needs — a tenant reporting a leak, a vendor booking the visit, an owner approving the cost — and for the harness that runs them.

We’re tuning the model and our harness together: the runtime that gives it its tools, its memory and its rules. That’s where it will run best.

03 · Where it will run

Solar-powered when we serve it. Yours to run when we open-source it.

  1. 01

    Served by K3YHOLD

    We plan to serve HueMann Mini ourselves: small-scale realtime inference on servers powered by solar, with battery backup — for maintenance coordination, and nothing else.

    Coming
  2. 02

    On your own GPU

    Open-source when testing is done. We're aiming to fit it on a single GPU with 12–16 GB of memory. It's optimized for our harness, so that's where it will run best.

    Coming
  3. 03

    On the device

    We're also working on a smaller on-device model, for conversations that should never leave the phone they're held on.

    Coming

The measuring has started before the model has: our voice cost studies price real maintenance calls on today’s realtime voices, and the last study in the series measures HueMann Mini on our own servers, counted all the way down to the panels. Read the first study.

04 · Measured in the open

Judged on how it holds a conversation.

We test HueMann Mini’s turn-taking with Full-Duplex-Bench, a public benchmark from university researchers. We’ll publish our results when we release it — the numbers, and how we got them.

We’ve built and measured more than one design against the same benchmark, and we keep what holds a real conversation best. Every round of training aims smaller and more efficient. That’s the Mini.

  • Pauses

    Does it wait while the caller thinks, instead of jumping in?

  • Backchannels

    Does it say “mm-hm” at the right moments without taking the floor?

  • Taking the turn

    Does it answer when the caller is done — not before, and not after a long silence?

  • Being interrupted

    Does it stop and listen when the caller cuts in?

05 · The research

The study that started our current voice work.

arXiv:2503.04721 · Version 3,

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, Hung-yi Lee. First published .
Read the paper on arXiv (opens in a new tab)

K3YHOLD did not sponsor, conduct or take part in this study. We use its published benchmark to test our own work.

06 · Quick start

What the quick start will cover.

It ships with the open-source release. Nothing here can be downloaded yet.

  1. 01What you needA single GPU with 12–16 GB of memory — the size we're aiming for.At release
  2. 02Get the modelThe open weights, and the licence they ship under.At release
  3. 03Make a test callTalk to it. Pause. Interrupt it. Ask it to book something.At release
  4. 04Measure it yourselfRun Full-Duplex-Bench and check our published numbers against your own.At release

The rules it’s built under

  1. 01Every call in our harness starts by saying it's an AI.
  2. 02Anything that commits money waits for a person — the same rule as everything else we build.
  3. 03We publish how we tested it before we ask anyone to trust it.
  4. 04No release date until we're confident in one.

Questions, answered plainly.

What is HueMann Mini?

A realtime voice model K3YHOLD is building for maintenance coordination: speech in, speech out, built to listen while it talks, think a request through and call the tools a repair needs. It's in development.

How do you say it?

Like “human mini”.

Is HueMann Mini open source?

Not yet. We plan to open-source the model when testing is done. Only the model: the harness it runs in stays ours. It's optimized for that harness, so that's where it will run best.

When can I use it?

There's no release date. We'll announce one when we're confident in it.

Will MaintenanceTech run on HueMann Mini?

MaintenanceTech was designed to be model-agnostic and capable of using open-source models, so the model underneath can change without the product changing — and HueMann Mini is being optimized to run on it.

Where will it run?

Three places, as we plan it: served by K3YHOLD from small-scale servers powered by solar, with battery backup; on your own GPU once it's open-sourced; and, later, a smaller model on the device itself.

How will you show it works?

We test its turn-taking with Full-Duplex-Bench, a public research benchmark, and we'll publish our results — and how we got them — with the release.

Want to hear it first?

Join the HueMann Mini list for updates and early testing. Joining commits you to nothing.