Ambient is DuoDuo in a room. It listens around the clock, with no app to open, no button to hold and no wake word to recite before every sentence. You talk, it is already listening, and it judges whether you were talking to it.
A smart speaker asks did this person stop talking? Ambient asks a harder question, was that said to me?, and answers it the way someone sitting in the corner would.

Being talked about is not being talked to#
Two colleagues at a whiteboard spend four minutes arguing about whether DuoDuo handled yesterday's rollout well. Its name comes up a dozen times. It says nothing: no chime, no "did you mean me?", not half a second of a listening animation.
Then one of them turns to the corner and asks what it thinks, and it answers about the argument itself: both positions, who said which, the rollout they were describing. It had been listening the whole time and had no standing to speak until someone gave it one.
The judgment also runs the other way. Command-shaped speech is not automatically an order: people dictate to their phones in the phrasing they would use on DuoDuo, and being overheard is not being instructed. A retold conversation, thinking out loud, a rhetorical question and an instruction to a third person are not addressed to it either; someone in the room would answer none of them. Hearing its name counts as evidence, not as being called.
It records either way#
Everything anyone says becomes part of the room's record, addressed to DuoDuo or not. That record is how it reads the room later. The addressee judgment decides only whether it acts; recording is not a behavioural choice.
Transcript and speaker attribution arrive together, so a five-person conversation is already sorted by who said what, and the same person keeps one identity through the day without anyone enrolling a voice sample.
Voice presence is continuous rather than push-to-talk, so you can interrupt it mid-word. What it says next is grounded in how much of its answer reached you, not in the paragraph it meant to say.
A room is a directory#
Each room is an instance directory alongside the other channels, and what it knows about itself is files:
| File | Written by | What it holds |
|---|---|---|
notes.md | you and DuoDuo | What has been worked out about the room: whose voice is whose, the phrases it keeps mishearing |
transcript-*.jsonl | the runtime | What was said, as it was heard |
imlog-*.jsonl | the runtime | The cleaned record of the room |
events-*.jsonl | the runtime | What happened in the room that nobody said |
An unfamiliar voice gets a temporary badge for that room, never a name of its own accord. You name it by writing one line in notes.md, the way you would introduce a person, or by telling DuoDuo who it was. Deciding unprompted that two voices are the same person is a confident, undetectable error that can collapse a whole room into one identity, so the code is not allowed to make it.
There is no alias table and nothing to configure. notes.md is prose, and a misidentification, or a word the room keeps getting wrong, is corrected by editing a line of it. The file is read fresh for each exchange, so a correction applies to the next thing said, with no restart and no reprocessing.
Tuning a room#
Ambient follows the same layering as the other channels: a kind descriptor for defaults across every room, an instance descriptor for one room, and the host environment for addresses and credentials. The name a room answers to, how long it waits before deciding a sentence is finished, and how much of the room's recent talk it carries into an answer all live in the descriptors, so one loud room can be tuned without touching the others. To change one, tell DuoDuo what the room should do differently.
A channel, and the cerebellum it hears with#
Ambient is a channel, like Feishu/Lark or Tether, but a channel cannot hear. The hearing happens in a backend service, the cerebellum, and the two live apart:
- The Ambient channel runs beside your DuoDuo. It serves the room's capture page, holds the microphone, hands each sentence addressed to DuoDuo to a session, and brings the answer back as text and speech.
- The cerebellum runs beside the models it calls. It does the hearing: voice presence, transcription, who said what, voiceprints, and the judge that decides whether a sentence was said to DuoDuo.
Something in the room holds the microphone: a tablet or any browser with the capture page open, an edge device, or your iPhone. The DuoDuo app is Ambient's app; see DuoDuo Pocket.
The channel serves its own room page. Open it on a tablet in the room and it is the ears and the face: DuoDuo listening, what it heard, the step it is on, the answer it spoke, and the room's record, including what it heard and judged was not said to it. Its own text is Chinese today; these drawings show its states in English.
The cerebellum is a reference, not a requirement#
openduo/ambient is source-available: the channel, the wire protocol and the cerebellum. The cerebellum itself is one small process. Everything that needs a model sits behind five legs, each a contract that can run on your own machine or in the cloud:
- Ears transcribe each stretch of speech together with who said it. The reference runs MOSS-Transcribe-Diarize; anything that returns the same speaker-labelled rows fits.
- The diarizer follows speaker tracks through the continuous stream.
- Voiceprints keep each voice on one anonymous number through the day.
- The judge decides whether a sentence was said to DuoDuo. Any OpenAI-compatible endpoint can serve it: a small local model, a 27B model on your own cards, or a hosted API.
- The mouth speaks the replies. The reference uses a cloud realtime voice because no local one has met the bar yet; the contract does not care where it runs.
The repository ships reference deployments for the model legs (how one GPU machine ran them, with the reason for each setting), and each doubles as a conformance test for a replacement: the same model on another engine, a different model, or a hosted API.
The cerebellum can be all local, all cloud, or split. The balanced split keeps the legs that hear everything (the ears, the diarizer and the voiceprints) on machines you run, and lets the judge and the mouth, which only handle words, go wherever quality and cost are best. Which split fits depends on your machine and your budget; wiring it is your agent's job.
With the listening legs and the judge on your own card, the reference fits two shapes, and the judge decides which:
| Profile | Ears, voiceprints, judge | Free VRAM |
|---|---|---|
| Constrained | Lightweight ears, voiceprints on the CPU, a small local judge | about 4.5 GB on one card |
| Ample | Faster ears, voiceprints on the GPU, a 27B judge across two cards | 29 GB of judge weights, plus a pool you size |
For a single RTX 4090- or 3090-class card, openduo/ambient-engine builds an optional judge server for a ternary 27B model. The judge's behaviour was measured on the 27B reference model; a small local judge or a hosted endpoint works, but its quality is unmeasured.
Where the room's audio and words go#
A room full of people talking all day is about the most sensitive recording there is, which is why the balanced split keeps the listening legs at home. Whatever stands behind the ears, the diarizer and the voiceprints receives that audio; run them yourself and the room's audio stays on your machines. The judge receives the room's words: host it yourself and they stay too, or point it at a hosted API and they go there. The mouth receives the text of what DuoDuo says, never the room's audio.
Have your agents bring it up#
Nothing on this page is meant to be done by hand. There are two jobs, and each has an agent-ready path.
The cerebellum. docs/deploy.md in openduo/ambient is a runbook written to be followed top to bottom. It reads the machine first (free memory per card, driver, ports), chooses a profile, lists what only you can supply, and ends every step with the command that proves it. It walks the simplest shape, DuoDuo and the cerebellum on one machine with the card. The channel and the cerebellum talk over one authenticated connection, so the cerebellum can also live on another machine, and the agent adapts the runbook to that. Give the agent on that machine, your DuoDuo or any coding agent, one sentence:
Bring up an Ambient cerebellum on this machine by following docs/deploy.md in https://github.com/openduo/ambient.
If someone already runs a cerebellum for you, skip this: they give you its address and token, or a Tailscale share link.
The channel and your rooms. Give your DuoDuo the duoduo-ambient skill, then ask it to set up Ambient:
npx -y skills add https://github.com/openduo/duoduo --skill duoduo-ambient
DuoDuo installs the channel, connects it to the cerebellum, creates a room, and publishes the room to your tailnet. Later it restarts and upgrades the channel, adds rooms, and works through errors. Your part is what only you can do: type the cerebellum's token into the host terminal yourself, never into a chat; approve each change; and do the taps on your phone.
On your tailnet#
Ambient listens on loopback. For devices on your tailnet, tailscale serve terminates HTTPS at the host's ts.net name and forwards to the channel; the DuoDuo app and a room's tablet reach it there. Do not use Funnel for Ambient: it would publish the room page to the internet.
- 01Hold OK on the Passport. Your voice streams to the phone over Bluetooth LE.
- 02The phone dials your host over your tailnet: HTTPS to its ts.net address, then Ambient and DuoDuo.
- 03The answer comes back the same way, to the phone and the Passport screen.
- 04A room device opens the Ambient page at the same address, and listens.
- 05Grok, Dots or Muse knock on Tether. Only your passkey opens the door.