A model call ends. The project, the task it was in the middle of, the people it was talking with, and what it worked out should not end with it. DuoDuo exists to give an agent that continuity—and it does so without being another harness. Claude, Codex, Grok, and Pi are the harnesses, and they get smarter with every release. DuoDuo is the environment they live in: an operating system for agent work. It gives an agent a durable body of files, a clock that keeps ticking, and doors to the outside world—then stays out of the way.
An agent is a model inside a harness#
The model reasons. The harness is everything around it that turns reasoning into work: the loop that decides when to call a tool, what the model can see, how a turn ends. Neither half is an agent alone.
That is why "which model" is the wrong unit. Claude Code, Codex, Grok, and Pi are four harnesses, and putting the same model in a different one gives you a different agent. DuoDuo runs all four, and adds the layer neither half has: somewhere for that agent to live between turns.
Why four runtimes instead of one#
DuoDuo runs work through Claude, Codex, Grok, and Pi as peers. Not four models behind one loop—four harnesses, each the one its makers ship.
Most of what an agent can do travels between them, by design. MCP is a protocol, so DuoDuo mounts one tool surface into all four. Skills are files with frontmatter. A model is an API call. DuoDuo widens that portable layer further: whichever runtime a session runs on, it sees the same DuoDuo tools, the same prompt layering, and the same interruption behavior.
What does not travel is a short list, and a decisive one:
| Stuck to its runtime | Why it cannot move |
|---|---|
| First-party tools built into the vendor's own agent | Grok Build searches X natively. There is no protocol for that to travel over. |
| The loop itself | How it steers mid-turn, compacts, spawns subagents, and calls tools |
| Extensions running inside the harness process | A Pi extension hooks Pi's internals; it is not a file you can copy elsewhere |
| Your subscription | A plan is tied to how you authenticate, not to which model answers |
Choosing one runtime for everyone would mean choosing which of those to give up.

Why not sign every model into one harness#
Because that harness becomes the environment, and DuoDuo becomes a wrapper around it. An environment cannot live inside another environment; that is the same layer twice.
It does not get you the capability either. Signing in to Grok from inside another harness gets you the Grok model—not Grok Build's native X search, which lives inside Grok's own agent.
The proof that the four are peers is how DuoDuo fails. Ask for a runtime that is not usable on this machine and you get an error, never a silent fall back to Claude. A fallback chain would mean one of them was the real one.
Your plan follows the client, not the model#
Entitlement attaches to how you authenticate, so the runtime you pick decides what you pay.
Anthropic grants Claude plan usage to third-party apps that authenticate through the Agent SDK. DuoDuo's Claude runtime embeds that SDK at a pinned version, so your Claude plan applies. A harness that reimplements the loop instead of using the SDK falls outside it—Pi documents this about itself: third-party harness usage "draws from extra usage and is billed per token, not against Claude plan limits."
Codex and Grok run through the CLI you already installed, under the login you already have.
Why not build our own harness#
The coding inner loop is solved four times over by teams who do only that. Rebuilding it would mean redoing someone else's layer, and inheriting the job of keeping pace with four vendors at once. Instead their releases arrive with nothing to rebuild: a Codex or Grok release the day you upgrade their CLI, a Claude or Pi release with the DuoDuo release that bumps the pinned embed.
What is not solved is the layer underneath: state that outlives a process, a clock that runs without a message, doors to the outside world, and memory you can read as a diff.
Pi makes the rule literal. Pi is not here as another swappable loop—Claude, Codex, and Grok already cover that. It is here so that the Pi you have already tuned comes along whole: your extensions, your skills, your prompt templates, your AGENTS.md. DuoDuo supplies what a coding harness has no way to have on its own—sessions that survive, channels, a cadence, and jobs.
Adopt the harness whole. Do not absorb it.
Four harnesses, one workload#
Runtime is a per-actor choice, not a host setting. A channel can answer on one, a scheduled job can run on another, and a subconscious partition can use a third—at the same time, on the same machine, against the same files.
So the shape of the work is yours to arrange. A Claude session plans and hands a task to a job that runs on Codex; the job leaves its result on disk; a Grok partition picks it up on the next tick because that is where the room's own knowledge is. None of them call each other. They meet in files and in the event history, which is what lets the arrangement change without anything being rewired.
This is the part that a single harness cannot offer no matter how many models you sign into it. Inside one harness, everything is one process's idea of a conversation. Here they are separate actors with their own ownership and lifecycle, which is exactly what makes composing them safe.
Why this does not fight your harness's own memory#
Claude Code has its CLAUDE.md. Codex and Pi read AGENTS.md. Those are harness-layer memory: instructions you wrote, scoped to that harness and that project, read by that harness for itself. DuoDuo does not touch them, and they keep working exactly as they do today.
What the subconscious produces is a different thing at a different layer. It is not configuration you maintain—it is distilled from the one event history that every surface and every runtime feeds. A correction that actually held up becomes a lesson; a procedure that converged through repeated use becomes a groove; what holds up is folded into a compact intuition layer, assembled by DuoDuo and handed to the session at the point where every runtime is treated the same.
The two compose because they answer different questions. Your CLAUDE.md says how this project wants to be worked on. The intuition layer says what has actually been learned working on it—across every harness, on a clock, without you writing it down.
And it travels because of what it is made of. Fine-tuning locks a lesson into one model's weights. Harness memory locks it into one harness's config. A lesson written as text is bound to neither: it applies to any model in any harness that can read. That is the whole reason DuoDuo distills into prose instead of anything cleverer—it is the only representation that crosses all four.
The filesystem is the body#
Everything DuoDuo does ends up in ordinary files on your machine: conversations, events, jobs, outputs, prompts, memory. The filesystem is the database. The daemon is just a replaceable coordinator around that durable state.
Every incoming message is written to the append-only event history before execution. That history is the record of what happened—what you audit, what duoduo spine reads, what the subconscious learns from, and what the daemon rebuilds its routing from after a restart. The conversation itself continues from each runtime's own session state, kept beside it; the history is how you know what happened to that conversation.
One identity, many actors#
You meet one DuoDuo through the terminal, Feishu, an ACP-enabled editor, or a room you speak in. Inside the daemon, work is separated into actors with explicit ownership:
| Actor | What it owns |
|---|---|
| Foreground session | One continuing conversation attached to one workspace |
| Job session | One-shot or recurring work with its own schedule and delivery target |
| Subconscious partition | One bounded background responsibility, prompt, and schedule |
| Channel | Ingress, presentation, and delivery for one user surface |
Separating these actors keeps a busy channel, a scheduled job, and memory maintenance from collapsing into one giant transcript.
Two cognitive loops#
The foreground cortex responds to live messages, reasons with whichever runtime the session is on, uses tools, and changes the project in front of you.
The background subconscious wakes on cadence without waiting for a message. Each partition owns one job: one reads new event evidence, one tends the intuition layer and dossiers, one watches for patterns that keep repeating, one commits what changed to the kernel's Git history.
The two loops meet through files. Whatever surface the experience came through—terminal, Feishu, editor, a room, a scheduled job—and whichever runtime did the work, it lands in the same event history. The subconscious consumes that one timeline in order, distills selected evidence into intuition and behavioral rules, and every new foreground session starts with that compact intuition layer already loaded. What Claude worked out yesterday sharpens what Codex starts with tomorrow, and the same in every other direction.
Code for what is certain, a model for what needs understanding#
DuoDuo's code does not decide anything on the agent's behalf. Where a question has an exact answer—whether an event was recorded, which day of history has no fragments yet, whether a watched feed changed, how long a tool call has been open—the code turns that certainty into a tool and hands it over. duoduo spine reads the history. A lint measures the memory tree and posts a bounded signal into the inbox of the partition that owns the gap; the partition, not the lint, decides what to do about it. A watcher script calls duoduo session notify when data arrives, so no model turn is spent waiting. duoduo daemon status lists what is running and passes no judgement on it.
The rest needs understanding: what an exchange meant, whether a correction actually held, whether a line of intuition is earning its place, whether the person in the room was talking to it. The agent decides those, with the tools in hand. The boundary is deliberate—code is not asked to fake judgement with thresholds, and the agent is not asked to poll. Agent work stays with the agent; the code's job is to make the certain parts cheap to reach.
A conversation is not a request queue#
People add a detail while the answer is being worked out, cut in, change their minds, and talk about DuoDuo without talking to it. Forcing every message to produce exactly one reply would misread all of that.
So a follow-up sent while a session is working reaches the agent at its next tool boundary rather than waiting for the turn to end, and one reply can take several messages into account. In a room, the judgement goes further: Ambient records everything said and answers only what was said to it—its name is evidence, not a trigger. See Channels.
What it learns is inspectable#
What the subconscious learns is stored as files: lessons, grooves, dossiers, and the intuition layer every new session starts with. The kernel is tracked in Git, so each change to that memory has a diff and a rollback path rather than disappearing inside model weights or a proprietary database.
Read Subconscious & Intuition for the full loop and its operating boundaries.
Who owns what#
| Layer | Responsibility |
|---|---|
| DuoDuo runtime | durability, routing, lifecycle, scheduling, delivery, and concurrency boundaries |
| The harness runtime | reasoning, planning, tool use, and project work |
| Filesystem kernel | the intuition layer and dossiers, prompts, configuration, subconscious partitions, and their history |
| Workspace | the actual project files the agent is reading and changing |
Models do the thinking and keep getting better at it. DuoDuo is where that thinking gets a body, a clock, and a history.