Agent Voice Mode turns the AI agent you already run into a phone call. You talk. It answers, shows its work, and gets on with it.
Works with the agent you already have — not a chatbot of ours.
It has no knowledge of its own. Every real question goes to your agent — on your machine, with your files, your context and your tools. The app is the voice.
Then it's just a phone call.
Running somewhere reachable? Paste its URL and key. Running on your own Mac or home box behind NAT? Install a small connector that dials out — no port forwarding, no public IP.
One scan signs the phone in with your account and your agent already attached. Nothing to fill in, no password to invent.
Every call is linked to a project, so the agent starts with the context instead of asking for it.
Voice alone loses anything with structure. When the answer is a table, a list or a picture, it appears on screen while the voice explains it — so you get the shape and the summary at once.
The things that make a voice tool usable past the first five minutes.
Cut in mid-sentence the way you would with a person. It stops, listens, and picks up from what you actually said.
Take a picture mid-call and ask about it. Say what you want done with it and the caption writes itself.
"Switch to best quality" or "use the economy model" — said out loud, mid-conversation. The countdown re-reads itself at the new rate.
Switch projects and the conversation follows, history and all. The agent knows which one you mean.
Native iOS throughout. It follows your phone, or you pin it — and the transcript stays readable either way.
Thirty seconds muted and quiet closes the session. A closed session bills nothing, so forgetting to hang up costs you nothing.
No subscription. Talk time is metered per minute against a prepaid balance, and you choose the rate.
Any agent you can reach over HTTP, and any agent running on a machine you control. Claude Code, OpenClaw and Hermes are the ones we run against; the protocol is deliberately agent-agnostic, so a new one is an adapter rather than a rewrite.
The audio doesn't. Speech flows directly between your phone and the speech model; we mint the short-lived token that lets that happen and meter the minutes. When a question is substantive it's relayed as text to your agent, which is the only way your agent can answer it.
No. You scan a code and you're signed in, with your agent already attached. There's no password, and the credential lives in the iPhone's Keychain.
That's the common case and it's the one the connector exists for. It dials out from your machine and holds the connection open, so nothing has to be reachable from the internet — no port forwarding, no static IP, no tunnel service in the middle.
Not yet — it's in TestFlight while we finish the release build. Ask for an invite below and you'll get one.
Not today. iOS first.
TestFlight invites are going out now. Tell us what you run your agent on and we'll send one.