Skip to content
Everruns Cloud is open in early access. Run agents without operating the platform.

A voice channel lets people talk to an Agent and hear its answers. Voice is a channel like Slack or AG-UI, so the same Agent can be reachable by text and by voice at once, with one set of instructions, tools and memory. A call and a text chat on the same session share one transcript, so a conversation can move between typing and talking.

Status: Voice is an Adoption feature. An organization administrator turns on Voice in Settings → Features. Speech uses an OpenAI provider.

A voice channel runs in delegated mode. A speech-to-speech model (gpt-realtime-2 by default) only listens and speaks; the Agent, with its own model and tools, writes every answer.

  1. The browser connects to the speech provider over WebRTC. Everruns places the call server-side, so the provider key never reaches the browser.
  2. When the caller finishes talking, their words become a user message on the session, marked with metadata.source = "voice".
  3. The Agent runs a normal turn. Its answer is spoken sentence by sentence as it streams, so speech starts before the answer is complete.
  4. If the turn says nothing for a moment, a short filler line (“One moment.”) is spoken so the line does not feel dead.
  5. When the caller talks over the answer, speech stops at once. The transcript records what the caller actually heard.

Tools, approvals, budgets and audit apply to a voice turn exactly as to a text turn. Raw audio is never stored; transcripts are kept like chat messages.

  1. Open the Agent and select the Integrations tab.
  2. Select Add channel and choose Voice.
  3. Pick a voice and, optionally, a greeting, a language hint and a speaking style, then select Save channel.
  4. Select Talk to this channel to test it from the browser.

Over the API, create it like any channel:

Terminal window
curl -X POST "$EVERRUNS_API/v1/agents/$AGENT_ID/channels" \
-H "Authorization: Bearer $EVERRUNS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"channel_type": "voice",
"channel_config": {
"voice": "marin",
"greeting": "Hi, you are talking to an AI assistant. How can I help?"
}
}'
FieldDefaultWhat it does
modedelegatedHow listening, thinking and speaking are split. Only delegated today.
modelgpt-realtime-2Speech model that listens and speaks.
voicemarinProvider voice.
languagenoneLanguage hint for transcribing the caller, such as en.
greetingnoneSpoken when the call connects. Say that the caller is talking to an AI.
turn_detectionserver_vadserver_vad ends a turn on a pause; semantic_vad waits for a finished thought.
interruptionsteerWhat talking over the answer does to the running turn: steer keeps it running and the next utterance steers it, cancel stops it.
fillerOne moment.Line spoken while the Agent works.
filler_after_ms1500Silence before the filler is spoken. 0 turns fillers off.
speaking_stylenoneDelivery instructions for the speech model, such as pace or tone. Business rules belong in the Agent.

The speech provider is the organization’s default for realtime speech. Pass provider_id when placing a call to use another configured OpenAI provider.

Your app creates a WebRTC offer with the microphone track, sends it to Everruns, and applies the SDP answer it gets back.

MethodRouteUse
POST/v1/agents/{agent_id}/channels/{channel_id}/voice/callsCall a voice channel. Starts a new session unless session_id is given.
POST/v1/sessions/{session_id}/voice/callsTalk on an existing session. Uses the Agent’s voice channel, or channel_id if given.
POST/v1/sessions/{session_id}/voice/{voice_connection_id}/endHang up.
const pc = new RTCPeerConnection();
const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
pc.addTrack(mic.getTracks()[0]);
const audio = new Audio();
audio.autoplay = true;
pc.ontrack = (event) => (audio.srcObject = event.streams[0]);
pc.createDataChannel("oai-events");
const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
const response = await fetch(
`${api}/v1/agents/${agentId}/channels/${channelId}/voice/calls`,
{
method: "POST",
headers: { Authorization: `Bearer ${token}`, "Content-Type": "application/json" },
body: JSON.stringify({ sdp: offer.sdp }),
},
);
const { session, voice } = await response.json();
await pc.setRemoteDescription({ type: "answer", sdp: voice.answer_sdp });

The response carries the session, so your app can show the transcript by streaming its events. A call lasts at most one hour.

A call emits voice.session.started, voice.session.ended and voice.session.failed, transcript events for the caller (voice.input_transcript.*) and for what was spoken (voice.output_transcript.*), and voice.output.interrupted when the caller talks over an answer. See the event reference.