Back to Copilotkit

Voice

showcase/shell-docs/src/content/docs/voice.mdx

1.64.13.9 KB
Original Source
<InlineDemo demo="voice" />

You have a working chat surface and you want users to be able to speak instead of type. By the end of this guide, the chat composer will sprout a mic button, recorded audio will be transcribed by the runtime, and the transcript will auto-send to the agent like any other message.

When to use this

  • Hands-free or accessibility flows where typing isn't the right input modality.
  • Mobile or kiosk surfaces where a long voice query is faster than thumb-typing.
  • Demo and test loops where you want canned audio to drive the chat without a microphone.

If you only need file uploads (audio, images, video, documents), use Multimodal Attachments instead. Voice is specifically about live transcription of recorded speech into chat input.

Frontend

<CopilotChat /> from @copilotkit/react-core/v2 renders the mic button automatically when the runtime advertises audioFileTranscriptionEnabled: true on its /info endpoint. There's nothing to wire up on the chat surface itself:

<Snippet region="voice-page" title="frontend/src/app/page.tsx — chat surface" />

The runtimeUrl="/api/copilotkit-voice" points the browser to your Next.js API route. When the user clicks the mic, the chat captures audio, POSTs it to that runtime route's /transcribe endpoint, drops the resulting transcript into the composer, and submits.

Driving the demo without a mic

For Playwright runs, screenshots, or any flow where prompting for mic permissions is awkward, ship a button that emits a canned sample phrase through an onTranscribed callback, bypassing the transcription endpoint entirely:

<Snippet region="sample-audio-button" title="frontend/src/app/sample-audio-button.tsx" />

The parent chat component can then drop that text into the composer's textarea (matched via data-testid="copilot-chat-textarea") using the native value setter and a synthetic input event so React's managed state updates correctly.

Backend

Next.js API route

Create a dedicated API route at app/api/copilotkit-voice/[[...slug]]/route.ts. The [[...slug]] catch-all pattern lets the V2 runtime handle its internal URL routing (/info, /agent/:id/run, /transcribe, etc.) under the /api/copilotkit-voice base path.

Wire up the V2 runtime with a TranscriptionService. The V1 wrapper drops the transcriptionService option, so use createCopilotRuntimeHandler from @copilotkit/runtime/v2 directly:

<Snippet region="voice-runtime" title="app/api/copilotkit-voice/[[...slug]]/route.ts" />

The basePath: "/api/copilotkit-voice" in createCopilotRuntimeHandler must match the API route's directory path. With transcriptionService set, the runtime advertises audioFileTranscriptionEnabled: true on /info (which is what tells the chat to render the mic button) and routes POST /transcribe to the service.

<WhenFrameworkHas flag="voice_backend_pattern" equals="adk-fastapi-agent-path"> For the Google ADK showcase, agent runs take one more hop: this Next.js route registers the `voice-demo` agent with an `HttpAgent` pointed at `${AGENT_URL}/voice`. The Python `agent_server.py` mounts registered ADK agents with `add_adk_fastapi_endpoint(app, ..., path=f"/{agent_name}")`, so the browser talks to `/api/copilotkit-voice` while the Next.js runtime forwards voice-demo agent runs to the backend `/voice` endpoint. </WhenFrameworkHas>

Custom transcription backends

TranscriptionService from @copilotkit/runtime/v2 is an abstract class. Subclass it to plug in any transcription provider — Whisper, AssemblyAI, Deepgram, your own model. The library ships TranscriptionServiceOpenAI as the canonical reference implementation.

A useful pattern is wrapping your service in a guard that returns a clean 4xx when credentials aren't configured, instead of an opaque 5xx from the underlying SDK:

<Snippet region="transcription-service-guard" title="backend — guarded transcription service" /> <IntegrationGrid path="voice" />