Thevideo,voiceandconversationlayerforAI.

Video, voice, and real-time conversation for any being you can imagine, humans, mascots, game avatars, and brand personalities. Built in one studio. Shipped anywhere.

Built for production

Proof, not just a demo.

  • 2,400+teams building
  • 180msmedian voice latency
  • 32languages live
  • SOC 2Type II
  • TypeScriptPython · Go
Sunesis Studio

One canvas. Every layer.

Voice, video, and conversation, generate, preview, and ship from a single screen. No tabs. No context switching. Just flow.

app.sunesis.ai/studio

Welcome to Sunesis, I can see, hear, and talk back in real time.

Narrate a storyRecord an adGuide a meditation
Sunesis Voice v2
Generate
Generate a character. Talk to it live.

A digital character that sees, hears, and talks back.

Build a character in seconds, choose its face, its look, and its voice, then start a real video call right in your browser. It listens with live speech recognition, replies out loud with natural text-to-speech, and holds the whole conversation in a running chat thread. No clips, no pre-renders. Press the mic and say hello.

  • Hears you in real time, replies out loud, no server round-trip
  • Full chat thread and live captions as they speak
  • Turn on your camera for a face-to-face call, right in the browser
1 · Face
2 · Look
3 · VoiceAria's voice · Warm · Female
Photorealistic · Voice · VideoReady
Aria
Aria · Photorealistic
Call Aria, live, right in your browser
How it connects

One character. A complete interaction layer.

Build your first character
01

Give it an identity

Choose a face, define personality, and assign a voice that sounds unmistakably yours.

02

Let people talk to it

Open a low-latency voice or video call that listens, responds, and takes turns naturally.

03

Give it work to do

Connect knowledge, memory, and tools so every conversation can lead to an outcome.

04

Put it where it matters

Ship on the web, in your product, or over the phone from one real-time API.

Voice, video, conversation, and action work from the same character definition.SUNESIS / REAL-TIME STACK
Sunesis Voice

Make every interaction feel considered.

Explore voice, live interaction, and developer tools when this section comes into view.

Flexible ways to build

Start with the moment you want to improve.

Choose the path that fits your team today. Add characters, agents, and developer tooling later without changing the foundation underneath.

{ }
S
</>
01For product teams

Turn one idea into a live interaction.

Move from a brief to a character, voice, or guided conversation without assembling five separate tools.

  • Prototype a character in Studio
  • Preview voice and video together
  • Share a working flow before you ship
1 workspace
to shape the experience
184ms
streaming round-trip
How can I change the delivery address?
I found the order and kept the context for your team.
Memory attached · handoff ready
02For customer-facing teams

Give every conversation a useful next step.

Ground an agent in your knowledge, give it a memory, and let it hand off with the context your team needs.

  • Voice, chat, and video in one agent
  • Scoped tools and knowledge access
  • Human handoff with a clean summary
32 languages
for global conversations
24/7
without losing the human tone
TSPYGOWS
03For developers

One foundation for every modality.

Use the same credentials, events, usage controls, and SDK conventions across speech, characters, live rooms, and agents.

  • TypeScript, Python, Go, REST, and WebSocket
  • Stream speech and video as the user waits
  • Add webhooks, memory, and tools when ready
1 API key
across every surface
Sub-200ms
for real-time turns

One workspace · shared credentials · usage you can see.

Developers

Voice, video, and conversation APIs built for production.

Fast to start. Open by design.

Build speech, cloning, avatars, and live transcription into one API workflow. TypeScript, Python, and Go SDKs are ready out of the box, with sub-200ms response times.

184ms
first audio
32
languages
99.99%
uptime
Quickstart

Your first response in five minutes.

import { SunesisClient } from "@sunesis/sunesis-js";
const sunesis = new SunesisClient();

// Stream lifelike speech, token by token
const audio = await sunesis.textToSpeech.convert({
  text: "Welcome to Sunesis.",
  voiceId: "aria",
  stream: true,
});
Use cases

What teams ship on Sunesis.

Four production paths. One API key, one event model, and one usage view underneath them all.

Use case 01

Avatar video

Lip-syncable, emotion-aware TTS for AI avatar products. Your digital character speaks with natural prosody and visible emotion, video + voice fused.

SceneVoiceStatus
Welcome + product tourAria · WarmRendered
Feature walkthroughAria · WarmRendered
Pricing explainerAtlas · DeepQueued
Closing + next stepsAria · WarmDraft
Localised cut, FrenchÉtienne · ClearDraft
Use case 02

Voice agent

Sub-200ms turn-taking over WebSocket. Streaming TTS and ASR so your agent can listen, think, and respond without the awkward pause.

TurnLatencyOutcome
Caller states the problem168msUnderstood
Agent checks the order184msTool call
Caller interrupts142msBarge-in
Handed to a person176msEscalated
Summary written back159msLogged
Use case 03

Audio content & companions

Notes-to-audio, narration, and AI companions that actually talk back. Per-character pricing that scales with usage, not seats.

TrackVoiceLength
Chapter one, narrationNova · Bright12:04
Weekly notes digestSage · Calm04:38
Companion, evening check-inLuna · Soft02:11
Ad read, 30 second cutAtlas · Deep00:30
Meditation, sleep setLuna · Soft18:52
Use case 04

Conversational apps

Voice cloning from 15 seconds of audio. Persistent memory, function calling, and RAG, build agents that remember and act.

SessionToolsMemory
Onboarding assistant3 scopedPersistent
Support triage5 scopedPersistent
In-app tutor2 scopedSession
Account concierge4 scopedPersistent
Renewal outreach3 scopedPersistent
Character, voice & live conversation

Make every interaction feel like it has a point of view.

Create the character and voice together, then deliver a response quickly enough to keep the moment moving.

You define

Start with someone worth talking to.

Role & point of view
What it is there to do, and the opinion it holds while doing it.
Voice direction
Pace, warmth, and how it should sound when the answer is no.
Context & guardrails
What it may act on, and the lines it will not cross.

Aria

Warm · Female · English

Role
Product guide
Voice
Warm, measured
First response
<200ms
Languages
32
Voice sample
15s

Everything above travels together. Change the voice and the character keeps its role; change the role and it keeps its voice.

4.6 MOS voice qualityMeasured naturalness, not a promise on a slide.
One foundation. Many ways to build.

Compose the interaction your product deserves.

Start with speech, a character, a live session, or an agent, and connect them when the experience calls for more. One workspace, one set of credentials, one usage view.

184ms
to first audio, streaming
32
languages live in production
4.6
MOS measured naturalness
99.99%
uptime across 14 regions

Typed SDKs

TypeScript, Python, Go, REST, and WebSocket examples that match production behavior.

Built to operate

Scoped API keys, audit-friendly workspace controls, and predictable error handling.

Human developer support

Talk to the engineers building the platform when you are designing at scale.

Talk to us
Real people

Made for the humans on both sides.

Creators

Clone your voice from one short take, then narrate anything — and keep that same voice across every language you publish in.

Live conversations

Put a face on the conversation. A visible character that listens, answers, and holds the thread in real time.

Support & sales

One agent across phone, chat, and video — listening properly, and handing over with full context when a person should take it.

Sub-200ms
latency across video and voice
32
languages one voice can carry
4.6 MOS
measured, not promised on a slide
Digital Characters

Real faces. Real conversations.

Click any face to hear them speak.

240+ voices, infinite possibilities, voices uploaded and built by the Sunesis community.

The difference

Not another laggy AI platform.

Sunesis Labs compared with other voice and video platforms
CapabilityThe othersSunesis Labs
Latency500ms+ latency180ms latency
Voice qualityRobotic voicesMOS 4.6 naturalness
MemoryNo memoryPersistent memory
Pricing / lock-inVendor lock-inOpen API, no lock-in
Regions / uptimeSingle-region14 regions · 99.99% uptime

Trusted by teams shipping voice and video in production

  • Northwind
  • Koru Health
  • Helio
  • Northwind
  • Koru Health
  • Helio
  • Northwind
  • Koru Health
  • Helio
  • Northwind
  • Koru Health
  • Helio
Customer outcomes

Powering voice and video for 2,400+ teams.

From new products finding their voice to public companies redesigning a customer journey, teams use Sunesis to make the interaction count.

Voice agents
We replaced our entire IVR with Sunesis agents. Call deflection jumped overnight and our support costs dropped by 64%.

Mara Lindqvist

VP Engineering, Northwind

64% cost reduction
Digital characters
Onboarding completion went from 41% to 89% once we added a face-to-face digital character to the flow.

Dev Patel

Head of Product, Koru Health

41% → 89% onboarding
Real-time API
We ran 40k concurrent voice sessions during launch week without a single dropped call. The streaming latency is unreal.

Sofia Marchetti

Staff Engineer, Helio

40k concurrent sessions
In the wild

The product is part of the conversation.

Creators and developers share the moments that make a voice experience feel different.

Cloned my voice in 15 seconds and it nailed my cadence on the first try. Wild.

@creatorkai@YouTube

The streaming TTS feels instant, my viewers think I'm live when I'm not.

@novastreams@TikTok

Built a customer support agent in an afternoon. It's better than our offshore team.

@indiehacker@Twitter

Finally a voice API that doesn't sound like a robot from 2015.

@devjules@YouTube
Plans that grow with you

Start building now. Choose a plan when it matters.

Explore credits, production features, and team options on the pricing page, without interrupting the product story.

Explore pricing
The real-time backbone

The details that make an interaction feel instant.

Every layer is measured for the moment between a person speaking and the system responding.

2.4k
teams building on Sunesis
180ms
median latency, voice generation
32
languages for live conversations
4.6
MOS, human-level naturalness
One turn, end to end14 regions online
  1. 01

    Speech in

    Live transcription as they talk

  2. 02

    Understand

    Grounded reasoning and tools

  3. 03

    Voice out

    Streaming speech, first token

  4. 04

    On screen

    Character video, lip synced

Voice in to response out

180msmedian, measured in production

Streaming voice · character video · grounded agents

FAQ

Questions, answered.

Sunesis Voice v2 speaks 32 languages out of the box with any of our 240+ voices. The API supports 30+ languages including English, Spanish, Japanese, French, German, Mandarin, and more, all with native-quality prosody and the ability to clone a voice and use it across languages.

AI Voice, Video & Conversation Agents | Sunesis Labs