Peer-reviewed work, openly published
Our research appears at ICASSP, CVPR, NeurIPS, and ACL. We publish our methods, not just our benchmarks — because progress in the open moves everyone forward.
Prosody-Aware Diffusion for Emotional Speech Synthesis
L. Chen, M. Okafor, R. Diaz
We introduce a prosody-conditioned diffusion model that generates emotionally expressive speech by modeling pitch and timing contours as explicit latent variables, improving MOS on emotional benchmarks by 0.41.
Neural Talking Heads with Phoneme-Aware Lip Sync
S. Yamamoto, K. Iyer
A phoneme-conditioned rendering pipeline that produces frame-accurate lip synchronization across 32 languages from a single reference image, reducing lip-sync error by 38%.
Memory-Augmented Agents for Long Conversations
A. Bauer, P. Novak, T. Wei
A hierarchical memory architecture that lets conversational agents maintain coherent state across multi-session dialogues, with sub-linear retrieval cost.
On the Calibration of Emotional Classifiers
M. Okafor, R. Diaz
We study the miscalibration of emotion classifiers on out-of-distribution inputs and propose a temperature-scaling method that reduces expected calibration error by 62%.
Voice Cloning from 3 Seconds of Audio
L. Chen, S. Yamamoto
A meta-learning approach to few-shot voice cloning that achieves speaker similarity on par with 30-second baselines using only 3 seconds of enrolment audio.
Where our work lands
Sunesis research has been cited over 2,400 times across industry and academia, and underpins production deployments at four of the top ten tech companies.
Read along as we publish.
Free to start — 10,000 credits every month, no credit card required.
SOC 2 Type II · GDPR · HIPAA-ready · No vendor lock-in