Fish Audio raises $52 million for voice AI models and services

Fish Audio just turned a bedroom GPU experiment into a $52 million seed. The Palo Alto voice startup is keeping older speech models open while parking its newest quality behind a paid API โ and telling enterprises that expressive delivery, not just correct words, is the product.
Who put $52M into Fish Audio
Fish Audioโs own funding post says Coreline Ventures and Capital Today led the round on the companyโs first anniversary, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphaist Partners. TechCrunchโs write-up matches the lead and most of the syndicate, and adds Bayhouse Ventures to the list.
The company claims more than 8 million people using open-source or hosted versions of its models, and $21 million in annual recurring revenue.
The origin story is concrete. Former Nvidia researcher Shijia Liao trained a voice model on a single GPU after getting tired of lifeless synthetic speech; the Fish Speech GitHub repo later passed 31,000 stars. CEO and co-founder Rissa Cao says the open-source-plus-creator-plans phase ran efficiently enough that the company did not need outside capital โ until it wanted harder models and enterprise distribution while investor interest was heating up.
What is open vs locked behind S2.1 Pro
Fish Audio says it shipped five models in a year: four speech-generation models and one speech-to-text model. Three speech-generation models are open-sourced. The latest, S2.1 Pro, sits behind the paid API.
That split is the business. Creators and indie builders get a public stack and a community library Fish Audio puts at 2 million-plus voice models. Teams that need uptime, latency, on-prem, zero-data-retention, or HIPAA-style configurations pay for the frontier voice. Fishโs blog says enterprise plus developers now make up two-thirds of revenue, and that S2.1 Pro was preferred by 66% of listeners over leading competitors in blind tests, with 83-plus languages and more than 15,000 natural-language controls.
Monthly paid plans unlock generation minutes and voice cloning. There is also an enterprise API and platform tier. The company says it plans an audio understanding model this year and is building speech-to-speech.
Why voice startups still raise at seed scale
The category is crowded โ ElevenLabs, WellSaid, Cartesia, Speechify, Async, Krisp, and others are already fighting for the same creator and contact-center budgets.
Fishโs bet is fine-grained control and cheap enough inference to give the model away strategically. The company says it rebuilt its inference stack with custom FP8 kernels and can run S2.1 Pro at 8,000-plus tokens per second on a single H200 โ cheap enough that S2.1 Pro was free for developers via API through the end of August after the raise.
Consent is the soft underbelly. TechCrunch reports creators previously alleged voices were uploaded without consent, and that takedowns were slow even with a DMCA process. Cao told TechCrunch takedowns are now automated โ a short voice sample or contract can pull a voice in under three minutes. Corelineโs Osuke Honda put the investor standard on the record: community models only work if consent, transparency, and attribution are product features, not afterthoughts.
Customers to watch: HeyGen and Sanas
Fish Audio says organizations including HeyGen and Sanas already use the enterprise stack. Cao told TechCrunch HeyGen wants realism for AI avatars; gaming studios want expressive character voices; voice-agent shops like LiveKit want natural, low-latency speech that still sounds alive on a call. Fishโs blog also names Retell as a partner for putting expressive voice into more agent stacks.
Those names matter more than the seed headline. Avatar video, gaming, and voice agents are three different quality bars. If Fish can hold all three without fracturing the model line, the open/paid split becomes a real moat.



