shower thought:
the first industrial revolution caused a transition away from personalized stuff (eg. clothes, furniture, food) and towards mass-produced.
this AI one is enabling a return to personalization (eg. software, education, entertainment)
Very, very dangerous AI usage by me today: giving it access to both a wide array of local burger chains as well as my credit card.
Last week, @DoorDash gave me early access to their CLI. Wrangled it into HTTP endpoints for my @FishAudio voice agent. 🧑🍳
In just 1 year, Fish Audio reached $21M ARR, 8M users, and raised a $52M seed.
When we started out every investor and friend told me it was too late to compete in voice AI.
Turns out, it wasn't. Here's the story of our first year.
.@FishAudio is evolving… 🐟🐠🦈🐋
Today we’re celebrating our successful $52M seed round, and tomorrow we’ll continue building the future of audio understanding, multimedia creation, and (keyboard-free) human-computer interaction.
Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro.
>It can clone a voice from 5 seconds of audio
>2x faster than Cartesia & 1/6th the cost of Eleven Labs
>most expressive model with word level control over emotion, intonation, pacing etc
We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production.
If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free.
Book a demo: s.fish.audio/tmapke
To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.
Our best voice model is now free for developers.
S2.1 Pro- the same TTS on our paid tier, free API, 83 languages, no hard usage cap, same endpoint you already call.
Already integrated? Set model: "s2.1-pro-free" and you're on S2.1 Pro.
With Inworld's TTS-2 launch yesterday, we now support a ton more languages. I was able to update my language learning app to support Swedish, which I'm starting to learn 🇸🇪
App is free for now if you want to try it out.
.@inworld launched our new TTS model today! We've made some big improvements in latency, emotionality, and stability, and also added phoneme alignment to support lip sync.
Here's a cheeky all-CSS lip sync demo. Does it look like me?
github.com/cshape/lipsync
. @inworld Runtime and TTS are optimized for realtime, but people have been asking about how to create stuff like audiobooks or long article narration.
I created some scripts in Python and JS to help: github.com/inworld-ai/inworl…
Note: Inworld TTS is still FREE til Dec 31!
This example / demo with @inworld is probably the most fun I've ever had building one.
Tabletop-like w/ skill checks, combat, gearing, barter system, XP, all AI generated in realtime. Kept wanting to add more—scope def creeped a lot. Code + key decisions below ...
We're making Inworld TTS free until the end of the year (!)
We were feeling in the holiday spirit today, and after seeing the community rate our TTS at #1 on leaderboards and help us grow 100% week on week, we wanted to gift something and give every builder the chance to try what's topping the benchmarks.
Merry Christmas, Happy holidays. Inworld TTS is free this month.
Look out next week: we're going to redefine what #1 means.
In 2023, I used @rpgjs_dev for an @inworld hackathon project where we made a survivor game with AI players.
Looking forward to trying out the new version!
After leaning on LLMs pretty hard to code efficiently at work, it's refreshing to work through Advent of Code in a vibeless fashion. Bless you, Eric Wastl.
We just made state-of-the-art TTS 20x more affordable.
$5 per million characters.
And we're open sourcing the training and modeling code (built on Llama).
Because scaling voice AI shouldn't break your budget.
Technical Details → Inworld.ai/blog/introducing-…
Why and how we did it 🧵
If you are contemplating joining a coding bootcamp in 2020 then let me give you a list of the top scams they use to steal your money (yes, even with an ISA, the #1 scam):