The team claims Sonilo isn't just another AI toy but a sophisticated 'Sound World Model' that understands narrative pacing and emotion, offering a seamless, legally clean solution for creators and studios alike.
Insiders whisper that the real drama here is the talent raid. Pulling the head of TikTok's global music operations and their top multimodal AI lead is a massive flex. Are they building a tool or a Trojan horse to disrupt traditional licensing?
1. B Capital led an $11 million round on Oct 8, 2026, with Redpoint participating. 2. CEO Shawn Song explicitly stated: 'You give us the footage, and Sonilo understands what’s happening frame by frame... and generates music and sound effects that naturally fit the scene.'
TikTok’s brain trust is exiting the app to build the next layer of the creator economy, and they’re bringing $11 million and a very specific vendetta against bad audio sync.
Let’s be clear: when the architects of TikTok’s AI and music divisions decide to pack their bags and start a new venture, you don’t just yawn and scroll past. You pay attention. Sonilo, a San Francisco-based generative audio startup founded by former TikTok heavyweights, has just secured an $11 million funding round led by B Capital, with participation from Redpoint.
This isn't some garage hobby project; this is a coordinated exit from the most influential short-form video platform on the planet to build what they call the "Sound World Model." The drama lies in the pedigree. CEO Shawn Song, who trained at Carnegie Mellon University, previously led multimodal AI technology work at TikTok. He’s joined by CTO Alex Yin, who also worked alongside him at the ByteDance giant and holds a doctorate in computer music from the University of York.
But the real tea? COO Keli Li, the woman who built and ran TikTok’s global music operation, is now steering this ship. With a background in rights deals and partnerships, her presence signals that Sonilo isn’t just trying to generate noise; they’re trying to fix the broken economics of music licensing.
They’ve already snapped up Shutterstock as an early licensing partner, promising that artists and rights owners will receive fees and revenue shares—a direct shot at the current chaotic state of AI training data. Song’s vision is unapologetically ambitious. "We built Sonilo to understand the visual story," Song said in a statement.
"You give us the footage, and Sonilo understands what’s happening frame by frame, including the pacing, emotion, and story, and generates music and sound effects that naturally fit the scene." Unlike competitors that treat music and effects as standalone assets, Sonilo’s model ties video comprehension directly to audio generation. Users can upload a clip or type a description, and the system spits out a synced soundtrack ready for commercial use.
They’re even rolling out AI dubbing with lip-sync, targeting the booming short-drama market that needs localized content without the post-production headache. With B Capital managing over $9 billion and teams in nine locations, the financial backing is serious. The new capital will fuel the computing power needed to scale their models and expand their footprint with creators and developer platforms.
A consumer app is slated for launch in the fourth quarter of 2026, putting this tech directly into the hands of individual creators. As Daisy Cai, General Partner at B Capital, noted, they believe generative audio is becoming an "essential part of the AI-native video stack." Whether Hollywood embraces this TikTok exile’s creation or fights it remains to be seen, but with major rights organizations reportedly closing agreements within weeks, Sonilo is playing to win.