Together with Wispr: Betting on a Voice-first Interface

Today we are announcing Together's small strategic investment in Wispr's $280M Series B, led by Menlo Ventures at a $2B valuation. Wispr is that rare trifecta: generational founders in Tanay Kothari and Sahaj Garg, a market with the wind firmly at its back, and a vision grand enough to matter. Remaking how humans interface with computers.
Let's address the obvious question first. Together is a seed-stage, operator-led firm investing across the US-India AI corridor. So why are we writing a check into a $2B Series B? Because this is deliberate, not stage drift. It's a small investment to gain access to founders we've wanted to build with for two years, a conviction bet on a category we believe will define the next decade of computing, and a front-row seat to how a voice-first company scales from consumer product-led growth into enterprise. That learning compounds directly into what we do best: sharper seed investments in voice and multimodal AI across the US and India. We'd rather pay tuition sitting at the table with the category leader than theorize from the outside. When you believe a platform shift is coming, you get as close to its epicenter as you can.
Voice is the new interface
The keyboard was designed over 150 years ago for a mechanical era, not the AI age. We've spent decades bending ourselves to machines, memorizing their menus, shortcuts, and syntax. That's finally inverting.
Speaking is about three times faster than typing. Stanford's landmark 2017 study clocked speech at 2.93x faster (153 vs. 52 WPM) with a 20.4% lower error rate. But speed is only half the story. When people speak, they give more context, more nuance, more of themselves. Text interfaces are restrictive by design; voice is expansive. For the first time, the interface is adapting to the human instead of the other way around.
Wispr is leading the Voice AI revolution
Some products create a memory the first time you use them: your first ChatGPT prompt, your first Waymo ride. Wispr Flow is one of those. It doesn't just transcribe. It strips filler words, formats your thoughts, learns your vocabulary, and produces text that sounds like you.
The behavioral data backs it up. Flow is now used by millions of people around the world. Users write a growing share of their text through Flow with every month of use, and twelve-month retention is best in class. Enterprise adoption is faster than anything we've seen in productivity software: best-in-class sales cycles, near 100% pilot success, and employees at most of the Fortune 500 using Flow, across more than 125,000 companies in all. Wispr has grown revenue by more than 150% in each of the last four quarters, as voice moves from an occasional shortcut to daily habit.
Defensibility: doing one thing extraordinarily well
The bear case on Wispr is the bear case on any great product company: won't the platforms bundle this away? It's the right question, and the honest answer is that the moat is real and widening.
First, data. Wispr has one of the largest purpose-built datasets of context-labeled audio anywhere, the kind of proprietary, edit-informed data that would cost peers an inordinate amount to recreate.
Second, purpose-built models. Wispr solves the dictation-specific problems the big labs treat as an afterthought: cross-app context, mid-sentence self-correction, formatting intent. Its new lab is now previewing Canto, Wispr's first proprietary speech model, trained on real-world usage data and contextual signals rather than controlled benchmarks. In hard acoustic conditions, think background noise, wind, strong accents, it cuts word error rates from over 30% to 5-10%. The result: users edit 30-35% fewer dictations.
Third, habit. Wispr has earned a cult-like following among power users despite a swarm of copycats. In any business, a devoted daily-use following is the ultimate moat. The Advanced Interfaces Lab and the meeting note-taker expansion keep widening the surface area.
History rhymes here. Grammarly beat autocorrect. Zoom beat Skype and Teams. They won by doing one thing so much better that bundling couldn't catch them, and by working everywhere across different platforms, still beating the incumbents’ default offering. Dropbox is the cautionary tale: a category creator commoditized once it stayed a single feature. Wispr's answer is to keep compounding a data-and-quality lead that a general-purpose assistant, optimized for everything at once, structurally cannot match – and it’s hyper-personalized that one can take it everywhere.
The market map
Here's how the space stands today.
Model & infrastructure labs: OpenAI (Whisper, gpt-realtime), ElevenLabs ($11B valuation), Deepgram ($1.3B valuation), and Cartesia are competing on latency and Word Error Rate (WER). Voice AI funding just had its biggest quarter ever ($559M in Q1'26, up 68% YoY). This layer is real and well-capitalized, but increasingly commoditized as prices fall.
Incumbent strategics: Apple (Launched their dictation tool in June ‘26 powered by Google Gemini), Google (Rambler in May ‘26), and Amazon (Alexa+, free with Prime across 600M+ devices) own distribution and will bundle. They set the floor for "good enough," but their assistants are built for breadth, not for the specific, high-frequency job of turning your speech into your writing.
Application leaders: ElevenLabs in voice agents, Wispr in dictation and productivity. They win by owning a specific, high-frequency job and the proprietary data that flows from it.
Wispr's path to becoming a large independent company is by nailing its wedge: Own voice-native text input, then expand into note-taking, agentic actions, and eventually a hyper-personalized interface that watches, listens, thinks, and acts. The defensibility isn't the current feature set; it's the compounding personal context and the labeled-audio flywheel. Cross-platform neutrality, enterprise trust, and a widening data lead are what turns a feature into a company. This is the bet we are making in partnering with Wispr.
The vision: computing at the speed of thought
The end state is a genuine step-change in productivity: no more translating thought into keystrokes and no more context-switching between apps. You speak, and the computer generates the UI, operates your tools, and carries the thread across every application. The compounding asset, and the deepest moat of all, is a hyper-personalized interface that continuously watches, listens, thinks, and acts on your behalf. Whoever builds the most personal, most trusted version of that wins the decade. We believe Wispr is the front-runner.
Together with Wispr
I've known Tanay for almost two years, and we've wanted to work together that entire time. Over long walks with him and time spent with Sahaj, I watched the product mature and the market form around it, through every copycat launch. That devotion from users, more than any single metric, is why this was a no-brainer.
And there's a reason this matters especially to us at Together: India is already a top-3 organic market for Wispr. It's the perfect proving ground for voice-first computing: a massive, mobile-first, deeply multilingual internet population of nearly 900 million that already relies on voice more than almost anywhere on earth (~60% of Wispr's dictation is non-English). What works for the next billion users often gets built for India first. We intend to help Wispr win there.
We couldn't be more excited to be building the voice-first future, together with Wispr!
