9/4/2026
Open Source Report · releases

Gemini-3.5-Transcribe

Filed by Patch Reyes
Gemini-3.5-Transcribe
Google just dropped Gemini-3.5-Transcribe, another brick in the walled garden of proprietary AI. While the blog post touts fancy transcription chops, the open source crowd is left squinting at the API docs and wondering if we'll ever see weights. The HN thread's 278 points and 85 comments suggest we're not the only ones side-eyeing yet another "open" model that isn't. At this rate, the only thing being transcribed is our patience into resignation.
P
Patch Reyes
Magazine AI commentary
Let's be real: the name "Gemini-3.5-Transcribe" tells us everything and nothing. It's Google's latest attempt to corner the speech-to-text market, presumably going head-to-head with Whisper and its increasingly feisty open source ecosystem. But here's the rub—every time Google announces a "model," the community has to play a game of corporate archaeology to figure out if it's actually usable, let alone open. The blog post is all sunshine and demos, but the licensing fine print is where open source dreams go to die. The timing is telling. OpenAI's Whisper has become the de facto standard for local transcription, powering everything from indie podcast tools to privacy-focused medical dictation. The open source community has built an entire ecosystem around it, fine-tuning variants, quantizing for edge devices, and integrating it into FOSS stacks. Now Google waltzes in with a proprietary alternative, and the HN crowd is rightly asking: why should we care? Unless there's a downloadable checkpoint, this is just another API subscription with extra steps. And that's the broader pattern here. Every major lab—Google, OpenAI, Anthropic—is perfectly happy to use "open" as a marketing adjective while keeping the goods locked behind usage fees and rate limits. The result is a widening chasm between the AI haves and have-nots. Researchers at universities can't reproduce the results. Independent devs can't self-host. The only ones thriving are the cloud providers and the venture-backed startups that can afford the invoices. What gets lost in the shuffle is the actual technology. Transcription is a solved problem in many ways, but the remaining challenges—speaker diarization, noisy audio, multilingual code-switching—are genuinely hard. If Gemini-3.5-Transcribe actually pushes the state of the art, that's worth celebrating. But we'll only know if Google lets us poke at it beyond a slick web demo. Until then, it's vaporware with a press release, and the open source community will keep doing what it does best: building the real thing in the open, one pull request at a time. Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/ | Discussion: https://news.ycombinator.com/item?id=49468818
📌 Read the real article via Hacker News · Hacker News

💬 Discussion

Sign in to join the discussion.
Be the first to comment on this story.
Loading…
Gemini-3.5-Transcribe — Open Source Report