IBM released Granite Speech 5.0 Turbo CTC and 5.0 Turbo CTC NC on 25 August, the same day as Granite 4.2. Two English speech-recognition models, 470 million parameters each, and no language-model backbone. Earlier Granite Speech stacks bolted a Conformer encoder onto a Granite LLM. This pair leaves the LLM off. Connectionist temporal classification maps audio to text in one non-autoregressive pass.

IBM Research says the models are aimed at laptops, smartphones, and other edge devices. The Hugging Face post adds native transformers support, a WebGPU streaming demo, Apache 2.0 on the commercial checkpoint and CC-BY-NC-SA-4.0 on the extra-data twin. Training data on the Apache model is about 60,000 hours of public English audio.

The speed claim is IBM's. Current Hugging Face Open ASR leaders sit around 6,000 RTFx. Granite Speech 5.0 Turbo CTC came in closer to 12,600 on a single H200, which the company frames as more than 3.5 hours of speech in one second, batched. The Apache model reports 5.00 percent aggregate WER on the public short-form tests; the non-commercial twin reports 4.85 percent. IBM labels some of those Open ASR charts unofficial. Keep the label on.

IBM cuts the LLM out of Granite Speech and ships 470 million parameters
Will Fisher from Richmond, VA, United States / Wikimedia Commons (CC BY-SA 2.0)Download

The date is 25 August, not a later bake. Granite Speech 5.0 Turbo CTC and Granite Speech 5.0 Turbo CTC NC landed the same calendar day as Granite 4.2. The reasoner family is a separate file. This file is the speech pair. Two English automatic-speech-recognition checkpoints. Four hundred and seventy million parameters each. No language-model backbone. The earlier Granite Speech stacks bolted a Conformer encoder onto a Granite LLM. This pair leaves the LLM off. That is the deletion the title is pointing at. Connectionist temporal classification maps audio to text in one non-autoregressive pass. IBM Research says the target is laptops, smartphones, and other edge devices. The Hugging Face post is the serving note: native transformers support, a WebGPU streaming demo, Apache 2.0 on the commercial checkpoint, CC-BY-NC-SA-4.0 on the extra-data twin. Training data on the Apache model is about 60,000 hours of public English audio.

Whiteboard cut
A Granite Speech Model diagram on a whiteboard, scissors on a Speech Decoder card, a crate stencilled Granite Speech.

The second still is the cut made physical. A whiteboard labelled Granite Speech Model. Audio in. An encoder path. A Speech Decoder card under a pair of scissors. A crate in the foreground stencilled Granite Speech. The picture is the metaphor it is. The live objects are two 470-million-parameter English transcribers that IBM put on Hugging Face on 25 August. They are not a decoder module in a wooden box. They are CTC models. Connectionist temporal classification. One non-autoregressive pass from audio to text. No LLM in the loop. Earlier Granite Speech stacks kept the LLM. This pair does not.

IBM's speed number is the one that will travel, so the attribution stays glued to it. Hugging Face Open ASR leaders, as IBM cites them, sit around 6,000 RTFx. Granite Speech 5.0 Turbo CTC, on a single H200, came in closer to 12,600 RTFx. The company's Hugging Face write-up frames that H200 figure as more than 3.5 hours of speech in one second, batched. Those are IBM's measurements. The board is not rerun here. It is printing the claim with the label IBM already put on some of the Open ASR charts: unofficial. Keep the label on. The commercial checkpoint and the extra-data twin are still the two files that shipped.

Word error rate is the other IBM table. The Apache model reports 5.00 percent aggregate WER on the public short-form tests. The non-commercial twin, Granite Speech 5.0 Turbo CTC NC, reports 4.85 percent. Two licences. Two numbers. Same 470 million parameters. Same English ASR job. The NC checkpoint is the extra-data twin under CC-BY-NC-SA-4.0. The commercial checkpoint is Apache 2.0. Training hours named for the Apache model: about 60,000 hours of public English audio. A training-hour figure for the twin is not on this page. The Apache hours IBM named stay on this page.

On the FFASR far-field board, as of 25 August, IBM said the non-commercial model ranked fifth in accuracy and the Apache model ninth, and that both were the two fastest entries, measured on an L4. Fifth and ninth are accuracy ranks. Fastest is a speed rank on that board, on that GPU, on that date. IBM said so. Keep the date on. 25 August. L4. Far-field. The same company that put 12,600 RTFx on an H200 also put a speed rank on an L4. Two chips. Two sentences. Do not collapse them.

The architecture IBM is willing to print for this pair is small on purpose. Sixteen Conformer blocks. Eightfold temporal subsampling. Greedy decode. That is the stack that remains once the LLM is off. Earlier Granite Speech bolted a Conformer encoder onto a Granite language-model backbone. The backbone is gone. Speech translation is gone with the decoder. Keyword extras that lived on the old decoder path are not this drop. What shipped is a 470-million-parameter English transcriber that fits on a phone. IBM Research is explicit about the devices: laptops, smartphones, other edge hardware. Native transformers support is how the Hugging Face post says a builder loads it. A WebGPU streaming demo is how the same post says a browser can try it. Those two serving facts sit next to the licences, not instead of them.

Apache 2.0 on the commercial checkpoint is the sentence that makes the file a tool a company can ship. CC-BY-NC-SA-4.0 on the extra-data twin is the sentence that keeps the other file in the non-commercial column. Two English ASR models. Same parameter count. Same CTC pass. Same 25 August. Same day as Granite 4.2. The speech pair is not merged into the reasoner family. Granite 4.2 is the same-day neighbour. Granite Speech 5.0 Turbo CTC and 5.0 Turbo CTC NC are the objects.

Audio desk
A dark audio desk with an IBM Granite Speech card over a waveform, a Shipped badge on the card.

The third still is a dark audio desk: a waveform, a session list, an IBM Granite Speech card with a Shipped badge over the timeline. The still is a picture of a product card, not a spec sheet. The live checkpoints are 470 million parameters each. INT8 on a mock card is not added to the 25 August file. Size on a mock card is not added either. The Hugging Face post and IBM Research are the sources for the pair that actually shipped: Granite Speech 5.0 Turbo CTC, Granite Speech 5.0 Turbo CTC NC, English ASR, 470 million parameters, CTC, no LLM backbone.

Edge is the point of leaving the backbone off. A Conformer encoder bolted onto a Granite LLM is a speech stack that still wants a language-model decoder in memory. CTC does not. One non-autoregressive pass. Audio in, text out. Sixteen Conformer blocks. Eightfold temporal subsampling. Greedy decode. Aimed at the machines people already carry. Laptops. Phones. Other edge devices. A WebGPU streaming demo is the browser version of that sentence. Native transformers support is the library version. About 60,000 hours of public English audio is the Apache training sentence. Those hours are not turned into a multilingual claim. IBM named English. The models are English ASR.

Speed, again, because it is the number IBM put next to the Open ASR board. Leaders around 6,000 RTFx. Granite Speech 5.0 Turbo CTC closer to 12,600 on one H200. More than 3.5 hours of speech in one second, batched, in the company's framing. Unofficial on some of the charts, by IBM's own label. 5.00 percent aggregate WER Apache. 4.85 percent NC. FFASR far-field, 25 August: NC fifth, Apache ninth, both the two fastest on an L4. Those are the vendor tables. They are printed as vendor tables.

What 25 August did not ship is a speech translator. Translation left with the decoder. What it did ship is a pair of English transcribers small enough, in IBM Research's telling, for a laptop and a phone. Four hundred and seventy million parameters is the size both times. Apache 2.0 is the commercial licence. CC-BY-NC-SA-4.0 is the twin. Native transformers. WebGPU streaming demo. CTC. No LLM. Same day as Granite 4.2.

The date is 30 August 2026, a Sunday, five days after the drop. The pair has not been withdrawn. The H200 number has not been replaced by an independent rerun. The unofficial label has not been taken off the charts IBM labelled unofficial. Until a later board moves, the honest objects remain the two files: Granite Speech 5.0 Turbo CTC and Granite Speech 5.0 Turbo CTC NC. English. 470 million parameters. No language-model backbone. CTC in one non-autoregressive pass. Aimed at laptops, smartphones, and other edge devices. About 60,000 hours of public English audio on the Apache model. 12,600 RTFx on an H200, IBM's claim, against Open ASR leaders around 6,000. Keep the claim on IBM. Keep the stills as stills. Keep the decoder off the stack that shipped.