Pruned CTC Cuts 180K-Vocabulary ASR Training Memory
Restricting CTC alignment states to batch target tokens preserves full-vocabulary loss and gradients while cutting full-step memory 5.1× at 180K tokens.
Underlying Paper
Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
CTC is attractive for speech recognition because it supports both offline and streaming decoding from utterance-level transcripts, without requiring token-level speech alignments. Its usual implementation, however, materializes activations over every frame and every vocabulary item. That cost becomes difficult to absorb when an ASR system adopts the native vocabulary of a large language model, where the output space can reach hundreds of thousands of tokens. The paper introduces Pruned CTC, an exact reformulation of the CTC alignment computation intended to remove that vocabulary-driven activation-memory term.
Core Contribution
The central observation is narrow but consequential: any valid CTC path for an utterance can emit only the transcript's target tokens and the blank symbol. For a batch, the union of those tokens is generally far smaller than the complete vocabulary. Pruned CTC performs forward-backward alignment computation on this batch-specific subset, while retaining the full-vocabulary normalization required by the original objective.
The authors prove equivalence to ordinary full-vocabulary CTC in both loss and gradients. This distinction matters. The method is not an approximation that changes the training target by discarding competing vocabulary items from the softmax; rather, it reduces which output entries must participate in the dynamic program after normalization. The resulting savings target the head-and-loss activations whose size otherwise grows linearly with vocabulary size.
Technical Approach
For each batch, Pruned CTC gathers the transcript-token union plus blank and computes CTC alignment quantities only for those retained symbols. Full-vocabulary logit normalization remains in place, preserving the contribution of all vocabulary items to the denominator. The paper also adds finite-beam alignment pruning, which limits the active alignment states during computation.
The authors use this objective to build LLM-CTC, a non-autoregressive ASR formulation that adapts pretrained LLMs without replacing their native tokenizer or causal-attention structure. This is a practical contrast with cross-entropy LLM ASR decoding: CTC can recognize in parallel, while the LLM vocabulary no longer makes the CTC loss prohibitively expensive. For streaming, the paper extends LLM-CTC to bounded-history attention. The stated advantage is that training does not require chunk-level speech-text alignment annotations, which are otherwise an operational burden for streaming systems.
Results and Analysis
The clearest systems result uses a Zipformer-M encoder with a 180K-token vocabulary. Pruned CTC reduces full-step memory by 5.1× relative to standard CTC, with a 17% step-time overhead. That is a favorable exchange when memory is the binding constraint: a modest increase in per-step time can make a vocabulary setting feasible that would otherwise force smaller batches, model changes, or distributed-memory workarounds.
Accuracy tests across three corpora report that Pruned CTC matches standard CTC accuracy. The paper therefore supports its primary claim that the memory reduction does not require an observed recognition-quality trade-off in those experiments. The evidence is stronger than a memory-only benchmark because it compares against the original objective, but the supplied results do not establish how the trade-off changes for other encoders, tokenizers, or much larger training configurations.
For LLM-CTC, the GigaSpeech study spans six Qwen3 sizes from 0.6B to 32B parameters. Recognition remains within 7% relative word error rate of LLM-CE while running 7–10× faster. That gap means CTC does not fully match autoregressive cross-entropy recognition quality, but the speed difference is large enough to make the approach relevant where latency or throughput dominates. In bounded-history streaming fine-tuning of Qwen3-ASR, LLM-CTC stays within 3% relative WER of matched offline models on the test set. The results support a useful deployment claim: causal, native-vocabulary LLM ASR can retain much of its offline accuracy under bounded history, without introducing chunk-alignment supervision.
Limits of the Evidence
The paper evaluates recognition comparisons on three corpora and reports relative WER gaps, so the available evidence does not show absolute error rates or behavior on broader multilingual, noisy, or domain-shifted speech. Finite-beam pruning also introduces a practical configuration choice, but the reported summary does not quantify sensitivity to beam settings. Finally, the 17% step-time cost should be treated as part of the method's operating point rather than ignored: Pruned CTC trades compute time for a substantial reduction in activation memory.
Evidence Box
strongKey Claims
- •Batch token-union CTC preserves full-vocabulary loss and gradients
- •Pruned alignment computation removes vocabulary-linear head-and-loss memory growth
- •LLM-CTC enables non-autoregressive native-vocabulary LLM ASR
- •Bounded-history streaming avoids chunk-level speech-text alignments
Key Results
- •5.1× lower full-step memory with Zipformer-M and a 180K vocabulary, at 17% step-time overhead
- •Matched standard CTC accuracy across 3 corpora
- •Within 7% relative WER of LLM-CE across six 0.6B–32B Qwen3 models on GigaSpeech
- •7–10× faster recognition than LLM-CE on the GigaSpeech comparison
Limitations & Caveats
- •Reported recognition evaluation covers 3 corpora rather than broad multilingual or domain-shifted speech
- •LLM-CTC retains up to a 7% relative WER gap versus LLM-CE on GigaSpeech
- •Finite-beam alignment pruning adds a beam-setting dependency without reported sensitivity details
- •Memory savings come with 17% step-time overhead in the 180K-vocabulary experiment