Kandinsky 6.0 Improves Its Predecessor’s Synchronized Video

A bidirectional dual-stream diffusion transformer couples a pretrained video model with a new audio stream to generate five-second clips with 44 kHz sound.

Editorial Desk·October 8, 2026·5 min readmoderate

Underlying Paper

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.

arXiv:2610.05608Submitted: Oct 6, 2026v1

Generating video and sound together is harder than generating either independently: a plausible clip still fails if speech drifts from lips, an impact lands off-beat, or camera motion contradicts the soundtrack. Kandinsky 6.0 Video addresses that coupling directly rather than attaching audio after video generation. The paper presents a 3B-parameter Lite model and a 29B-parameter Pro model for text-to-audio-video and image-to-audio-video generation, followed by a separate latent-space super-resolution pipeline for Full-HD output.

Core Contribution

The central design choice is CrossDiT, a dual-stream diffusion transformer. Kandinsky 6.0 inherits its video stream from Kandinsky 5.0, trains an audio stream from scratch, then connects the streams with bidirectional cross-modal attention. Video queries can attend to audio keys and values, and audio queries can attend back to video representations. That is a more explicit synchronization mechanism than sequential video-then-audio generation, while retaining independently useful visual and acoustic representations.

The training program is also unusually broad for a single paper: unimodal continuous pretraining precedes joint paired audio-video training; curated supervised fine-tuning is merged through model soups; reinforcement-learning post-training follows; and distillation reduces sampling cost. The authors report that the released models support five-second clips with synchronized 44 kHz audio, including lip synchronization. This makes the system most relevant to teams building short-form audiovisual generation, prototyping, and creative tools rather than long-duration video workflows.

Technical Approach

Figure 1 summarizes the main transformer block: each stream performs self-attention, text cross-attention, cross-modal attention, and a feed-forward update. The two streams keep their own text-conditioning and timestep pathways, while cross-modal layers provide the communication channel. That separation matters because it avoids forcing audio and video into one token space before each modality has formed its own representation.

Figure 1. Simplified scheme of a single Kandinsky 6.0 Video transformer block. The video (top) and audio (bottom) streams each apply self-attention, cross-attention to the text embeddings (T2V and T2A), cross-modal attention, and a feed-forward layer. In the V2A block, video queries attend to audio keys and values; in the A2V block, audio queries attend to video keys and values.

For high-resolution output, the paper splits restoration from generation. A Latent Upscaler first brings a low-quality latent to the working resolution, then an SR-DiT restores details using flow matching. SR-DiT operates without text conditioning and uses a noisy video latent, a first-frame high-quality anchor, and a binary anchor mask as concatenated inputs. Its concatenated input has 129 channels: 64 noisy-latent channels, 64 anchor channels, and one mask channel. The architecture has 32 transformer blocks, a 1792-dimensional hidden size, 28 attention heads of dimension 64, and about 1.413B parameters. Figure 3 depicts this input arrangement and its timestep-modulated transformer blocks.

Figure 3. SR-DiT architecture. The noisy video latent, first-frame HQ anchor, and binary anchor mask are concatenated along the channel dimension and processed by timestep-modulated DiT blocks. The output head predicts the flow-matching velocity. The model operates without text conditioning.

The paper uses a dense-attention training phase, then substitutes sparse NABLA attention and sliding-tile attention to reduce cost. A final π-Flow distillation stage trains a student to make two SR-DiT calls, down from a 10-function-evaluation teacher process. This is a practical engineering decision: the more expensive restoration model is kept separate from the multimodal base generator and optimized for few-step inference.

Results and Analysis

The evidence is strongest for improvement over the authors’ previous model. In side-by-side human evaluation for text-to-video, Kandinsky 6.0 Video Pro received a 0.58 preference share for overall visual quality against Kandinsky 5.0 Video Pro’s 0.16, with 0.26 ties. It also led on artifacts, 0.49 versus 0.19. Motion realism and physics is less decisive: Pro received 0.28 preference versus 0.13 for its predecessor, but 0.59 of comparisons were ties. In image-to-video, its overall-visual-quality preference was 0.64 versus 0.11, with 0.25 ties. Those results support a real visual upgrade, but they do not establish a broad absolute quality ranking.

The comparisons with other systems are mixed, which is the paper’s most useful finding. Against LTX 2.5 in text-to-audio-video, Kandinsky Pro had a 0.49 overall-visual-quality preference share versus 0.26, while LTX led audio-video synchronization, 0.13 versus 0.03. The paper describes a similar split against Veo 3.1 Fast: Kandinsky is preferred for cleaner visuals and image preservation in image-conditioned generation, while Veo is preferred for prompt following and several audio dimensions. MiniMax H3 is preferred on most reported criteria. The authors also report the best VABench scores among the evaluated models for speech quality, audio aesthetics, text-video alignment, audio-video alignment, and lip-sync desynchronization; reinforcement learning reduced the Pro model’s generated-speech word error rate by 47%.

The editorial reading is therefore narrower than a blanket leadership claim. CrossDiT appears to improve the team’s own audiovisual stack and produces competitive speech and visual quality, but the human comparisons show criterion-specific trade-offs rather than dominance. The strongest deployment argument is the open release of weights, source code, and diffusers integration under MIT terms; the strongest technical caveat is that the demonstrated product remains short, super-resolved video rather than native long-form Full-HD generation.

Limits and Safety

The authors identify a five-second duration limit, standard-definition base generation before super-resolution, and remaining gaps to the strongest proprietary systems in visual and overall audio quality. The system can also create persuasive synchronized speech and faces, creating impersonation, deception, and non-consensual-content risks. The paper asks downstream users to label synthetic media, obtain appropriate rights and consent, moderate deployment, and treat outputs as generated rather than factual depictions.

Evidence Box

moderate

Key Claims

  • •Bidirectional CrossDiT aligns pretrained video and newly trained audio streams
  • •Continuous pretraining and post-training preserve unimodal fidelity while learning paired audiovisual generation
  • •Two-call π-Flow distillation makes the super-resolution stage suitable for few-step inference
  • •Released weights, source code, and diffusers integration support open use

Key Results

  • •0.58 visual-quality preference for Pro versus 0.16 for Kandinsky 5.0 Pro in text-to-video
  • •0.64 visual-quality preference for Pro versus 0.11 for Kandinsky 5.0 Pro in image-to-video
  • •0.49 visual-quality preference for Pro versus 0.26 for LTX 2.5 in text-to-audio-video
  • •47% reduction in generated-speech word error rate for Pro after reinforcement learning

Limitations & Caveats

  • •Generation is limited to 5-second clips
  • •Base generation uses standard definition before a separate Full-HD super-resolution stage
  • •MiniMax H3 is preferred on most reported comparison criteria
  • •Strongest proprietary systems retain gaps in visual and overall audio quality

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.