DreamX-Creator 1.0은 video와 audio에 별도 Transformer backbone, token rate, positional encoding을 유지하고 shared text encoder의 조건을 modality-specific path로 전달합니다. 네트워크 전반부는 두 stream을 독립적으로 처리하며, 후반부 cross-modal block은 modality-specific self-attention과 text cross-attention 뒤에 배치됩니다. A2V는 video Query와 audio Key·Value를 사용하고 V2A는 audio Query와 video Key·Value를 사용하며, 두 경로가 동시에 활성화되면 같은 pre-fusion state에서 병렬 residual update를 계산합니다.
Cross-modal attention에서 target X와 source Y를 각각 LN한 뒤 Q, K, V projection을 적용하고 Q·K에는 공통 temporal coordinate 기반 temporal RoPE를 적용합니다. head h의 attention은 scaled dot-product와 padding mask를 거쳐 source Value를 집계합니다. 그 결과 Cᵢ,ₕ에 target token xᵢ의 hidden state 정보를 함께 넣어 sigmoid gate gᵢ,ₕ를 계산하고, gᵢ,ₕ ⊙ Cᵢ,ₕ를 head concatenate·output projection·residual addition 순서로 처리합니다. gate는 token과 head마다 달라지며, 전체 경로를 선택하는 direction mask와 달리 각 head 출력의 세기를 조정합니다.
Training mode는 noise level σᵥ와 σₐ의 비교로 방향성을 만듭니다. A2V는 σᵥ > σₐ 및 mask 1,0, V2A는 σₐ > σᵥ 및 mask 0,1, Joint는 σᵥ = σₐ 및 mask 1,1을 사용합니다. σ가 클수록 corruption이 강하므로 A2V와 V2A는 noisy target과 cleaner conditioning stream을 구성합니다. A2V·V2A에서는 stop-gradient를 Key·Value projection 전에 적용해 conditioning backbone으로 target loss가 전파되는 것을 막고, conditioning stream 자체는 자신의 Flow Matching loss로 계속 학습합니다.
Flow Matching에서는 clean latent z₀ᵐ와 독립적으로 샘플링한 Gaussian noise εᵐ를 사용해 noisy latent를 만들고, 목표 velocity는 εᵐ − z₀ᵐ입니다. modality별 손실은 유효 token 수와 feature 차원으로 정규화되며 audio·video 손실의 가중합으로 전체 학습 신호를 구성합니다. 두 pre-training 단계는 λᵥ = λₐ = 0.5, High-Quality Finetuning은 λᵥ = 0.5와 λₐ = 0.1을 사용합니다.
Stage 1의 LoRA는 rank 256, scaling factor 128, dropout 0이며 후반부 LoRA learning rate는 1 × 10⁻⁴, cross-modal·gate module은 2 × 10⁻⁵입니다. Stage 2에서는 video·audio backbone learning rate가 1 × 10⁻⁵, cross-modal module이 2 × 10⁻⁵이며, AdamW의 β₁·β₂는 0.9·0.99, ε는 10⁻¹⁰입니다. bfloat16 mixed precision, 200-step linear warm-up, gradient norm 1.0 clipping을 사용합니다.
RL에서는 normalized advantage를 video·audio·cross-modal로 분해하고 video stream에는 Aᵥ + Aₐᵥ, audio stream에는 Aₐ + Aₐᵥ를 전달합니다. frame별 cross-modal relevance score를 평균과 표준편차로 정규화한 뒤 sigmoid와 warmup 계수로 token weight를 만들므로 초기에는 거의 균등하게, 이후에는 입·소리 발생 물체처럼 동기화 반응이 큰 영역에 집중합니다. policy는 pretrained base와의 regularization을 유지해 reward over-optimization과 생성 다양성 저하를 억제합니다.
2K Refiner는 LR video와 이전 HR chunk를 입력으로 한 chunk 단위 autoregressive 구조입니다. bidirectional teacher는 과거와 미래 frame을 모두 참조해 refinement prior를 만들고, autoregressive refiner는 teacher forcing으로 과거 HR context를 사용합니다. 최종 1-step student는 자신의 과거 출력을 context로 삼아 DMD distribution matching, frozen VAE decoder를 통한 DISTS perceptual loss, ℓ2 reconstruction loss를 함께 학습하며, 각 temporal chunk마다 한 번의 denoising evaluation을 수행합니다.