InfinityEdit은 14B Helios-Distilled를 frozen backbone으로 사용하고 Transformer layer마다 Edit-Ignition Adapter를 삽입한다. Backbone의 per-layer hidden state는 history token과 current-chunk token을 이어 붙인 H=[Hhist;Hcur]이며, adapter는 history token을 수정하지 않고 current token에 세 가지 residual update를 적용한다. Noise level σ에서 AdaLN 기반 modulation mσ(⋅)=(1+sσ)⊙(⋅)+bσ의 scale과 shift를 공유 sigma-embedding module이 예측하고, 각 stage output projection은 zero-initialize한다.
History Cross-Attention은 앞쪽 La개 frame의 current token만 Query로 사용한다. Query는 Hcur1:LaWQh, Key는 HhistWKh, Value는 HhistWVh로 만들며, Attention은 softmax(QKT/√d)V를 계산한 뒤 WOh로 투영한다. 따라서 입력은 denoising chunk 앞부분과 clean history이고, 출력은 history에 맞춰 보정된 앞부분 feature이며, 이 신호가 다음 temporal attention을 통해 chunk 전체로 전달된다.
Temporal Causal Self-Attention은 각 spatial position p에 대해 시간축 길이 L의 token sequence를 독립적으로 처리한다. Causal mask는 t′≤t인 위치에 0을 두고 그 외 위치에 −∞를 두므로 frame t의 Query가 미래 frame을 참조하지 못하며, 별도의 one-dimensional temporal rotary position encoding을 Q와 K에 적용한다. 입력은 현재 chunk의 시간별 spatial token이고 출력은 앞 시점의 motion cue가 뒤 시점으로만 흐르는 temporal feature이며, 이 방향성이 streaming autoregressive 순서와 일치한다.
Edit Cross-Attention은 모든 current denoising token을 Query로 사용하고 edit instruction token을 Key·Value로 사용한다. Q=HcurWQe, K=ceditWKe, V=ceditWVe를 계산한 후 Attention 결과를 WOe로 투영해 모든 current token에 residual로 더한다. 따라서 입력은 현재 영상 feature와 자연어 지시이고 출력은 instruction-conditioned feature이며, 앞선 두 attention이 보존한 history·temporal 구조 위에 편집 의미를 입힌다.
Flow Matching은 clean target latent z0와 ∼N(0,I)를 사용해 zσ=(1−σ)z0+σ를 만들고 목표 velocity v=−z0를 학습한다. 모델은 vθ(zσ,σ,c)를 예측하며 c에는 scene prompt, edit instruction, history condition이 포함되고, 예측 velocity와 목표 velocity의 제곱 오차를 w(σ)로 가중한다. 예를 들어 σ가 0에 가까우면 latent가 clean target에 가까워져 최종 세부 묘사 보정에 해당하고, 큰 σ에서는 noise에서 layout과 큰 구조를 복원하는 방향을 학습한다.
훈련 noise는 N개 Gaussian component의 혼합에서 샘플링하며, 각 중심 μk는 추론에서 사용되는 이산 noise level이고 마지막 중심은 μN≈0이다. 첫 phase는 πk=1/N으로 모든 noise 중심을 균등하게 다루고 ωt=1로 모든 frame을 동일하게 감독한다. 두 번째 phase는 중간·저 noise 중심과 뒤쪽 frame의 loss weight를 키워, 편집 결과의 fine detail과 history anchor에서 멀어질수록 증가하는 품질 저하를 집중적으로 줄인다.
추론의 첫 ignition chunk는 adapter와 single-stage Euler schedule을 사용하고, 이후 continuation chunk는 adapter 없이 Helios의 pyramid scheduler를 사용한다. 새 chunk가 history window에 들어오면 backbone이 이를 조건으로 삼으므로 edit instruction을 계속 직접 주입하지 않아도 편집 상태가 전달된다. History window는 최신 구간만 남겨 token budget을 일정하게 유지하고, anchor는 첫 edited frame으로 재설정한 뒤 다음 edit instruction이 도착할 때까지 고정해 원본 appearance로의 회귀를 막는다.