Modern multimodal systems usually maintain separate representations for understanding and generation. FLAT instead resamples either images or text into the same one-dimensional sequence of continuous tokens, jointly trained for cross-modal alignment, image synthesis, and captioning.
Nested dropout organizes information from coarse to fine. At inference time, selecting a prefix of length K provides a direct compute–detail trade-off: the first token already captures global semantics, while longer prefixes recover composition and fine-grained attributes.















