Wink Pings

NVIDIA's PixelUMM Eliminates Both VAE and ViT: Understanding and Generating Images/Videos Directly in Pixel Space

NVIDIA and the University of Waterloo have proposed PixelUMM, which removes VAEs and visual encoders and uses a single decoder-only Transformer to understand and generate images and videos in pixel space. The 8B model, along with its code and weights, has been open-sourced.

NVIDIA and the University of Waterloo have released PixelUMM. Just a few days after the paper was published, Niels Rogge reposted it on X: it's already the number one trending project on Papers with Code.

PixelUMM directly removes both VAEs and visual encoders. VAEs compress images into latent variables, and visual encoders convert images into tokens—these two components are relied on by nearly all mainstream multimodal models. Bypassing both, PixelUMM uses a single decoder-only Transformer to perform understanding and generation directly on raw pixels. Images are split into 16×16 patches, while videos are split into 4-frame tubelets, which are then projected via a single linear layer into the Qwen3-8B backbone. Understanding tasks and generation tasks each have their own set of expert parameters, but text, clean pixels, and noisy pixels all share the same self-attention module. This is a Mixture-of-Transformers design.

![PixelUMM promotional image](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHTxiSTJWIAA13lM%3Fformat%3Djpg%26name%3Dlarge)

The paper notes that existing Unified Multimodal Models (UMMs) generally rely on two separate visual representation modules, which lengthens visual context and makes the training pipeline more complex. By removing these two external components, PixelUMM extends the encoder-free image approach of SenseNova-U1 to image-to-video generation and video editing.

To identify which design choices deliver practical benefits, the authors conducted eight ablation studies labeled F1 through F8. Several key conclusions are worth noting:

- 16×16 image patches are easier to train than 32×32 patches, even though 32×32 patches can fit four times more image content at the same sequence length.

- The more aggressively video tokens are compressed, the harder the model is to train. The final choice is p16/t4, where each token covers 16×16 pixels and 4 frames, with a compression rate matching that of Wan2.2's VAE.

- The gradient norm in pixel space is almost identical to that in VAE latent space. Training in pixel space is not significantly faster—it just operates in a different loss space.

- To reach the same generation and text loss levels, an 8B model only requires roughly one-third the training steps of a 1.7B model.

- More GPUs primarily accelerate the convergence of text loss. For the same 1.7B model, using 128 GPUs instead of 8 reduces the number of steps needed to lower text loss (CE) to 1.0 by roughly 9 times, but only reduces the steps needed to lower generation loss (MSE) to 0.065 by around 1.3 times.

- A linear output head will leave faint 16-pixel grid boundaries in smooth regions like the sky. Switching to a convolutional output head alleviates this issue, but the officially released checkpoint and all demos still use a linear head because changing the head requires full retraining.

- For video understanding, 4 FPS tubelet input achieves almost identical scores to 1 FPS frame-by-frame input.

For understanding capability, the official team evaluated 21 image tasks via LMMS-Eval, totaling 64,750 generation runs. The 8B checkpoint scored 90.42 on DocVQA, 82.96 on ChartQA, 70.53 on MVBench, and 57.33 on Video-MME (no subtitles). When compared to models including BAGEL, TUNA, Qwen2.5-VL, and LLaVA-OV in the paper's tables, overall performance is comparable, but it still lags behind understanding-specialized models like Qwen3-VL on complex reasoning tasks such as MMMU. The official team notes that different models use different training data, so no direct conclusion on architecture quality can be drawn.

![Image understanding benchmark](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHTlUooJX0AAXEos%3Fformat%3Dpng%26name%3Dlarge)

![Video understanding benchmark](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHTlUrHVXkAAGQSn%3Fformat%3Dpng%26name%3Dlarge)

The official team also lists known failure modes: object count drifts when multiple similar objects overlap, hands still frequently have extra fingers, and physical interactions occasionally violate common sense. Removing the VAE does not automatically solve these longstanding problems.

However, the value of PixelUMM does not lie in another benchmark leaderboard refresh. It simplifies the architecture of multimodal models by one key step: one single Transformer, no dependency on external VAEs or visual encoders, and reads raw pixels directly. Code, weights, and demos are all publicly available.

- Project page: https://nv-tlabs.github.io/PixelUMM

- Paper: https://arxiv.org/abs/2609.38597

- Code: https://github.com/nv-tlabs/PixelUMM

- Model: https://huggingface.co/nvidia/PixelUMM

Another official demo: a white off-road vehicle driving on a forest dirt road:

发布时间: 2026-10-05 03:55