Official PyTorch implementation of:
BIGFix: Bidirectional Image Generation with Token Fixing
Victor Besnier, David Hurych, Andrei Bursuc, Eduardo Valle
arXiv:2510.12231
TL;DR: Parallel (masked-token) generation is fast but, unlike autoregressive decoding, it can't revisit a token
once committed. So early sampling mistakes get "baked in" and corrupt the rest of the generation. BIGFix trains
the model to expect this by injecting random/self-sampled tokens into the visible context (p_resample /
bigfix/utils/masking_scheduler.py), so it learns to fix errors it (or an earlier step) already committed instead of
blindly trusting them. This preserves the speed of parallel decoding while closing much of the quality gap with
slower, error-correcting samplers. On top of it, the Halton Scheduler spreads which tokens get decoded at each
step uniformly across the image, further reducing sampling errors. The paper reports substantial gains from this
combination on image generation with up to an order-of-magnitude inference speedup from
multi-token parallel prediction.
Welcome to the official implementation of BIGFix ! π
This repository includes:
- Class-to-Image Model: Generates high-quality 384x384 images from ImageNet class labels.
- Text-to-Image Model: Generates realistic images from textual descriptions, up to 512x512
(see
bigfix/txt2img_pipeline.pyfor a minimal inference example, and the BigFIX xlarge/512 checkpoint on Hugging Face).
Both models share the same self-correcting training recipe (p_resample in bigfix/utils/masking_scheduler.py) and the
Halton sampling schedule at inference time.
Explore, train, and extend our easy to use generative models! π
β BigFix/
| βββ bigfix/ <- The installable Python package (`import bigfix`)
| | βββ config/ <- Base config files (package data)
| | | βββ base_cls2img.yaml
| | | βββ base_txt2img.yaml
| | | βββ xlarge_txt2img_gpic.yaml
| | βββ dataset/ <- Data loading utilities
| | | βββ dataset.py <- PyTorch dataset class
| | | βββ dataloader.py <- PyTorch dataloader
| | βββ metrics/
| | | βββ inception_metrics.py <- Inception score and FID evaluation
| | | βββ sample_and_eval.py <- Sampling and evaluation
| | | βββ *.tsv <- Evaluation prompts
| | βββ network/
| | | βββ ema.py <- EMA model
| | | βββ transformer.py <- Transformer for class-to-image
| | | βββ txt_transformer.py <- Transformer for text-to-image
| | | βββ vq_model.py <- VQGAN architecture
| | βββ sampler/
| | | βββ confidence_sampler.py <- Confidence scheduler
| | | βββ halton_sampler.py <- Halton scheduler
| | βββ trainer/ <- Training classes
| | | βββ abstract_trainer.py <- Abstract trainer
| | | βββ cls_trainer.py <- Class-to-image trainer
| | | βββ txt_trainer.py <- Text-to-image trainer
| | βββ utils/
| | | βββ masking_scheduler.py <- BIGFix's masking + token-fixing/self-correction training recipe
| | | βββ rewards.py <- Optional reward models (CLIP, PickScore, ImageReward, HPSv2, ...)
| | βββ main.py <- Training / evaluation entry point (`python -m bigfix.main`)
| | βββ txt2img_pipeline.py <- Minimal, diffusers-style text-to-image inference wrapper
| βββ scripts/ <- Data preparation (run once before training)
| | βββ extract_vq_features.py <- Pre-extract the VQGAN codes of ImageNet
| | βββ extract_vq_txt_features.py <- Extract VQGAN codes + text embeddings from text-image WebDataset shards
| | βββ extract_train_fid.py <- Precompute FID stats for ImageNet
| βββ launch/
| | βββ run_c2i_{base,xlarge}.sh <- Training scripts for class-to-image
| | βββ run_t2i_{base,xlarge}.sh <- Training scripts for text-to-image
| βββ statics/ <- Sample images and assets
| βββ saved_networks/ <- placeholder for the downloaded models
| βββ demo_txt2img.ipynb <- Text-to-image inference demo (Jupyter / Colab)
| βββ LICENSE.txt <- MIT license
| βββ THIRD_PARTY_NOTICES.md <- Licenses of the third-party code and data used here
| βββ pyproject.toml <- Package metadata and dependencies
| βββ README.md <- This file! π
Get started with just a few steps:
git clone https://github.com/valeoai/BigFix.git
cd BigFixconda create -n bigfix python=3.10 -y
conda activate bigfix
pip install -e .This installs the bigfix package (import bigfix), including its YAML configs, so it can be used from any
working directory. The reward models (bigfix/utils/rewards.py) used for reward-conditioned text-to-image
training and the RL fine-tuning are optional: pip install -e ".[reward]".
Both examples below train with BIGFix's token-fixing recipe (--p-resample 0.2) and the Halton sampler for the
periodic visualizations. Adapt the paths, --global-bsize and --grad-cum to your hardware. Single node shown;
ready-to-edit versions are in launch/ (run_c2i_*.sh, run_t2i_*.sh). The training entry point is the module
bigfix.main, so it is started with python -m bigfix.main (or torchrun ... -m bigfix.main for several GPUs).
Download the pretrained VQGAN, then pre-extract the VQGAN codes of ImageNet once (the transformer is trained on the codes, not on the pixels):
python scripts/extract_vq_features.py --data-folder="/path/to/ImageNet/" --dest-folder="/path/to/ImageNet_codes/" \
--vqgan-folder="/path/to/vq_ds16_c2i.pt" --bsize=256 --f-factor 16 --img-size 384 --compileTrain (--data-folder must contain the Train/ and Eval/ folders written above):
torchrun --standalone --nnodes=1 --nproc_per_node=gpu -m bigfix.main \
--mode "cls-to-img" --data "imagenet_feat" --nb-class 1000 \
--data-folder "/path/to/ImageNet_codes/" --vqgan-folder "/path/to/vq_ds16_c2i.pt" \
--vit-folder "/path/to/saved_networks/imagenet_large/" --writer-log "/path/to/logs/imagenet_large/" \
--vit-size "large" --img-size 384 --f-factor 16 --codebook-size 16384 --mask-value 16384 \
--register 1 --proj 1 --dropout 0.1 --dtype "bfloat16" \
--global-bsize 256 --lr 1e-4 --warm-up 2500 --max-iter 1000000 --grad-clip 1 \
--p-resample 0.2 --resample-method shuffle \
--sampler "halton" --sched-mode "arccos" --step 32 --cfg-w 1.5 \
--resume --compileThe text-to-image model reads WebDataset shards directly: images
(jpg/png/webp) and their captions (txt), no pre-extraction needed. Put the GPIC shards in
/path/to/GPIC/train/*.tar. Download the pretrained
VQGAN beforehand; the text encoder
(google/flan-t5-xl) is fetched from Hugging Face on first use.
torchrun --standalone --nnodes=1 --nproc_per_node=gpu -m bigfix.main \
--mode "txt-to-img" --data "gpic" --nb-class 3 \
--data-folder "/path/to/GPIC/" --vqgan-folder "/path/to/vq_ds16_t2i.pt" \
--vit-folder "/path/to/saved_networks/gpic_xlarge/" --writer-log "/path/to/logs/gpic_xlarge/" \
--vit-size "xlarge" --img-size 256 --f-factor 16 --codebook-size 16384 --mask-value 16384 \
--register 1 --proj 1 --dropout 0.0 --dtype "bfloat16" \
--global-bsize 256 --grad-cum 2 --lr 1e-4 --warm-up 2500 --max-iter 200000 --grad-clip 1 \
--p-resample 0.2 --resample-method shuffle \
--sampler "halton" --sched-mode "arccos" --step 32 --cfg-w 5 --sm-temp 1.1 --top-p 0.8 \
--resumeTo go up to 512x512, raise --img-size 512, lower --global-bsize (e.g. 16) and raise --grad-cum
accordingly (see launch/run_t2i_xlarge.sh).
To quickly verify the functionality of our model, you can try this Python code:
from bigfix import MaskGITPipeline
# BigFIX text-to-image, 512x512, straight from Hugging Face
pipe = MaskGITPipeline.from_pretrained(
repo_id="llvictorll/BigFIX",
vit_filename="BigFix_XLarge_aplha02_res512_MiroFT.pth",
vqgan_filename="vq_ds16_t2i.pt",
config="xlarge_txt2img_gpic.yaml", # one of the configs shipped in bigfix/config/, or a path to your own
img_size=512,
).to("cuda")
image = pipe(
prompt="A tiny astronaut hatching from an transparent egg on the moon.",
num_inference_steps=32,
guidance_scale=5.0,
seed=42,
).images[0]
image.save("t2i_example.png")π¨ Want to try the model, but you don't have a gpu? Check out the Colab Notebook for an easy-to-run demo!
Class-to-image models are on Hugging Face, and the BIGFix text-to-image model is on Hugging Face. Use them to jump straight into inference or fine-tuning.
| Model | # Params | # Input | VQGAN | MaskGIT |
|---|---|---|---|---|
| Halton-MaskGIT-Large | 480M | 24x24 | π Download | π Download |
| BigFIX-XLarge (txt2img, 512x512) | 627M | 32x32 | π Download | π Download |
We welcome contributions and feedback! π οΈ If you encounter any issues, have suggestions, or want to collaborate, feel free to:
- Create an issue
- Fork the repository and submit a pull request
Your input is highly valued. Letβs make this project even better together! π
This project is licensed under the MIT License. See the LICENSE file for details. Some code and data adapted from other projects remain under their own licenses, listed in THIRD_PARTY_NOTICES.md.
We are grateful for the support of the IT4I Karolina Cluster in the Czech Republic for powering our experiments.
The pretrained VQGAN ImageNet (f=16/8, 16384 codebook) is from the LlamaGen official repository
If you find our work useful, please cite us and add a star β to the repository :)
BIGFix: the token-fixing / self-correction training recipe (ArXiv), the main paper this repository now implements:
@article{besnier2025bigfix,
title={BIGFix: Bidirectional Image Generation with Token Fixing},
author={Besnier, Victor and Hurych, David and Bursuc, Andrei and Valle, Eduardo},
journal={arXiv preprint arXiv:2510.12231},
year={2025}
}
The Halton Scheduling (ICLR2025): the decoding schedule used at inference time:
@inproceedings{besnier2025iclr,
title={Halton Scheduler for Masked Generative Image Transformer},
author={Victor Besnier, Mickael Chen, David Hurych, Eduardo Valle, Matthieu Cord},
booktitle={International Conference on Learning Representations (ICLR)},
year={2025}
}
Or simply the code repository for reproduction of MaskGIT (ArXiv):
@article{besnier2023pytorch,
title={A pytorch reproduction of masked generative image transformer},
author={Besnier, Victor and Chen, Mickael},
journal={arXiv preprint arXiv:2310.14400},
year={2023}
}

