🌐 Project Page | 📄 arXiv
Alicja Polowczyk*, Agnieszka Polowczyk*, Piotr Borycki, Joanna Waczyńska, Jacek Tabor, Przemysław Spurek
(*equal contribution)
DIAMOND is a training-free, inference-time guidance framework that tackles one of the most persistent challenges in modern text-to-image generation: visual and anatomical artifacts.
While recent models such as FLUX achieve impressive realism, they still frequently produce distorted structures, malformed anatomy, and visual inconsistencies. Unlike existing post-hoc or weight-modifying approaches, DIAMOND intervenes directly during the generative process by reconstructing a clean sample estimate at each step and steering the sampling trajectory away from artifact-prone latent states.
The method requires no additional training, no finetuning, and no weight modification, and can be applied to both flow matching models and standard diffusion models, enabling robust, zero-shot, high-fidelity image synthesis with substantially reduced artifacts.
- Feb. 2026: Initial codebase released with support for FLUX models (FLUX.1-dev, FLUX-schnell, FLUX-2-dev).
- Feb. 2026: Paper is available on arXiv.
- Oct. 2026: Added support for SDXL.
We provide two separate environment configurations depending on the model variant.
Create and activate the Conda environment:
conda create -n diamond python=3.11 -y
conda activate diamondInstall PyTorch and remaining dependencies:
pip install torch==2.6.0 torchvision==0.21.0 --index-url https://download.pytorch.org/whl/cu126
pip install -r requirements.txtRequires a newer version of diffusers installed directly from GitHub.
conda create -n diamond-flux2 python=3.11 -y
conda activate diamond-flux2
pip install torch==2.6.0 torchvision==0.21.0 \
--index-url https://download.pytorch.org/whl/cu126
pip install -r requirements2.txtDownload the required detector and place it in checkpoints/:
- DiffDoctor artifact detector (
ad_pytorch_model.bin) - RichHF baseline detector (
ad_richhf_baseline_model.bin)
The artifact detector was originally released with DiffDoctor.
We provide the model weights used in our evaluation of state-of-the-art artifact mitigation methods.
| Base Model | DiffDoctor | HPSv2 | HandsXL |
|---|---|---|---|
| FLUX.1 [dev] | Download | Download | Download |
| FLUX.1 [schnell] | Download | Download | — |
| SDXL | — | — | Download |
Select the appropriate animals, people, or words checkpoint from the DIAMOND Model Weights v1.0 release.
Download the selected LoRA checkpoint and place it in checkpoints/lora/.
HandsXL weights originate from the official HandsXL repository.
Evaluation datasets are not distributed with the repository. Generate them locally using the provided prompt files and the instructions in Generate Custom Evaluation Dataset below.
Move to the repository root:
cd DIAMONDYou can select the base model using model=dev (FLUX.1 [dev]) or model=schnell (FLUX.1 [schnell]).
Setting guidance.enabled=true enables DIAMOND guidance during sampling. To run without DIAMOND (baseline), set guidance.enabled=false.
You can also modify the loss type and the lambda_schedule to explore different guidance behaviors.
python src/generate_single_image.py \
model=dev \
'prompt="Luxury crystal blue diamond, premium brand mark, vector style, simple and iconic, 4k resolution"' \
seed=100285 \
guidance.enabled=false \
loss=power \
lambda_schedule=power \
lambda_schedule.start=25 \
lambda_schedule.end=1 \
lambda_schedule.power=2 \
output.run_name=example_runFor FLUX.2 [dev], use the separate script:
python src/generate_single_image_flux2.py \
model=flux2_dev \
'prompt="Luxury crystal blue diamond, premium brand mark, vector style, simple and iconic, 4k resolution"' \
seed=100285 \
output.run_name=example_runFor SDXL, use:
python src/generate_single_image_sdxl.py \
model=sdxl_base \
'prompt="A sample prompt."' \
seed=100285 \
guidance.enabled=false \
output.run_name=example_sdxlImportant
Activate the correct Conda environment before running (see Environment Setup).
Outputs are saved to the outputs/ directory.
See the 📦 SOTA Method Weights table for model support. Enable LoRA and set the appropriate checkpoint in lora.path.
python src/generate_single_image.py \
model=dev \
'prompt="A South Asian man, 35 years old, with a visual impairment, reading braille books in a library."' \
seed=100283 \
lora=enabled \
lora.path="checkpoints/lora/people_handv1.safetensors" \
lora.scale=0.1 \
guidance.enabled=false \
output.run_name=lora_examplepython src/generate_single_image_sdxl.py \
model=sdxl_base \
'prompt="A sample prompt."' \
seed=100283 \
lora=enabled \
lora.path="checkpoints/lora/people_handv55.safetensors" \
lora.scale=0.1 \
guidance.enabled=false \
output.run_name=handsxl_sdxlImportant
When using LoRA-based SOTA methods, always set guidance.enabled=false.
The generation setup is identical to single-image generation. DIAMOND can be enabled or disabled using guidance.enabled=true/false.
LoRA-based SOTA methods can be used by setting lora=enabled and specifying lora.path.
For FLUX.1 [dev], FLUX.1 [schnell], use:
python src/generate_images_csv.py \
model=schnell \
csv_path=/path/to/prompts.csv \
loss=power \
lambda_schedule=power \
lambda_schedule.start=25 \
lambda_schedule.end=1 \
lambda_schedule.power=2 \
output.run_name=example_runFor FLUX.2 [dev], use:
python src/generate_images_csv_flux2.py \
model=flux2_dev \
csv_path=/path/to/prompts.csv \
loss=power \
lambda_schedule=power \
lambda_schedule.start=25 \
lambda_schedule.end=1 \
lambda_schedule.power=2 \
output.run_name=example_runFor SDXL, use:
python src/generate_images_csv_sdxl.py \
model=sdxl_base \
csv_path=/path/to/prompts.csv \
output.run_name=example_sdxlThis script computes quantitative evaluation metrics for generated images.
Results are saved to outputs/metrics/results.txt by default and can be customized if needed.
The following metrics are computed: CLIP-T, MeanArtifactFreq (%), ArtifactPixelRatio (%), MAE, MAE(A), MAE(NA).
python src/generate_metrics.py \
metrics.generated_dir=/path/to/generated/images \
metrics.reference_dir=/path/to/reference/images \
metrics.prompts_csv=/path/to/prompts.csv For computing ImageReward, please refer to the official repository: https://github.com/zai-org/ImageReward
Note
Generate the prompt-and-seed CSV files locally using Generate Custom Evaluation Dataset below.
Generate a dataset by searching for valid seeds and saving prompts + seeds into a CSV file.
Prompts are provided as .txt files (one per line). Example files are in prompts/.
The script also saves generated images and corresponding artifact masks.
The seed parameter specifies the starting seed from which the search begins
python src/generate_dataset.py \
model=dev \
seed=100000 \
dataset.prompts_file=prompts/animals_100_eval.txt \
dataset.name=my_dataset \
output.run_name=dataset_genNote
Dataset generation is supported for FLUX.1 [dev], FLUX.1 [schnell], FLUX.2 [dev], and SDXL.
To switch models, only the script name and the model value need to be changed:
generate_dataset.py→ dev/schnellgenerate_dataset_flux2.py→ flux2_devgenerate_dataset_sdxl.py→ sdxl_base
If you find this work useful, please cite:
@misc{polowczyk2026diamonddirectedinferenceartifact,
title={DIAMOND: Directed Inference for Artifact Mitigation in Flow Matching Models},
author={Alicja Polowczyk and Agnieszka Polowczyk and Piotr Borycki and Joanna Waczyńska and Jacek Tabor and Przemysław Spurek},
year={2026},
eprint={2602.00883},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.00883},
}