Image-to-Image
Diffusers
Safetensors
LDMPipeline
computed-tomography
ct-reconstruction
diffusion-model
latent-diffusion
inverse-problems
dm4ct
sparse-view-ct
Instructions to use jiayangshi/lodoind_latent_diffusion with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use jiayangshi/lodoind_latent_diffusion with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("jiayangshi/lodoind_latent_diffusion", dtype=torch.bfloat16, device_map="cuda") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Improve model card: add pipeline tag, paper link, and fix usage snippet
Browse filesHi! I'm Niels from the community science team at Hugging Face.
I've opened this PR to improve the model card for this artifact:
- Added `pipeline_tag: image-to-image` to the YAML metadata to improve discoverability.
- Added a link to the Hugging Face paper page.
- Fixed the `diffusers` usage snippet to be syntactically correct and use the `LDMPipeline` as specified in your `model_index.json`.
This helps researchers find and cite your work more easily!
README.md
CHANGED
|
@@ -1,14 +1,15 @@
|
|
| 1 |
---
|
| 2 |
-
license: mit
|
| 3 |
library_name: diffusers
|
|
|
|
|
|
|
| 4 |
tags:
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
---
|
| 13 |
|
| 14 |
# Latent Diffusion Model โ LoDoInd (DM4CT)
|
|
@@ -16,9 +17,9 @@ tags:
|
|
| 16 |
This repository contains the pretrained **latent-space diffusion model** used in the
|
| 17 |
**DM4CT: Benchmarking Diffusion Models for CT Reconstruction (ICLR 2026)** benchmark.
|
| 18 |
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
|
| 23 |
---
|
| 24 |
|
|
@@ -26,60 +27,71 @@ This repository contains the pretrained **latent-space diffusion model** used in
|
|
| 26 |
|
| 27 |
This model learns a **prior over CT reconstruction images in a compressed latent space** using a denoising diffusion probabilistic model (DDPM).
|
| 28 |
|
| 29 |
-
Unlike
|
| 30 |
|
| 31 |
- **Architecture**:
|
| 32 |
- VQ-VAE (image encoder/decoder)
|
| 33 |
- 2D UNet operating in latent space
|
| 34 |
- **Input resolution (image space)**: 512 ร 512
|
| 35 |
-
- **Latent resolution**: (insert latent size, e.g., 64 ร 64)
|
| 36 |
- **Channels**: 1 (grayscale CT slice)
|
| 37 |
- **Training objective**: ฮต-prediction (standard DDPM formulation)
|
| 38 |
- **Noise schedule**: Linear beta schedule
|
| 39 |
- **Training dataset**: Industry CT dataset (LoDoInd)
|
| 40 |
- **Intensity normalization**: Rescaled to (-1, 1)
|
| 41 |
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
This model is intended to be combined with data-consistency correction for CT reconstruction.
|
| 45 |
|
| 46 |
---
|
| 47 |
|
| 48 |
## ๐ Dataset: LoDoInd
|
| 49 |
|
| 50 |
-
|
| 51 |
-
https://zenodo.org/records/10391412
|
| 52 |
|
| 53 |
-
|
| 54 |
-
-
|
| 55 |
-
- Rescale reconstructed slices to (-1, 1)
|
| 56 |
-
- No geometry information is embedded in the model
|
| 57 |
-
|
| 58 |
-
The model learns an unconditional latent prior over CT slices.
|
| 59 |
|
| 60 |
---
|
| 61 |
|
| 62 |
## ๐ง Training Details
|
| 63 |
|
| 64 |
-
- Optimizer: AdamW
|
| 65 |
-
- Learning rate: 1e-4
|
| 66 |
-
-
|
| 67 |
-
- Training
|
| 68 |
-
- Hardware: NVIDIA A100 GPU
|
| 69 |
-
|
| 70 |
-
Training scripts:
|
| 71 |
-
- Latent diffusion: https://github.com/DM4CT/DM4CT/blob/main/train_latent.py
|
| 72 |
-
- Autoencoder training: (insert if separate)
|
| 73 |
|
| 74 |
---
|
| 75 |
|
| 76 |
## ๐ Usage
|
| 77 |
|
|
|
|
|
|
|
| 78 |
```python
|
| 79 |
from diffusers import LDMPipeline
|
|
|
|
| 80 |
|
| 81 |
-
|
| 82 |
"jiayangshi/lodoind_latent_diffusion"
|
| 83 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 84 |
|
| 85 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
|
|
|
| 2 |
library_name: diffusers
|
| 3 |
+
license: mit
|
| 4 |
+
pipeline_tag: image-to-image
|
| 5 |
tags:
|
| 6 |
+
- computed-tomography
|
| 7 |
+
- ct-reconstruction
|
| 8 |
+
- diffusion-model
|
| 9 |
+
- latent-diffusion
|
| 10 |
+
- inverse-problems
|
| 11 |
+
- dm4ct
|
| 12 |
+
- sparse-view-ct
|
| 13 |
---
|
| 14 |
|
| 15 |
# Latent Diffusion Model โ LoDoInd (DM4CT)
|
|
|
|
| 17 |
This repository contains the pretrained **latent-space diffusion model** used in the
|
| 18 |
**DM4CT: Benchmarking Diffusion Models for CT Reconstruction (ICLR 2026)** benchmark.
|
| 19 |
|
| 20 |
+
- **Paper:** [DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction](https://huggingface.co/papers/2602.18589)
|
| 21 |
+
- **Project Page:** [https://dm4ct.github.io/DM4CT/](https://dm4ct.github.io/DM4CT/)
|
| 22 |
+
- **Codebase:** [https://github.com/DM4CT/DM4CT](https://github.com/DM4CT/DM4CT)
|
| 23 |
|
| 24 |
---
|
| 25 |
|
|
|
|
| 27 |
|
| 28 |
This model learns a **prior over CT reconstruction images in a compressed latent space** using a denoising diffusion probabilistic model (DDPM).
|
| 29 |
|
| 30 |
+
Unlike pixel diffusion models, diffusion is performed in the latent space of a pretrained autoencoder (VQ-VAE).
|
| 31 |
|
| 32 |
- **Architecture**:
|
| 33 |
- VQ-VAE (image encoder/decoder)
|
| 34 |
- 2D UNet operating in latent space
|
| 35 |
- **Input resolution (image space)**: 512 ร 512
|
|
|
|
| 36 |
- **Channels**: 1 (grayscale CT slice)
|
| 37 |
- **Training objective**: ฮต-prediction (standard DDPM formulation)
|
| 38 |
- **Noise schedule**: Linear beta schedule
|
| 39 |
- **Training dataset**: Industry CT dataset (LoDoInd)
|
| 40 |
- **Intensity normalization**: Rescaled to (-1, 1)
|
| 41 |
|
| 42 |
+
This model is intended to be combined with data-consistency correction for CT reconstruction tasks.
|
|
|
|
|
|
|
| 43 |
|
| 44 |
---
|
| 45 |
|
| 46 |
## ๐ Dataset: LoDoInd
|
| 47 |
|
| 48 |
+
The model was trained on the industrial CT dataset [LoDoInd](https://zenodo.org/records/10391412).
|
|
|
|
| 49 |
|
| 50 |
+
- Reconstructed slices were rescaled to the range (-1, 1).
|
| 51 |
+
- The model learns an unconditional latent prior over CT slices; no specific geometry information is embedded in the weights.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
---
|
| 54 |
|
| 55 |
## ๐ง Training Details
|
| 56 |
|
| 57 |
+
- **Optimizer**: AdamW
|
| 58 |
+
- **Learning rate**: 1e-4
|
| 59 |
+
- **Hardware**: NVIDIA A100 GPU
|
| 60 |
+
- **Training scripts**: Available in the [DM4CT GitHub repository](https://github.com/DM4CT/DM4CT/blob/main/train_latent.py).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 61 |
|
| 62 |
---
|
| 63 |
|
| 64 |
## ๐ Usage
|
| 65 |
|
| 66 |
+
You can load and use this model with the `diffusers` library:
|
| 67 |
+
|
| 68 |
```python
|
| 69 |
from diffusers import LDMPipeline
|
| 70 |
+
import torch
|
| 71 |
|
| 72 |
+
pipeline = LDMPipeline.from_pretrained(
|
| 73 |
"jiayangshi/lodoind_latent_diffusion"
|
| 74 |
)
|
| 75 |
+
pipeline.to("cuda")
|
| 76 |
+
|
| 77 |
+
# Generate a sample (unconditional prior)
|
| 78 |
+
image = pipeline().images[0]
|
| 79 |
+
image.save("generated_ct_slice.png")
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
Note: For actual CT reconstruction, this prior is typically used with data-consistency guidance as described in the paper.
|
| 83 |
+
|
| 84 |
+
---
|
| 85 |
|
| 86 |
+
## Citation
|
| 87 |
+
|
| 88 |
+
```bibtex
|
| 89 |
+
@inproceedings{
|
| 90 |
+
shi2026dmct,
|
| 91 |
+
title={{DM}4{CT}: Benchmarking Diffusion Models for Computed Tomography Reconstruction},
|
| 92 |
+
author={Shi, Jiayang and Pelt, Dani{\in}l M and Batenburg, K Joost},
|
| 93 |
+
booktitle={The Fourteenth International Conference on Learning Representations},
|
| 94 |
+
year={2026},
|
| 95 |
+
url={https://openreview.net/forum?id=YE5scJekg5}
|
| 96 |
+
}
|
| 97 |
+
```
|