πŸ–ΌοΈ POCKET-Image-Zimage β€” 4-bit (NF4) Z-Image for on-device

A 4-bit (NF4) quantized build of Z-Image (Apache-2.0), packaged by VIDRAFT for low-VRAM, on-device image generation β€” part of the POCKET line.

  • πŸ“¦ ~6 GB on disk (transformer + text encoder in NF4, VAE in fp16)
  • ⚑ Runs from ~8.6 GB VRAM (β‰ˆ4.5 GB with CPU offload) β€” vs 23.3 GB for bf16
  • 🎯 ~2.7–5Γ— smaller footprint, quality on par with the bf16 base

Usage

import torch
from diffusers import ZImagePipeline   # or ZImageImg2ImgPipeline / ZImageInpaintPipeline

pipe = ZImagePipeline.from_pretrained(
    "FINAL-Bench/POCKET-Image-Zimage", torch_dtype=torch.bfloat16
).to("cuda")
img = pipe("a serene mountain lake at sunrise, photorealistic", num_inference_steps=20).images[0]
img.save("out.png")

Requires bitsandbytes (CUDA). Measured reload + generate peak: 10.9 GB VRAM. For Apple Silicon / CPU, an optimum-quanto int8 build (13.4 GB) is the portable option.

🎨 The full POCKET-Image system

This repo hosts the quantized base model only. The headline character-perfect Korean & multilingual text feature is delivered by the POCKET-Image pipeline, not by these weights alone. Try the full system here:

Base model: Tongyi-MAI/Z-Image (Apache-2.0) Β· Quantization: bitsandbytes NF4 Β· By VIDRAFT.


🧩 The POCKET Family β€” On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

πŸ“š Full POCKET collection

Downloads last month
11
Safetensors
Model size
3B params
Tensor type
F32
Β·
BF16
Β·
U8
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FINAL-Bench/POCKET-Image-Zimage

Finetuned
(65)
this model

Space using FINAL-Bench/POCKET-Image-Zimage 1

Collection including FINAL-Bench/POCKET-Image-Zimage