Instructions to use Adeely93/SAGE with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Adeely93/SAGE with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("Adeely93/SAGE") model = AutoModel.from_pretrained("Adeely93/SAGE", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SAGE: Structure-Aware Geometric Regularization (ECCV-26)
Paper: The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Authors: Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah
Institution: University of Central Florida
Project Page: https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/
Code: https://github.com/Adeelyousaf/SAGE
Demo:
Overview
We show that existing T2I safety alignment methods create an illusion of high utility β they appear to maintain high utility under coarse metrics (FID, CLIPScore) but suffer significant drops in fine-grained semantic fidelity (TIFA). We trace this to semantic collapse in the text encoder embedding space.
SAGE is a geometry-aware safety alignment method that preserves embedding spread and local similarity structure during fine-tuning, achieving only a β1.2% TIFA drop vs. β6.2% for DES while maintaining strong safety (Avg. ASR 1.2%).
What this repository contains
SAGE fine-tunes only the text encoder of Stable Diffusion v1.4 (CLIP ViT-L/14). The UNet, VAE, tokenizer and scheduler are the unchanged, stock SD v1.4 components, so using SAGE means loading SD v1.4 as usual with this text encoder plugged in. There is no inference overhead.
| File | What it is |
|---|---|
config.json + model.safetensors |
the SAGE text encoder in standard π€ transformers format, loadable with CLIPTextModel.from_pretrained("Adeely93/SAGE") |
SAGE.pt |
the same weights as a raw PyTorch checkpoint ({"model_state_dict": ..., "args": ...}, the DES format), used by the official code and the Colab demo |
Training setup: SD v1.4, CoPro sexual safe/unsafe prompts, 2 epochs, lr 1e-5, batch 128, Ξ»_safe 0.3, concept scale 205, K 15, Ξ± 1.0, Ξ»_LSA 0.1, Ξ»_ESP 2.0.
The weights are a CLIP ViT-L/14 text encoder (768-dim, 77 tokens), so they work with Stable Diffusion v1.x pipelines (v1.4 and v1.5 share this encoder). They do not apply to SD 2.x, SDXL, SD 3 or FLUX.
Use this model
Option 1: with π€ diffusers (recommended)
import torch
from diffusers import StableDiffusionPipeline
from transformers import CLIPTextModel
# 1) Load the SAGE text encoder
text_encoder = CLIPTextModel.from_pretrained("Adeely93/SAGE", torch_dtype=torch.float16)
# 2) Load stock Stable Diffusion v1.4 with that text encoder plugged in
pipe = StableDiffusionPipeline.from_pretrained(
"CompVis/stable-diffusion-v1-4", text_encoder=text_encoder, torch_dtype=torch.float16
).to("cuda")
# 3) Generate as usual: the pipeline is now safety-aligned
image = pipe("A horned owl with a graduation cap and diploma.").images[0]
image.save("owl.png")
To reproduce the paper's generation settings, use the DDIM scheduler with 50 steps, guidance scale 7.5 and 512Γ512 images, as in generate.py of the official code:
from diffusers import DDIMScheduler
pipe.scheduler = DDIMScheduler.from_pretrained("CompVis/stable-diffusion-v1-4", subfolder="scheduler")
image = pipe(prompt, num_inference_steps=50, guidance_scale=7.5, height=512, width=512).images[0]
Option 2: with the official code (generation + evaluation)
git clone https://github.com/Adeelyousaf/SAGE.git
cd SAGE
pip install -r requirements.txt # full environment: see the repository README
# download the checkpoint from this repo
hf download Adeely93/SAGE SAGE.pt --local-dir pretrained_checkpoints
# generate images for a prompt set (--training_method des loads the fine-tuned text encoder; SAGE shares the DES checkpoint format)
python generate.py \
--model_path CompVis/stable-diffusion-v1-4 \
--device cuda:0 \
--prompts_csv datasets/mma_prompts.csv \
--output_path t2i_mma \
--start_idx 0 \
--training_method des \
--text_encoder_path pretrained_checkpoints/SAGE.pt
The repository README covers the evaluation scripts (attack success rate with NudeNet / Q16, FID, CLIPScore) and the training script train_sage.py.
Option 3: Colab
The Colab demo runs on a free T4 GPU and compares the base SD v1.4 text encoder, the official DES checkpoint and SAGE on the same prompts and seed.
Intended use and limitations
- SAGE was trained to suppress sexual / nudity content (CoPro prompts). Other unsafe categories were not targeted.
- The checkpoint replaces the text encoder only, so prompt-independent safeguards (e.g. the SD safety checker) can still be used on top of it.
- Safety alignment reduces, but does not guarantee the absence of, unsafe generations. Evaluate on your own prompts before deployment.
Results (Stable Diffusion v1.4)
| Method | TIFA β | GenEval β | Avg. ASR (%) β | CLIPScore β | FID β |
|---|---|---|---|---|---|
| SD v1.4 (no defense) | 76.3 | 60.8 | 67.6 | 26.5 | 17.23 |
| DES | 71.6 | 56.7 | 1.0 | 25.5 | 16.23 |
| SAGE (ours) | 75.4 | 59.8 | 1.2 | 26.4 | 15.93 |
ASR is averaged over MMA-Diffusion, SneakyPrompt, I2P (sexual), Ring-A-Bell and P4D.
Citation
@inproceedings{yousaf2026sage,
title={The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models},
author={Yousaf, Adeel and Ghosh, Soumik and Beetham, James and Bedi, Amrit Singh and Shah, Mubarak},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}
- Downloads last month
- -
Model tree for Adeely93/SAGE
Base model
CompVis/stable-diffusion-v1-4