SAGE: Structure-Aware Geometric Regularization (ECCV-26)

Paper: The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Authors: Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah
Institution: University of Central Florida
Project Page: https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/
Code: https://github.com/Adeelyousaf/SAGE
Demo: Open In Colab


Overview

We show that existing T2I safety alignment methods create an illusion of high utility β€” they appear to maintain high utility under coarse metrics (FID, CLIPScore) but suffer significant drops in fine-grained semantic fidelity (TIFA). We trace this to semantic collapse in the text encoder embedding space.

SAGE is a geometry-aware safety alignment method that preserves embedding spread and local similarity structure during fine-tuning, achieving only a βˆ’1.2% TIFA drop vs. βˆ’6.2% for DES while maintaining strong safety (Avg. ASR 1.2%).


What this repository contains

SAGE fine-tunes only the text encoder of Stable Diffusion v1.4 (CLIP ViT-L/14). The UNet, VAE, tokenizer and scheduler are the unchanged, stock SD v1.4 components, so using SAGE means loading SD v1.4 as usual with this text encoder plugged in. There is no inference overhead.

File What it is
config.json + model.safetensors the SAGE text encoder in standard πŸ€— transformers format, loadable with CLIPTextModel.from_pretrained("Adeely93/SAGE")
SAGE.pt the same weights as a raw PyTorch checkpoint ({"model_state_dict": ..., "args": ...}, the DES format), used by the official code and the Colab demo

Training setup: SD v1.4, CoPro sexual safe/unsafe prompts, 2 epochs, lr 1e-5, batch 128, Ξ»_safe 0.3, concept scale 205, K 15, Ξ± 1.0, Ξ»_LSA 0.1, Ξ»_ESP 2.0.

The weights are a CLIP ViT-L/14 text encoder (768-dim, 77 tokens), so they work with Stable Diffusion v1.x pipelines (v1.4 and v1.5 share this encoder). They do not apply to SD 2.x, SDXL, SD 3 or FLUX.


Use this model

Option 1: with πŸ€— diffusers (recommended)

import torch
from diffusers import StableDiffusionPipeline
from transformers import CLIPTextModel

# 1) Load the SAGE text encoder
text_encoder = CLIPTextModel.from_pretrained("Adeely93/SAGE", torch_dtype=torch.float16)

# 2) Load stock Stable Diffusion v1.4 with that text encoder plugged in
pipe = StableDiffusionPipeline.from_pretrained(
    "CompVis/stable-diffusion-v1-4", text_encoder=text_encoder, torch_dtype=torch.float16
).to("cuda")

# 3) Generate as usual: the pipeline is now safety-aligned
image = pipe("A horned owl with a graduation cap and diploma.").images[0]
image.save("owl.png")

To reproduce the paper's generation settings, use the DDIM scheduler with 50 steps, guidance scale 7.5 and 512Γ—512 images, as in generate.py of the official code:

from diffusers import DDIMScheduler
pipe.scheduler = DDIMScheduler.from_pretrained("CompVis/stable-diffusion-v1-4", subfolder="scheduler")
image = pipe(prompt, num_inference_steps=50, guidance_scale=7.5, height=512, width=512).images[0]

Option 2: with the official code (generation + evaluation)

git clone https://github.com/Adeelyousaf/SAGE.git
cd SAGE
pip install -r requirements.txt        # full environment: see the repository README

# download the checkpoint from this repo
hf download Adeely93/SAGE SAGE.pt --local-dir pretrained_checkpoints

# generate images for a prompt set (--training_method des loads the fine-tuned text encoder; SAGE shares the DES checkpoint format)
python generate.py \
    --model_path CompVis/stable-diffusion-v1-4 \
    --device cuda:0 \
    --prompts_csv datasets/mma_prompts.csv \
    --output_path t2i_mma \
    --start_idx 0 \
    --training_method des \
    --text_encoder_path pretrained_checkpoints/SAGE.pt

The repository README covers the evaluation scripts (attack success rate with NudeNet / Q16, FID, CLIPScore) and the training script train_sage.py.

Option 3: Colab

The Colab demo runs on a free T4 GPU and compares the base SD v1.4 text encoder, the official DES checkpoint and SAGE on the same prompts and seed.


Intended use and limitations

  • SAGE was trained to suppress sexual / nudity content (CoPro prompts). Other unsafe categories were not targeted.
  • The checkpoint replaces the text encoder only, so prompt-independent safeguards (e.g. the SD safety checker) can still be used on top of it.
  • Safety alignment reduces, but does not guarantee the absence of, unsafe generations. Evaluate on your own prompts before deployment.

Results (Stable Diffusion v1.4)

Method TIFA ↑ GenEval ↑ Avg. ASR (%) ↓ CLIPScore ↑ FID ↓
SD v1.4 (no defense) 76.3 60.8 67.6 26.5 17.23
DES 71.6 56.7 1.0 25.5 16.23
SAGE (ours) 75.4 59.8 1.2 26.4 15.93

ASR is averaged over MMA-Diffusion, SneakyPrompt, I2P (sexual), Ring-A-Bell and P4D.

Citation

@inproceedings{yousaf2026sage,
  title={The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models},
  author={Yousaf, Adeel and Ghosh, Soumik and Beetham, James and Bedi, Amrit Singh and Shah, Mubarak},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Adeely93/SAGE

Finetuned
(874)
this model

Paper for Adeely93/SAGE