Lotus & Lotus-2: Visual Foundation Models for Dense Geometry Estimation
This repository provides consolidated, plug-and-play model weights for both Lotus-2 (FLUX.1-dev SOTA) and Lotus-1 (SD 2.1 Fast) for monocular depth and surface normal estimation, pre-configured for native integration with ComfyUI-Lotus-2.
π¦ Repository Contents
All safetensors weights in this repository are verified, standalone, and organized for direct automatic downloading or manual placement.
1. Lotus-2 Weights (FLUX.1-dev Monocular Geometry)
Developed by EnVision-Research (2026), built on FLUX.1-dev DiT using multi-stage LoRA and Local Continuity Modules (LCM):
| File Name | Size | Task | Description |
|---|---|---|---|
lotus-2_core_predictor_depth.safetensors |
~1.43 GB | Depth | LoRA adapter for core global depth prediction |
lotus-2_detail_sharpener_depth.safetensors |
~1.43 GB | Depth | LoRA adapter for high-frequency depth refinement |
lotus-2_lcm_depth.safetensors |
~39 kB | Depth | Local Continuity Module for depth smoothness |
lotus-2_core_predictor_normal.safetensors |
~2.87 GB | Normal | LoRA adapter for core surface normal prediction |
lotus-2_detail_sharpener_normal.safetensors |
~2.87 GB | Normal | LoRA adapter for high-frequency normal refinement |
lotus-2_lcm_normal.safetensors |
~39 kB | Normal | Local Continuity Module for normal consistency |
ae.safetensors |
~335 MB | VAE | FLUX.1-dev native autoencoder |
2. Lotus-1 Weights (SD 2.1 Fast UNet Engine)
Developed by EnVision-Research, lightweight standalone diffusion UNet models (~1.7 GB):
| File Name | Size | Task | Description |
|---|---|---|---|
lotus-depth-g-v2-1-disparity-fp16.safetensors |
~1.74 GB | Depth | Standalone Generative Depth (disparity, FP16) |
lotus-normal-g-v1-1-fp16.safetensors |
~1.74 GB | Normal | Standalone Generative Surface Normal (FP16) |
vae-ft-mse-840000-ema-pruned.safetensors |
~335 MB | VAE | Stable Diffusion 2.1 fine-tuned MSE Autoencoder |
π Usage in ComfyUI
These models are natively supported by the ComfyUI-Lotus-2 custom node suite under category π§ͺAILab/Geometry:
cd ComfyUI/custom_nodes
git clone https://github.com/1038lab/ComfyUI-Lotus-2.git
Node 1: Lotus 2 (FLUX) (SOTA Quality)
- Minimal Inputs: Requires only
imageandmodel(FLUX.1-dev UNet / Checkpoint). - Zero VRAM Waste: VAE and text conditioning are handled internallyβno need to connect DualCLIPLoader or load 10GB T5 models!
- Auto-Download: If LoRA or VAE files are not found locally, the node automatically downloads them from this repository (
1038lab/Lotus-2).
Node 2: Lotus (SD 2.1 Fast Engine)
- Standalone: All-in-one node without requiring an external base model loader.
- Fast: Generates high-quality geometry maps in seconds.
π Credits & Acknowledgments
We express our sincere gratitude and full credit to the original research teams and creators:
Lotus-2 (FLUX-based)
- Authors: Jing He1, Haodong Li1,2*, Mingzhi Sheng1*, Ying-Cong Chen1,3β
- Affiliations: 1HKUST (GZ), 2HKUST, 3HKUST Dept. of CSE
- Original Code & Weights: EnVision-Research/Lotus-2 / jingheya/Lotus-2
Lotus-1 (SD-based)
- Authors: Jing He, Haodong Li, Weicai Ye, Wang Zhao, Suping Chen, Torsten Sattler, Guofeng Zhang, Shenghua Gao, Ying-Cong Chen
- Original Code & Weights: EnVision-Research/Lotus / jingheya/lotus-*
Base Models
- FLUX.1-dev & FLUX VAE: Black Forest Labs
- Stable Diffusion 2.1 & VAE: Stability AI
π Citation
If you find Lotus or Lotus-2 helpful in your research or projects, please cite the official papers:
@article{he2025lotus2,
title = {Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model},
author = {He, Jing and Li, Haodong and Sheng, Mingzhi and Chen, Ying-Cong},
journal = {arXiv preprint arXiv:2512.01030},
year = {2025}
}
@article{he2024lotus,
title = {Lotus: Diffusion-based Visual Foundation Model for Dense Geometry Estimation},
author = {He, Jing and Li, Haodong and Ye, Weicai and Zhao, Wang and Chen, Suping and Sattler, Torsten and Zhang, Guofeng and Gao, Shenghua and Chen, Ying-Cong},
journal = {arXiv preprint arXiv:2409.08272},
year = {2024}
}
