YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
CIS 6270 Course Code
Use this repository alongside the lecture notes for CIS 6270, Discrete Generative Models, at the University of Pennsylvania. We organized each method into a single file so that you can follow its corruption process, loss, and sampler together as you work through the corresponding mathematics. Throughout the code, we use the notation from the notes and include comments on almost every line to help you connect each implementation to the derivation.
With the default settings, you can run the examples on a laptop CPU in seconds. For offline use, run lecture_2/flow_matching_mnist.py with --dataset toy; its default configuration downloads MNIST. In Lecture 3, you can run the scripts offline by default and select --dataset mnist or --dataset esm2 when you want to work with MNIST or ESM-2.
To get started, clone the repository, create a Python environment, and install the dependencies. You can then train a masked diffusion model on DNA, generate sequences, and check the worked examples against the lecture notes.
git clone https://huggingface.co/ChatterjeeLab/CIS6270
cd CIS6270
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements.txt
python lecture_4/mdlm.py # train masked diffusion on DNA and generate, about 20 seconds
python -m pytest tests -q # the notes' worked numbers, checked against the code
If you use Windows PowerShell, activate the environment with .venv\Scripts\Activate.ps1.
Lectures
| Lecture | Topics you can work through | Guide | |
|---|---|---|---|
| 1 | Mathematical Foundations | Every worked example in the chapter, printed with its chapter example number | lecture_1 |
| 2 | Flow Matching | Velocity fields and the flow map, the conditional path, the objective, MNIST | lecture_2 |
| 3 | Score Matching and Diffusion | The forward SDE, the three score objectives, DDPM, three samplers, guidance, latent diffusion on protein embeddings | lecture_3 |
| 4 | Discrete Diffusion | Jump processes, MDLM, uniform replacement, D3PM, SEDD, block diffusion, guidance, planners, search | lecture_4 |
| 5 | Discrete Flow Matching | The mixture path and its velocity, correctors, kinetic-optimal paths, Edit Flows, Dirichlet, Fisher and Gumbel geometries, guidance, rectification | lecture_5 |
| 6 | Flow Maps | The three axioms, the three residuals, MeanFlow, Shortcut, consistency, the discrete flow maps, stochastic, expanding and posterior maps | lecture_6 |
| 7 | Optimal Transport and Schrödinger Bridges | Kantorovich and duality, Sinkhorn, displacement interpolation, static and dynamic bridges, DSB, DSBM, SF2M, the discrete bridges, branched and entangled | lecture_7 |
Working through a method
Start with the file for the method you are studying and read it alongside the corresponding chapter. In each script, we follow the organization of lecture_2/flow_matching_mnist.py. Begin with the docstring to find the method and chapter sections, then work through the numbered # %% sections in lecture order. Run the script through main() to produce a report and inspect the results.
We keep each lecture directory independent of the others so that you can run and study each method within its own file. Refer to cis6270/ when you want to inspect the shared utilities used across the lectures.
| Directory | Contents |
|---|---|
cis6270/ |
Shared utilities for data, four small networks, the training loop, and sample metrics |
lecture_1/ through lecture_7/ |
One standalone, runnable file per method |
tests/ |
Checks of the chapters' worked numerical examples against the code |
Data
| Data | Used by | Download |
|---|---|---|
| Synthetic DNA, four letters, a planted ACGT motif, a GC label | Lectures 4 to 6 | no |
| Four 2-D Gaussians, two moons, a checkerboard | Lectures 2, 3, 6, 7 | no |
Four-token phrases such as red circle moves left |
Lecture 6 | no |
| MNIST | Lectures 2, 3, with --dataset mnist |
yes |
| ESM-2 residue embeddings | Lecture 3, with --dataset esm2 |
yes |
We use synthetic DNA and two simple guidance objectives so that you can train a model quickly and inspect how guidance changes the generated sequences. Because we specify the sequence properties in advance, you can check whether the samples recover the intended structure by examining the output directly.
Shared flags
You can configure the scripts with the same flags throughout the repository, including --steps, --batch-size, --lr, --width, --length, --samples, --sample-steps, --seed, --device, --out, and --quiet. Use --out to save checkpoint.pt, config.json, report.json, losses.json, and samples.txt.
Start with the default settings to work through a short training run in well under a minute. Once you have followed the method, try --steps 1500 --width 128 to increase the training budget and network width, which usually improves the generated samples.
Tests
python -m pytest tests -q
Run the tests to compare the numerical results in the code with the worked examples in the notes. We check the posterior average of the reverse rates on one DNA base, the Sinkhorn iterates at epsilon = 2, the cocycle and Jacobian of the Gaussian flow, and the conditional velocity of the mixture path. You can use these checks as you work through the derivations or modify the implementations to explore their behavior.
License
You can use this code under the MIT license. When using ESM-2 weights, follow the terms of their original repository.
