RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing Paper • 2608.26101 • Published 22 days ago • 4
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published Jul 26 • 79
Running on Zero Agents 79 SAM2 Image Predictor 🔥 79 Generate object masks from an image with point guidance