Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Paper • 2501.04001 • Published 17 days ago • 41
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token Paper • 2501.03895 • Published 17 days ago • 48
The GAN is dead; long live the GAN! A Modern GAN Baseline Paper • 2501.05441 • Published 15 days ago • 82
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Paper • 2501.00958 • Published 22 days ago • 97