Collections
Discover the best community collections!
Collections including paper arxiv:2501.08332
-
Cosmos World Foundation Model Platform for Physical AI
Paper • 2501.03575 • Published • 66 -
Phi-4 Technical Report
Paper • 2412.08905 • Published • 106 -
MiniMax-01: Scaling Foundation Models with Lightning Attention
Paper • 2501.08313 • Published • 268 -
MangaNinja: Line Art Colorization with Precise Reference Following
Paper • 2501.08332 • Published • 55
-
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models
Paper • 2407.15642 • Published • 11 -
ashawkey/LGM
Text-to-3D • Updated • 117 -
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
Paper • 2407.19548 • Published • 25 -
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
Paper • 2408.12601 • Published • 29
-
Magic Insert: Style-Aware Drag-and-Drop
Paper • 2407.02489 • Published • 20 -
ZePo: Zero-Shot Portrait Stylization with Faster Sampling
Paper • 2408.05492 • Published • 7 -
CSGO: Content-Style Composition in Text-to-Image Generation
Paper • 2408.16766 • Published • 18 -
Style-Friendly SNR Sampler for Style-Driven Generation
Paper • 2411.14793 • Published • 36
-
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
Paper • 2311.17049 • Published • 1 -
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Paper • 2405.04434 • Published • 14 -
A Study of Autoregressive Decoders for Multi-Tasking in Computer Vision
Paper • 2303.17376 • Published -
Sigmoid Loss for Language Image Pre-Training
Paper • 2303.15343 • Published • 6
-
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 22 -
PALO: A Polyglot Large Multimodal Model for 5B People
Paper • 2402.14818 • Published • 23 -
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Paper • 2403.09611 • Published • 126 -
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Paper • 2404.06512 • Published • 30