CLIP

Contrastive Language-Image Pretraining (CLIP) model pre-trained on 2.5 billion data points of CommonCrawl at resolution 224x224. It was introduced in the paper Learning Transferable Visual Models From Natural Language Supervision and further reproduced in the follow-up paper Demystifying CLIP Data. The weights were converted from the b32_fullcc2.5b.pt file presented in the original repository.

Downloads last month: 2

Safetensors

Model size

151M params

Tensor type

I64

F32

Inference Examples

Zero-Shot Image Classification

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Collection including cs-giung/clip-vit-base-patch32-fullcc2.5b

MetaCLIP (CommonCrawl-2.5B)

Collection

5 items • Updated Jul 7