2 3 22

Ayaan Sharif

Ayaan-Sharif

https://shariif.tech

AI & ML interests

NLP, LLM, TEXT, Languages

Recent Activity

liked a model 11 days ago

MiniMaxAI/MiniMax-VL-01

liked a dataset 14 days ago

DAMO-NLP-SG/multimodal_textbook

liked a model 19 days ago

cognitivecomputations/dolphin-2.9-llama3-8b

View all activity

Organizations

Ayaan-Sharif's activity

liked a model 11 days ago

MiniMaxAI/MiniMax-VL-01

Image-Text-to-Text • Updated about 14 hours ago • 1.84k • 220

liked a dataset 14 days ago

DAMO-NLP-SG/multimodal_textbook

Updated 15 days ago • 12.5k • 130

liked 2 models 19 days ago

cognitivecomputations/dolphin-2.9-llama3-8b

Text Generation • Updated May 20, 2024 • 141k • 430

cognitivecomputations/Dolphin3.0-Llama3.2-1B

Updated 20 days ago • 7.66k • 21

replied to sanchit-gandhi's post 24 days ago

what if we segment the audio first and then transcribe tho its some extra compute to throw in but imo it would resul tin better result !

liked 3 Spaces 25 days ago

Running

🚀

Ebook2audiobook V2.0 Beta

Added improvements, 1107+ languages supported

liked a model 29 days ago

huggyllama/llama-7b

Text Generation • Updated Jul 2, 2024 • 168k • 314

commented a paper about 1 month ago

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Paper • 2412.10302 • Published Dec 13, 2024 • 11 •

liked a model about 1 month ago

deepseek-ai/DeepSeek-V3-Base

Updated 2 days ago • 21.2k • 1.33k

upvoted a collection about 1 month ago

IndicConformer

Collection

A collection of ASR models for 22 scheduled languages of India • 22 items • Updated Oct 15, 2024 • 7

liked 2 Spaces about 1 month ago

Running

528

🌍

THUDM/cogvlm2-llama3-caption

Video-Text-to-Text • Updated 4 days ago • 192k • 81

Neurazum/Xbai-Epilepsy-1.0

Video-Text-to-Text • Updated Nov 11, 2024 • 2

reacted to vladbogo's post with 👍 about 1 month ago

Post

Panda-70M is a new large-scale video dataset comprising 70 million high-quality video clips, each paired with textual captions, designed to be used as pre-training for video understanding tasks.

Key Points:
* Automatic Caption Generation: Utilizes an automatic pipeline with multiple cross-modality teacher models to generate captions for video clips.
* Fine-tuned Caption Selection: Employs a fine-tuned retrieval model to select the most appropriate caption from multiple candidates for each video clip.
* Improved Performance: Pre-training on Panda-70M shows significant performance gains in video captioning, text-video retrieval, and text-driven video generation.

Paper: Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers (2402.19479)
Project page: https://snap-research.github.io/Panda-70M/
Code: https://github.com/snap-research/Panda-70M

Congrats to the authors @tschen , @aliaksandr-siarohin et al. for their work!