Cohere For AI

community

https://cohere.for.ai/

CohereForAI

for-ai

Activity Feed Request to join this org

AI & ML interests

None defined yet.

Recent Activity

MarziehFadaee published a model about 22 hours ago

CohereForAI/c4ai-command-r7b-arabic-02-2025

alexrs new activity 1 day ago

CohereForAI/c4ai-command-r7b-arabic-02-2025:Add model card

kyle-cohere updated a model 1 day ago

CohereForAI/c4ai-command-r7b-arabic-02-2025

View all activity

Articles

A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality

Oct 24, 2024

• 59

Putting RL back in RLHF

Jun 12, 2024

• 77

CohereForAI's activity

davanstrien

posted an update about 5 hours ago

Post

132

📊 Introducing "Hugging Face Dataset Spotlight" 📊

I'm excited to share the first episode of our AI-generated podcast series focusing on nice datasets from the Hugging Face Hub!

This first episode explores mathematical reasoning datasets:

- SynthLabsAI/Big-Math-RL-Verified: Over 250,000 rigorously verified problems spanning multiple difficulty levels and mathematical domains
- open-r1/OpenR1-Math-220k: 220,000 math problems with multiple reasoning traces, verified for accuracy using Math Verify and Llama-3.3-70B models.
- facebook/natural_reasoning: 1.1 million general reasoning questions carefully deduplicated and decontaminated from existing benchmarks, showing superior scaling effects when training models like Llama3.1-8B-Instruct.

Plus a bonus segment on bespokelabs/bespoke-manim!

https://www.youtube.com/watch?v=-TgmRq45tW4

MarziehFadaee

published a model about 22 hours ago

CohereForAI/c4ai-command-r7b-arabic-02-2025

Text Generation • Updated 1 day ago • 80 • 18

davanstrien

posted an update 1 day ago

Post

1744

Quick POC: Turn a Hugging Face dataset card into a short podcast introducing the dataset using all open models.

I think I'm the only weirdo who would enjoy listening to something like this though 😅

Here is an example for eth-nlped/stepverify

1 reply

alexrs

in CohereForAI/c4ai-command-r7b-arabic-02-2025 1 day ago

Add model card

#1 opened 1 day ago by

kyle-cohere

updated a model 1 day ago

CohereForAI/c4ai-command-r7b-arabic-02-2025

Text Generation • Updated 1 day ago • 80 • 18

kyle-cohere

in CohereForAI/c4ai-command-r7b-arabic-02-2025 1 day ago

Add model card

#1 opened 1 day ago by

kyle-cohere

alexrs

updated a model 1 day ago

CohereForAI/c4ai-command-r7b-arabic-02-2025

Text Generation • Updated 1 day ago • 80 • 18

davanstrien

posted an update 8 days ago

Post

2521

Hacked together a way to log trl GRPO training completions to a 🤗 dataset repo. This allows you to:

- Track rewards from multiple reward functions
- Treat the completion and rewards from training as a "proper" dataset and do EDA
- Share results for open science

The implementation is super hacky, but I'm curious if people would find this useful.

To push completions to the Hub, you just need two extra parameters:

log_completions=True
log_completions_hub_repo='your-username/repo-name'

Example dataset: davanstrien/test-logs
Colab: https://colab.research.google.com/drive/1wzBFPVthRYYTp-mEYlznLg_e_0Za1M3g

alexrs

updated a model 8 days ago

CohereForAI/c4ai-command-r7b-12-2024

Text Generation • Updated 8 days ago • 8.97k • 362

Brittawnya

updated a Space 10 days ago

README

🏃

davanstrien

posted an update 12 days ago

Post

2206

Dataset descriptions for trending Hugging Face datasets? Powered by a Smol model davanstrien/Smol-Hub-tldr

davanstrien

posted an update 14 days ago

Post

1876

How do you make 1M+ Hugging Face models & datasets more discoverable?

davanstrien/Smol-Hub-tldr!

I fine-tuned HuggingFaceTB/SmolLM2-360M to generate one-line summaries from a model or dataset README.

Its own self-description?
"A model for generating concise summaries of model & dataset cards from the Hugging Face Hub"

The goal? Make it easier to find the right models and datasets for your specific needs. It's already powering a semantic search for datasets Space.

It's still a WIP but thanks to @loubnabnl , @anton-l , @eliebak et al, for cooking such a nice base model for fine-tuning small, efficient models for specific domains and tasks. 🙏

davanstrien

posted an update 15 days ago

Post

1342

Made some significant updates to my 🤗 semantic datasets search app. If you love falling into a wiki black hole, you might like this...

https://huggingface.co./spaces/librarian-bots/huggingface-datasets-semantic-search

maximevoisin1212

in CohereForAI/c4ai-command-r7b-12-2024 19 days ago

Opportunities `is_error` flag

#13 opened about 1 month ago by

DiTy

clefourrier

authored a paper 22 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 24 days ago • 195

davanstrien

posted an update about 1 month ago

Post

1823

Why choose between strong LLM reasoning and efficient models?

Use DeepSeek to generate high-quality training data, then distil that knowledge into ModernBERT answerdotai/ModernBERT-base for fast, efficient classification.

Blog post: https://danielvanstrien.xyz/posts/2025/deepseek/distil-deepseek-modernbert.html

alexrs

in CohereForAI/c4ai-command-r7b-12-2024 about 1 month ago

Update tokenizer_config.json

#14 opened about 1 month ago by

alexrs

davanstrien

posted an update about 1 month ago

Post

1914

Updated the ColPali Query Generator Space davanstrien/ColPali-Query-Generator to use Qwen/Qwen2.5-VL-7B-Instruct.

Given an input image, it generates several queries along with explanations to justify them. This approach can generate synthetic data for fine-tuning ColPali models.

davanstrien

posted an update about 1 month ago

Post

2032

🌍 Big step for multilingual AI data!

The Hugging Face community has rated educational content in languages spoken by 1.6 billion people! New additions:
• Japanese
• Italian
• Old High German

Learn more and contribute: https://huggingface.co./blog/davanstrien/fineweb2-community

These ratings can help enhance training data for major world languages.