leaderboard-pr-bot's picture
Adding Evaluation Results
330dc22 verified
|
raw
history blame
4.39 kB
metadata
library_name: transformers
tags:
  - mergekit
  - merge
base_model:
  - Tsunami-th/Tsunami-0.5x-7B-Instruct
  - Qwen/Qwen2.5-Math-7B
model-index:
  - name: 2_PRYMMAL-ECE-7B-SLERP-V3
    results:
      - task:
          type: text-generation
          name: Text Generation
        dataset:
          name: IFEval (0-Shot)
          type: HuggingFaceH4/ifeval
          args:
            num_few_shot: 0
        metrics:
          - type: inst_level_strict_acc and prompt_level_strict_acc
            value: 22.35
            name: strict accuracy
        source:
          url: >-
            https://huggingface.co./spaces/open-llm-leaderboard/open_llm_leaderboard?query=Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
          name: Open LLM Leaderboard
      - task:
          type: text-generation
          name: Text Generation
        dataset:
          name: BBH (3-Shot)
          type: BBH
          args:
            num_few_shot: 3
        metrics:
          - type: acc_norm
            value: 10.61
            name: normalized accuracy
        source:
          url: >-
            https://huggingface.co./spaces/open-llm-leaderboard/open_llm_leaderboard?query=Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
          name: Open LLM Leaderboard
      - task:
          type: text-generation
          name: Text Generation
        dataset:
          name: MATH Lvl 5 (4-Shot)
          type: hendrycks/competition_math
          args:
            num_few_shot: 4
        metrics:
          - type: exact_match
            value: 0
            name: exact match
        source:
          url: >-
            https://huggingface.co./spaces/open-llm-leaderboard/open_llm_leaderboard?query=Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
          name: Open LLM Leaderboard
      - task:
          type: text-generation
          name: Text Generation
        dataset:
          name: GPQA (0-shot)
          type: Idavidrein/gpqa
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            value: 0.89
            name: acc_norm
        source:
          url: >-
            https://huggingface.co./spaces/open-llm-leaderboard/open_llm_leaderboard?query=Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
          name: Open LLM Leaderboard
      - task:
          type: text-generation
          name: Text Generation
        dataset:
          name: MuSR (0-shot)
          type: TAUR-Lab/MuSR
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            value: 9.74
            name: acc_norm
        source:
          url: >-
            https://huggingface.co./spaces/open-llm-leaderboard/open_llm_leaderboard?query=Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
          name: Open LLM Leaderboard
      - task:
          type: text-generation
          name: Text Generation
        dataset:
          name: MMLU-PRO (5-shot)
          type: TIGER-Lab/MMLU-Pro
          config: main
          split: test
          args:
            num_few_shot: 5
        metrics:
          - type: acc
            value: 9.08
            name: accuracy
        source:
          url: >-
            https://huggingface.co./spaces/open-llm-leaderboard/open_llm_leaderboard?query=Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
          name: Open LLM Leaderboard

merged_model

This is a merge of pre-trained language models created using mergekit.

Merge Details

Merge Method

This model was merged using the SLERP merge method.

Models Merged

The following models were included in the merge:

Configuration

The following YAML configuration was used to produce this model:

slices:
  - sources:
      - model: Tsunami-th/Tsunami-0.5x-7B-Instruct
        layer_range: [0, 28]
      - model: Qwen/Qwen2.5-Math-7B
        layer_range: [0, 28]    
merge_method: slerp
base_model: Tsunami-th/Tsunami-0.5x-7B-Instruct
parameters:
  t:
    - filter: self_attn
      value: [0, 0.1, 0.2, 0.3, 0.4]  # Influence réduite pour Qwen sur les couches d'attention
    - filter: mlp
      value: [0, 0.15, 0.3, 0.45, 0.6]  # Influence légèrement accrue pour les couches MLP
    - value: 0.2  # Ajustement général pour favoriser Tsunami sur l'ensemble
dtype: bfloat16

Open LLM Leaderboard Evaluation Results

Detailed results can be found here

Metric Value
Avg. 8.78
IFEval (0-Shot) 22.35
BBH (3-Shot) 10.61
MATH Lvl 5 (4-Shot) 0.00
GPQA (0-shot) 0.89
MuSR (0-shot) 9.74
MMLU-PRO (5-shot) 9.08