Gemma RLAIF
Collection
9 items
•
Updated
This model is a fine-tuned version of lewtun/gemma-7b-sft-full-deita-10k-v0 on the argilla/dpo-mix-7k dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen |
---|---|---|---|---|---|---|---|---|---|---|---|
0.1371 | 1.9 | 100 | 0.5052 | -2.9796 | -5.1227 | 0.7708 | 2.1431 | -502.7766 | -483.2003 | 96.3785 | 102.3770 |
Base model
google/gemma-7b