Preference Leakage: A Contamination Problem in LLM-as-a-judge Paper • 2502.01534 • Published 25 days ago • 38
Lamarck-14B Qwen 2.5 and relatives Collection Lamarck's public releases, plus significant related merges and finetunes • 6 items • Updated 10 days ago • 1
Preference Datasets for DPO Collection This collection contains a list of curated preference datasets for DPO fine-tuning for intent alignment of LLMs • 7 items • Updated Dec 11, 2024 • 40