why the c-eval result is 76.8 for base model but only 38.9 for instruct model?

by xianf - opened 1 day ago

1 day ago

I use lm-eval to test the benchmark result, the base model performe well like the README said, but the instruct model is only 38.9 in this testset? What happened?

toothacher17

Moonshot AI org about 17 hours ago

Thanks for sharing. We actually evaluated C-Eval internally, and it performed reasonable. Could you please help check your log to see if the prompts are off or anything unexpected happened?

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment