Hi, thank you for your work on harness-r1!
I find the work really interesting and would like to explore other training strategies. But i find out if i use the vanilla Qwen model the performance isn't so good. So i'm wondering if you could release the sft-only chechpoint?
Thanks.
Hi, thank you for your work on harness-r1!
I find the work really interesting and would like to explore other training strategies. But i find out if i use the vanilla Qwen model the performance isn't so good. So i'm wondering if you could release the sft-only chechpoint?
Thanks.