Skip to content

fix: score think_format_reward when <think> is prefilled in the prompt - #7141

Open
behroozazarkhalili wants to merge 5 commits into
huggingface:mainfrom
behroozazarkhalili:reopen/6995
Open

behroozazarkhalili wants to merge 5 commits into
huggingface:mainfrom
behroozazarkhalili:reopen/6995

Conversation

@behroozazarkhalili

@behroozazarkhalili behroozazarkhalili commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

think_format_reward rejected valid completions when the chat template had already placed the opening thinking tag in the prompt. I made the opening tag optional while keeping the closing tag required and rejecting nested opening tags. The reward test suite passes apart from skipped tests, and lint and format checks pass on both files.
Fixes #6966.


Note

Medium Risk
Changes RL reward criteria for format shaping; broader acceptance (any completion with </think>) may shift GRPO training signal versus the old strict opening-tag rule.

Overview
think_format_reward no longer requires a leading <think> tag so GRPO runs that decode only generated tokens (with the opening tag prefilled in the chat template) can earn 1.0 on well-formed reasoning. The matcher still requires </think>, forbids nested <think>, and uses an optional (?:<think>)? prefix on the existing regex.

Docs and doctest examples describe this GRPO/prefill behavior. Tests add prefilled-prompt-style completions as valid, drop “closing tag without opening” from the invalid set, and document that a bare or text-mention </think> also scores 1.0.

Reviewed by Cursor Bugbot for commit d66dbb5. Bugbot is set up for automated code reviews on this repo. Configure here.

@bot-ci-comment

bot-ci-comment Bot commented Sep 9, 2026

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

…the prompt

GRPO decodes only the generated tokens, so models whose chat template
prefills <think> into the prompt scored 0.0 on every well-formed
completion. The opening tag is now optional, </think> is still required,
and completions that start with <think> keep scoring 1.0. Based on the
approach proposed in huggingface#6995. Fixes huggingface#6966.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

think_format_reward always returns 0.0 for models whose chat template prefills <think>

1 participant