Skip to content

Add PPO evaluation and greedy baseline reporting - #2

Open
chenmaike wants to merge 1 commit into
mainfrom
codex/implement-evaluate-function-for-mode-selection
Open

chenmaike wants to merge 1 commit into
mainfrom
codex/implement-evaluate-function-for-mode-selection

Conversation

@chenmaike

Copy link
Copy Markdown
Owner

Summary

  • refactor environment reward computation and add an evaluate() helper for PPO/greedy runs with optional temperature sampling
  • automatically run evaluations after training, log average WSR/DVP/constraint rate, and export comparison to JSON/CSV using the same seeded topology
  • overlay a greedy WSR baseline line on the observability plot for visual comparison

Testing

  • python -m compileall test.py

Codex Task

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant