Grpo Rlvr Training

by wshobson46891e7e60daNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.

Instructions onlyAI & Agents

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/grpo-memory.md3.5 KBtext/markdown
references/reward-functions.md11.5 KBtext/markdown
SKILL.md7.6 KBtext/markdown

Source and attribution

Source:wshobson/agentsinplugins/llm-finetuning/skills/grpo-rlvr-trainingat commit46891e7

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal