Discussion about this post

User's avatar
David Parker's avatar

The missing piece in most RL explainers is that the reward model is where the values actually live. I have seen the same shape in my own agent stack: the gate that says no is only as good as the criteria it refuses on. Trace the reward back far enough and you find the person who wrote it.

gutsun's avatar

the figure in PPO part: (High-level depiction of an actor-critic setup) seems to have a little mistake: where critic model is generating G_t and reward model is generatin V(s_t)?

2 more comments...

No posts

Ready for more?