Discussion about this post

User's avatar
Josh Green's avatar

Fascinating direction. Selfishly the thing I keep bumping into is that world-model-style agents want to run a lot of inference per decision, and once you are doing many model calls per step the cost stops being about the model and starts being about how cheaply you can serve it. That pushed me toward running the cheap stuff locally and only reaching for a big model at the decision points.

Emil Schmitz's avatar

Modeling future trajectories for RL reminds me of AlphaGo and MCMC. To what extent do you think its techniques are applicable here?

2 more comments...

No posts

Ready for more?