Beyond Reinforcement Learning: Sampling as Optimization for Reasoning Models
Imagine an LLM solving a math problem. It samples a full reasoning trace, you run a verifier, and the verifier returns a number: per...
Machine Learningenergy-based modelsimportance sampling