1Cademy - Best-of-N Sampling (BoN Sampling)

Learn Before

Rescoring and Reranking for Inference-Time Alignment

Activity (Process)

Best-of-N Sampling (BoN Sampling)

Best-of-N (BoN) sampling is a technique where a model generates multiple, or 'N', alternative outputs, and a reward model scores them to select the best one. While commonly used for inference-time alignment through reranking, the core mechanism of BoN sampling can also be adapted for training purposes, such as in rejection sampling.

Updated 2026-05-03

Contributors are:

Who are from:

References

Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course

Learn After

Input and Output Formulation in BoN Sampling
Generating N-Best Candidates in BoN Sampling
Reward Model Selection in BoN Sampling
Rejection Sampling for LLM Fine-Tuning
A company wants to improve the safety and helpfulness of its AI assistant without the high cost and time of retraining the entire base model. They propose a new system for handling user queries: for each query, the system will first generate 10 different potential responses. Then, a separate, fast-acting 'quality-scoring' model will evaluate all 10 responses based on pre-defined criteria. Finally, the system will present only the single response that received the highest score to the user. What
A system is designed to improve the quality of its generated text by producing multiple options and then picking the best one. Arrange the following steps of this process in the correct logical order.
Chatbot Response Quality Improvement

Learn Before

Related

Learn After