1Cademy - Rescoring and Reranking for Inference-Time Alignment

Learn Before

Inference-Time LLM Alignment

Concept

Rescoring and Reranking for Inference-Time Alignment

Rescoring, also known as reranking, is an inference-time alignment technique that evaluates and prioritizes a model's generated outputs. This method uses a scoring system, often a reward model, to select the best output from multiple candidates. Reranking has a history of use in NLP tasks like machine translation and is typically applied when training complex models is prohibitively expensive, as it offers a low-cost way to incorporate their capabilities.

Updated 2026-05-03

Contributors are:

Who are from:

References

Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course

Learn After

Using Scoring Systems for Inference-Time Rescoring
Best-of-N Sampling (BoN Sampling)
Use of Reranking to Explore Model and Search Errors
The Challenge of Candidate Diversity in Reranking Methods
A development team uses a large, pre-trained language model to generate summaries of news articles. To improve the factual accuracy of the final output, their system first generates five different summary candidates. Then, a separate, specialized scoring model evaluates each of the five summaries for factual consistency with the original article and selects the one with the highest score. Which statement best analyzes the trade-offs of this approach?
Improving Chatbot Responses on a Budget
A system is designed to improve the quality of its generated responses at inference time without altering the base model's parameters. It does this by producing several options and then choosing the best one. Arrange the following actions into the correct operational sequence.

Learn Before

Related

Learn After