Post-Pretraining Data Formatting Bug in a T5-Style Text-to-Text Service
You are rolling out an internal, single-model NLP service based on a T5-style text-to-text approach. The model is an encoder–decoder network that was pre-trained with span-based denoising using sentinel tokens (e.g., <extra_id_0>, <extra_id_1>) and then fine-tuned on multiple tasks using textual task prefixes.
After deployment, two issues appear:
- For classification-style requests (e.g., "sentiment:
"), the model often outputs strings that look like "<extra_id_0> positive" or "<extra_id_0> negative" instead of just "positive"/"negative". - For generation-style requests (e.g., "summarize:
"), the model sometimes inserts sentinel tokens into the summary.
A teammate proposes a quick fix: "Strip any <extra_id_*> tokens from the model output at inference time and ship." Another teammate argues the root cause is in how inputs/targets are being constructed for fine-tuning and that the fix should be in the text-to-text formatting and training pipeline.
As the reviewer, analyze which teammate is more correct and justify your decision by explaining (a) how span-based denoising trains an encoder–decoder model to use sentinel tokens, (b) how the text-to-text task prefix + target formatting should differ between denoising pretraining and downstream fine-tuning, and (c) one concrete change you would make to the fine-tuning data (source/target strings) to prevent sentinel-token leakage without relying on post-processing.

0
1
Tags
Ch.1 Pre-training - Foundations of Large Language Models
Foundations of Large Language Models
Foundations of Large Language Models Course
Computing Sciences
Data Science
Related
T5 Sample Format
Critique of the T5 Text-to-Text Approach
A developer is using a unified model that frames all natural language processing problems as a text-to-text task. The goal is to build a feature that extracts the main subjects from a sentence. Given the input text 'Instruction: Identify the subjects. Text: The cat and the dog played in the yard.', which of the following outputs best demonstrates the model's core operational principle?
A key principle of a unified text-to-text model is its ability to handle diverse natural language processing tasks by framing them as a transformation from an input text to an output text. Match each traditional NLP task with the most appropriate input/output text pair that represents how this type of model would process it.
Designing a Unified Text-to-Text Model and Pretraining Objective for Multiple NLP Features
Diagnosing a T5-Style Model That Ignores Task Prefixes After Span-Denoising Pretraining
Choosing Between Span-Denoising Pretraining and Task-Specific Fine-Tuning in a T5-Style Text-to-Text System
Selecting an Architecture and Pretraining Objective for a Unified Internal NLP Service
Post-Pretraining Data Formatting Bug in a T5-Style Text-to-Text Service
Root-Cause Analysis of a T5-Style Model Producing Fluent but Unfaithful Outputs
Your team is building a single internal T5-style t...
Your company wants one internal model to support m...
Your team is pretraining an internal T5-style mode...
Your team is pretraining an internal T5-style enco...
Training Process for Text-to-Text Models
T5 Model as a Text-to-Text System
A developer is using a single, unified model that processes all tasks by mapping an input text string to an output text string. The developer wants to perform a summarization task on the following article: 'Jupiter is the fifth planet from the Sun and the largest in the Solar System. It is a gas giant with a mass more than two and a half times that of all the other planets in the Solar System combined.' Which of the following input/output pairs correctly frames this task for such a model?
Evaluating a Unified NLP Approach
A key advantage of the text-to-text framework is its ability to represent a wide variety of Natural Language Processing (NLP) tasks using a single, unified format. Match each traditional NLP task with its corresponding text-to-text formulation.
Your team is pretraining an internal T5-style enco...
Your company wants one internal model to support m...
Your team is pretraining an internal T5-style mode...
Your team is building a single internal T5-style t...
Diagnosing a T5-Style Model That Ignores Task Prefixes After Span-Denoising Pretraining
Choosing Between Span-Denoising Pretraining and Task-Specific Fine-Tuning in a T5-Style Text-to-Text System
Designing a Unified Text-to-Text Model and Pretraining Objective for Multiple NLP Features
Root-Cause Analysis of a T5-Style Model Producing Fluent but Unfaithful Outputs
Selecting an Architecture and Pretraining Objective for a Unified Internal NLP Service
Post-Pretraining Data Formatting Bug in a T5-Style Text-to-Text Service
An encoder-decoder model is being trained with a span-based denoising objective. The encoder is given the following corrupted input text: 'To learn about the solar system, we first study <mask_0> and then move on to <mask_1> planets.' The original, uncorrupted text for the masked spans is '<mask_0>' = 'the Sun' and '<mask_1>' = 'the other'. What should the target output sequence for the decoder be in this training step?
Analysis of Denoising Training Objectives
Debugging a Span-Based Denoising Training Pipeline
Your team is pretraining an internal T5-style enco...
Your company wants one internal model to support m...
Your team is pretraining an internal T5-style mode...
Your team is building a single internal T5-style t...
Diagnosing a T5-Style Model That Ignores Task Prefixes After Span-Denoising Pretraining
Choosing Between Span-Denoising Pretraining and Task-Specific Fine-Tuning in a T5-Style Text-to-Text System
Designing a Unified Text-to-Text Model and Pretraining Objective for Multiple NLP Features
Root-Cause Analysis of a T5-Style Model Producing Fluent but Unfaithful Outputs
Selecting an Architecture and Pretraining Objective for a Unified Internal NLP Service
Post-Pretraining Data Formatting Bug in a T5-Style Text-to-Text Service
Encoder
Decoder
Context vector
Encoder-Decoder with Transformers
Multi-lingual Pre-training for Encoder-Decoder Models
Mathematical Formulation of an Encoder-Decoder Model
Seq2seq Models for Text Generation
Auto-Regressive Decoding in Machine Translation