Multiple Choice

In a standard self-attention mechanism, an input vector is transformed into three separate vectors (Query, Key, and Value) using three distinct, learned weight matrices. Imagine a modified self-attention layer where these three weight matrices are constrained to be identical. What would be the most direct consequence of this change?

0

1

Updated 2025-10-10

Contributors are:

Who are from:

Tags

Ch.5 Inference - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Analysis in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science