1Cademy - An activation function is designed to scale its input value by the probability that a randomly drawn value from a standard normal distribution (mean=0, variance=1) is less than or equal to that input. How does this functions output for a small negative input (e.g., -0.1) compare to the output of a function that simply sets all negative inputs to zero?

Learn Before

Gaussian Error Linear Unit (GELU)

Multiple Choice

An activation function is designed to scale its input value by the probability that a randomly drawn value from a standard normal distribution (mean=0, variance=1) is less than or equal to that input. How does this function's output for a small negative input (e.g., -0.1) compare to the output of a function that simply sets all negative inputs to zero?

Updated 2025-10-02

Contributors are:

Who are from:

Learn Before

Related