Back to Home

Introduction to Loss Functions: How Large Models Score Themselves with Cross-Entropy

September 26, 2026 at 01:33 PMSource: RunByAI0 comment(s)TechGuide

Training a large model is essentially letting it keep guessing the next word and then adjusting itself based on whether the guess was right. The ruler that measures how good the guess is, is the loss function. For language models the most common ruler is cross-entropy.

1. What a loss function does

A loss function compresses the gap between a model prediction and the true answer into a single number. The training goal is simple: make that number as small as possible. The optimizer (such as AdamW) takes the loss, computes the gradient of every parameter through backpropagation, and nudges the parameters in the direction that reduces the loss. The loss is the only learning signal, so its design directly shapes what the model learns.

2. What the language model predicts

A large language model performs next-token prediction: given the preceding text, it outputs a score (logit) for every token in the vocabulary, then turns those scores into a probability distribution via softmax. The true next token is unique, while the model gives a distribution over the whole vocabulary. Training aims to make the probability of the true token as high as possible.

3. Cross-entropy: punishing a miss

Cross-entropy measures the gap between the model distribution and the true distribution. In next-token prediction the true distribution is 1 on the correct token and 0 elsewhere, so cross-entropy simplifies to the negative log of the probability the model assigns to the correct token.

This form has intuitive properties: if the model gives the correct token a probability near 1, the loss is near 0; if the probability is tiny, the loss grows fast, because the logarithm plunges as it approaches zero. In other words, cross-entropy heavily penalizes the case where there is a definite correct answer and the model barely bet on it, which is exactly the mistake we want to avoid.

4. A training detail: compute loss only where needed

The input of a language model is a span of text, and the model predicts the next token at every position. But we usually do not want it to guess the user question; we want it to learn to generate answers. So in supervised fine-tuning, a common practice is to compute loss only on the answer part, masking out the question and padding so they do not contribute to the loss. This loss masking keeps the model from spending effort learning to repeat the input.

5. Beyond cross-entropy

Cross-entropy cares about how accurately each position predicts its token; it does not directly optimize the quality of a whole answer, nor does it penalize repetition or broken logic in long text. So in the LLM pipeline, pretraining and supervised fine-tuning build the base with cross-entropy, and later methods such as RLHF and direct preference optimization further align the model using preference signals about overall answer quality. The two measure goals at different levels and cannot replace each other.

Summary

Cross-entropy looks like just a formula, but it turns the fuzzy task of predicting the next word into a differentiable, optimizable, massively computable numeric goal. It penalizes hesitation (low probability) and rewards decisiveness (high probability), and masking lets us focus learning on the output we truly want to teach. Understanding the loss function is understanding the starting point of a large model self-correction.

Reference: compiled from publicly released industry information.

large modelLarge Language Model (LLM)
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment