When you train a model with supervised learning, you first need a large set of input–correct-answer pairs. But many problems have no standard answer—playing Go, making a robot walk, or recommending a user's next video. This is where reinforcement learning (RL) comes in.
RL is built around an agent–environment loop. An agent observes a state, chooses an action, and the environment moves to a new state and returns a reward signal. The agent's goal is to find a policy that maximizes cumulative reward over time through repeated trial and error.
The central tension is exploration versus exploitation. Exploitation means sticking with actions known to score well; exploration means trying new actions in the hope of finding something better. Pure exploitation traps the agent in a suboptimal solution; pure exploration never yields a stable policy. Nearly every RL algorithm balances the two.
Rewards are often delayed. In chess, individual moves give no reward—only the final result matters. To weigh how the current action affects the future, RL introduces a discount factor γ that discounts future rewards, producing the "return." The closer γ is to 1, the more patient the agent and the more it values the long term.
There are broadly two families of algorithms. Value-based methods learn how good it is to take an action in a state (the Q value) and pick actions accordingly—Q-learning is the classic example. Policy-based methods optimize the policy directly—policy gradients being the main example. After 2013, using deep neural networks for value estimation gave rise to deep RL: DeepMind's DQN reached human level on Atari games, and AlphaGo defeated the world Go champion in 2016.
RL also connects directly to large language models. RLHF (reinforcement learning from human feedback) asks humans to rank model outputs, trains a reward model on those rankings, and then uses RL to optimize the language model toward human preferences. The success of ChatGPT made RLHF a standard tool for aligning large models.
In practice, RL already powers game AI, robot control, autonomous-driving decisions, recommender systems, data-center energy savings, and power-grid scheduling optimization.
That said, RL is hard to train: sparse rewards, low sample efficiency, and instability are common obstacles. It typically needs vast interaction data and compute—which is why simulators and high-fidelity training environments matter so much in engineering practice.
In one sentence: RL frees AI from needing standard answers, letting it discover decision-making ability in complex environments through an act–feedback–adjust loop.
This article is compiled from publicly available industry information and classic textbooks.