From: The Carrot and the Stick for Computers: How AI Learns by Doing
evidenceacademic

The computer must constantly choose between using moves it already knows work and trying completely new ones.

98% confidence

This is called the exploration-exploitation trade-off. Think of it like deciding where to eat dinner. Do you go to your favorite restaurant because you know the food is great? Or do you try a brand-new place that might be even better, but could also be terrible? The AI faces this exact dilemma at every step. It must balance sticking to safe, known rewards with exploring new paths to find even bigger payoffs.

Read the full exploration
What else is in this exploration
3 perspectives2 visualizations2 insights2 media resources4 rabbit holes
evidence
Reinforcement learning trains computers through a feedback system of rewards and penalties.
evidence
Unlike other AI methods, reinforcement learning does not need pre-labeled data to learn.
perspective
Biologists and psychologists note that this computer method is deeply rooted in animal behavior. ...
Sign up to unlock
Continue exploring
The Carrot and the Stick for Computers: How AI Learns by Doing
Evidence, perspectives, rabbit holes, and more