top of page
1.png
3.png
3.png
1.png
1.png
3.png

Item Year

1989

Year Added

2026

Location

Cambridge, UK

Source

Watkins

Screenshot 2026-03-24 at 5.08.45 PM.png
calendar (1).png
calendar (2).png
circle-upload-512_edited.png
images_edited.png
Screenshot 2026-03-24 at 5.08.45 PM.png

Added By

Added By

Item Year

Item Year

Year Added

Item Year

Source

Item Year

Location

Item Year

LIVE LOUD

Q-learning is a reinforcement learning algorithm created by Christopher Watkins in his 1989 doctoral thesis, Learning from Delayed Rewards. It learns by trial and error, keeping a running estimate of how good it is to take each action in each situation, and updating those estimates from the rewards it receives plus the best expected future reward. Because it learns from experience without needing a model of the world, it became the theoretical backbone of reinforcement learning, the branch of AI that learns by doing rather than by being told. Its ideas run beneath the modern…




Help us find the original Alexa video!

brain.png

Q-learning

logo (1)_edited.jpg
ADDED BY:
Curators' Team

Item Year

1989

Year Added

2026

Location

Cambridge, UK

Source

Watkins

Screenshot 2026-03-24 at 5.08.45 PM.png
calendar (1).png
calendar (2).png
circle-upload-512_edited.png
images_edited.png
Screenshot 2026-03-24 at 5.08.45 PM.png

Added By

Added By

Item Year

Item Year

Year Added

Item Year

Source

Item Year

Location

Item Year

LIVE LOUD

Q-learning is a reinforcement learning algorithm created by Christopher Watkins in his 1989 doctoral thesis, Learning from Delayed Rewards. It learns by trial and error, keeping a running estimate of how good it is to take each action in each situation, and updating those estimates from the rewards it receives plus the best expected future reward. Because it learns from experience without needing a model of the world, it became the theoretical backbone of reinforcement learning, the branch of AI that learns by doing rather than by being told. Its ideas run beneath the modern era of AI. Deep Q-Networks, DeepMind's 2015 program that learned to play Atari games from raw pixels, was Q-learning upgraded with neural networks, and the same family of techniques powered the reinforcement learning systems that followed, including the lineage that led to AlphaGo. Q-learning is the foundation of reinforcement learning, the ground AlphaGo and everything that followed stand on. The ecosystem grew around it too. Q-learning gave reinforcement learning its core idea, and OpenAI Gym in 2016 gave it a standardized training ground, the ecosystem behind the modern RL boom.




Help us find the original Alexa video!


brain.png

Added By

Added By

Item Year

Added By

Year Added

Added By

Location

Added By

Source

Added By

Q-learning

Q-learning is a reinforcement learning algorithm created by Christopher Watkins in his 1989 doctoral thesis, Learning from Delayed Rewards. It learns by trial and error, keeping a running estimate of how good it is to take each action in each situation, and updating those estimates from the rewards it receives plus the best expected future reward. Because it learns from experience without needing a model of the world, it became the theoretical backbone of reinforcement learning, the branch of AI that learns by doing rather than by being told. Its…

logo (1)_edited.jpg
ADDED BY:

Curators' Team

Item Year

Item Year

calendar (1).png
calendar (2).png
circle-upload-512_edited.png
images_edited.png

ITEM YEAR

1989

Year Added

2026

Source

Watkins

Location

Cambridge, UK

brain.png
SOURCE
Watkins
calendar (1).png
calendar (1).png
circle-upload-512_edited.png
images_edited.png
Screenshot 2026-03-24 at 5.08.45 PM.png
ITEM YEAR
1989
YEAR ADDED
LOCATION
Cambridge, UK
2026

Q-learning

logo (1)_edited.jpg
ADDED BY:
Curators' Team

Q-learning is a reinforcement learning algorithm created by Christopher Watkins in his 1989 doctoral thesis, Learning from Delayed Rewards. It learns by trial and error, keeping a running estimate of how good it is to take each action in each situation, and updating those estimates from the rewards it receives plus the best expected future reward. Because it learns from experience without needing a model of the world, it became the theoretical backbone of reinforcement learning, the branch of AI that learns by doing rather than by being told. Its ideas run beneath the modern era of AI. Deep Q-Networks, DeepMind's 2015 program that learned to play…




Help us find the original Alexa video!




Help us find the original Alexa video!



Nothing to see yet. :(

Add media related to this item.


Nothing to see yet. :(

Add media related to this item.


Nothing to see yet. :(

Add media related to this item.


Nothing to see yet. :(

Add media related to this item.

Keep up with history
in the making.

Plus, get invited to curate, including telling your own stories, and receive new product alerts, and priority collab opportunities.

Keep up with history
in the making.

Plus, get invited to curate, including telling
your own stories, and receive new product alerts and priority collab opportunities. 
Customize preferences.
bottom of page