Computing 2028 Draft - AI Paradigms
Reinforcement Learning & State Tables
Authority: The Museum of AICore GCSE syllabus reference
1. Core Concept Description
Explore how agents learn actions in dynamic environments. Instead of programmed rules, agents use rewards (+10) and penalties (-10) to map states to actions.
Syllabus Key takeaway:
Trial and Error. Agents update their Q-value tables by exploring, leaking target rewards backward through the state map.
2. Interactive Museum Labs
Cabinet Laboratory
Launch the visual coding maze or attention simulator to test this concept.
Historical Timelines
Read the detailed curatorial dossiers on the chronological milestones.
2. Archival References
- Timeline dossier: Watkins' Q-Learning (1989) — https://museumofai.org.uk/nodes/node-q-learning-1989
- Interactive simulator: QLEARNING Lab — https://museumofai.org.uk/arcade/qlearning