The basal ganglia and reward
How deep brain circuits choose actions, build habits and learn from reward, and why the dopamine signal inspired a branch of AI.
Intermediate · about 8 min · updated 2026-10-02 · awaiting clinical review
See it in 3D:
The striatum, globus pallidus, nucleus accumbens and the direct and indirect pathways; action selection and habits; dopamine reward prediction errors and their distributional code; addiction as a three-stage cycle; Parkinson's and Huntington's diseases; levodopa, deep brain stimulation and optogenetics; and the reinforcement-learning mathematics shared by dopamine neurons, deep Q-networks and AlphaGo.
Contents
Choosing, wanting and learning from reward
Deep in each hemisphere lie the basal ganglia: the caudate and putamen (together the striatum), the globus pallidus, the nucleus accumbens and connected nuclei in the midbrain. They help decide which action to take, they turn repeated actions into habits, and their dopamine supply signals whether things turned out better or worse than expected.[1,2,3]
That last discovery became one of the strongest bridges between neuroscience and AI. The dopamine signal matches the error term of temporal-difference reinforcement learning. Reinforcement learning has since taught a deep network to play 49 Atari games and, combined with other methods, helped AlphaGo defeat a professional Go player.[3,4,5]
When the system breaks, movement and motivation break with it: Parkinson's disease, Huntington's disease and addiction all involve the basal ganglia. This reading explains the circuits, the mathematics of reward, the diseases, and the treatments from levodopa to deep brain stimulation.[6,7,8]
What the basal ganglia are
The striatum receives input from the cortex and sends it out through two parallel routes. In the classical model, the direct pathway, whose neurons carry dopamine D1 receptors, facilitates movement; the indirect pathway, whose neurons carry D2 receptors and project first to the lateral globus pallidus, inhibits it.[1,9]
Dopamine reaches the striatum from neurons in the substantia nigra; their loss, which causes striatal dopamine deficiency, is a hallmark of Parkinson's disease.[6]
Key numbers
- People aged 65 or over affected by Parkinson's disease
- 2–3%[6]
- Huntington's disease families with an expanded CAG repeat in the 1993 gene study
- all 75[7]
- Mean improvement in motor score off medication with subthalamic stimulation at 6 months (UPDRS-III)
- 19.6 points[10]
- AlphaGo's winning rate against other Go programs
- 99.8%[5]
Why we need them
Selection. At any moment many actions compete for the body. Redgrave, Prescott and Gurney proposed that the basal ganglia are the vertebrate brain's solution to this selection problem, letting one action through while holding the others back.[11]
Habits. Routines and rituals become automatic through learning. Graybiel suggests that many habitual and repetitive behaviours, including those seen in neuropsychiatric illness and addiction, emerge from experience-dependent plasticity in basal ganglia circuits that influence thought as well as action.[2]
Learning from outcomes. Dopamine neurons signal changes or errors in the prediction of future rewards, the teaching signal needed to learn which actions pay off.[3]
How the circuit works
Testing the model. The direct/indirect model was long untested in behaving animals. Kravitz and colleagues used optogenetics to switch on each pathway in mice: activating indirect-pathway neurons produced a parkinsonian state, with freezing, slowness and fewer movement starts, while activating direct-pathway neurons increased movement and, in a mouse model of Parkinson's disease, completely rescued these deficits.[9]
Prediction errors. Dopamine neurons fire more than usual when a reward is better than expected, stay at baseline when it is exactly as predicted, and dip below baseline when an expected reward fails to arrive. As learning proceeds the response moves from the reward to the cue that predicts it, exactly as temporal-difference learning predicts.[3]
Distributions, not averages. Inspired by AI research on distributional reinforcement learning, Dabney and colleagues proposed that dopamine neurons represent the whole range of possible future rewards rather than a single average, and found evidence for this code in recordings from the mouse ventral tegmental area.[12]
Text version of the diagram
- Cortex: proposes actions. Leads to Striatum, D1 neurons; Striatum, D2 neurons.
- Striatum, D1 neurons: direct pathway. Leads to Basal ganglia output (go).
- Striatum, D2 neurons: indirect pathway. Leads to Lateral globus pallidus (no-go).
- Lateral globus pallidus: first stop of the indirect pathway. Leads to Basal ganglia output.
- Basal ganglia output: inhibits the thalamus. Leads to Thalamus.
- Thalamus: back to cortex. Leads to Cortex (loop).
- Substantia nigra: dopamine to the striatum. Leads to Striatum, D1 neurons (dopamine).
When learning turns into habit, and habit into addiction
With repetition, behaviours that were once goal-directed can become habitual and stereotyped, a change attributed to plasticity in basal ganglia circuits.[2]
Koob and Volkow describe addiction as a cycle of three stages: binge and intoxication, involving dopamine and opioid changes in the basal ganglia that build incentive salience and drug-seeking habits; withdrawal and negative affect, with reduced dopamine function and recruitment of stress systems in the extended amygdala; and preoccupation and craving, driven by disrupted input from the prefrontal cortex and insula.[8]
When the basal ganglia fail
Parkinson's disease is the second most common neurodegenerative disorder. Loss of dopamine neurons in the substantia nigra and inclusions of aggregated α-synuclein are its hallmarks; diagnosis rests on slowness of movement (bradykinesia) and other motor signs, but many non-motor symptoms add to disability.[6]
Huntington's disease is caused by an expanded, unstable CAG trinucleotide repeat in a gene on chromosome 4, inherited as a dominant trait. Albin and colleagues proposed that hyperkinetic disorders such as this, with excess abnormal movements, result from impairment of the striatal neurons that project to the lateral globus pallidus, the start of the indirect pathway.[1,7]
Addiction dysregulates the motivational circuits of the basal ganglia, extended amygdala and prefrontal cortex.[8]
The mathematics of reward
Reinforcement learning, the branch of AI that learns from reward, uses the same quantities that dopamine neurons appear to signal.[3,4]
What an agent tries to maximise: the sum of future rewards, each discounted by for every step into the future, so sooner rewards count more.
| Symbol | Meaning | Unit |
|---|---|---|
| reward at time t | — | |
| discount factor between 0 and 1 | — |
Positive when things go better than predicted, negative when worse, zero when exactly as expected: the pattern of firing seen in dopamine neurons.
| Symbol | Meaning | Unit |
|---|---|---|
| predicted return from situation s | — | |
| prediction error | — |
Learns the value of taking action in situation by nudging it towards the reward plus the best value available next. The deep Q-network approximated with a deep neural network fed with game pixels.
| Symbol | Meaning | Unit |
|---|---|---|
| estimated return of action a in situation s | — | |
| learning rate | — | |
| the next situation | — |
If different dopamine cells weigh good and bad surprises differently (optimistic cells with large , pessimistic ones with large ), their predictions settle at different points of the reward distribution, so together they encode the whole distribution, as Dabney and colleagues proposed.
| Symbol | Meaning | Unit |
|---|---|---|
| the value predicted by cell (or channel) i | — | |
| learning rates for positive and negative errors | — | |
| that cell's prediction error | — |
Treatments, tools and AI
Replacing dopamine. Cotzias and colleagues showed in 1967 that high doses of the dopamine precursor given by mouth could markedly improve parkinsonism; treatment is still anchored on pharmacological substitution of striatal dopamine.[6,13]
Stimulation. Deep brain stimulation, for example of the subthalamic nucleus, is used when levodopa-related motor complications become intractable; as a surgical tool it can also record pathological brain activity directly.[6,10,14]
Light-controlled neurons. Optogenetics, which makes chosen neurons respond to light, let Kravitz and colleagues test the direct/indirect model directly in behaving mice.[9]
Reward-learning machines. A deep Q-network learned 49 Atari games from pixels and score, and AlphaGo combined networks trained by supervised and reinforcement learning with tree search to beat the European Go champion 5 games to 0.[4,5]
| Brain | Reinforcement learning |
|---|---|
| Dopamine burst for a better-than-expected reward | Positive temporal-difference error |
| Response moves from reward to predictive cue | Value learned for earlier states |
| Optimistic and pessimistic dopamine neurons | Distributional reinforcement learning |
Milestones
From levodopa to distributional dopamine
- 1967Oral levodopa transforms the treatment of parkinsonism.[13]
- 1989The direct/indirect pathway model of basal ganglia disorders.[1]
- 1993The Huntington's disease gene and its CAG repeat are found.[7]
- 1997Dopamine is linked to reward prediction errors.[3]
- 1999The basal ganglia are proposed as a solution to the selection problem.[11]
- 2006A randomised trial supports subthalamic deep brain stimulation for Parkinson's disease.[10]
- 2008Habits and rituals are linked to basal ganglia plasticity.[2]
- 2010Optogenetics tests the direct and indirect pathways in mice.[9]
- 2015A deep Q-network reaches human-level play on Atari games.[4]
- 2016AlphaGo defeats a professional Go player; addiction is mapped as a three-stage cycle.[5,8]
- 2020Evidence for a distributional reward code in dopamine neurons.[12]
Frontiers
Ideas now flow from AI back to the brain: distributional reinforcement learning was developed for artificial agents before it was tested, and supported, in dopamine neurons.[12]
Experimental Parkinson's therapies aim to restore striatal dopamine with gene-based and cell-based approaches, alongside work on the molecular causes, from α-synuclein handling to mitochondria and inflammation.[6]
Check yourself
Check yourself
- Which structures make up the striatum?
Show answer
The caudate and the putamen.
- In the classical model, what do the direct and indirect pathways do?
Show answer
The direct (D1) pathway facilitates movement; the indirect (D2) pathway inhibits it.
- What happened when Kravitz and colleagues activated indirect-pathway neurons?
Show answer
Mice became parkinsonian: more freezing, slowness and fewer movement starts.
- What do dopamine neurons signal?
Show answer
Errors in the prediction of reward: more firing when better than expected, less when worse.
- What causes Huntington's disease?
Show answer
An expanded, unstable CAG repeat in a gene on chromosome 4, inherited dominantly.
- What did the DBS trial find?
Show answer
Subthalamic stimulation improved quality of life and motor symptoms more than medication, with more serious adverse events.
- What is distributional reinforcement learning?
Show answer
Learning the whole distribution of possible rewards rather than only the average, which dopamine neurons may also do.
Glossary[1,3,4,6,9,10]
- Basal ganglia
- Deep nuclei including the striatum and globus pallidus that help select actions and form habits.
- Striatum
- The caudate and putamen, the main input stage of the basal ganglia.
- Direct pathway
- The striatal route, via D1 neurons, that facilitates movement.
- Indirect pathway
- The striatal route, via D2 neurons and the lateral globus pallidus, that inhibits movement.
- Dopamine
- A neurotransmitter whose release signals reward prediction errors.
- Substantia nigra
- A midbrain nucleus whose dopamine neurons supply the striatum.
- Bradykinesia
- Slowness of movement, a cardinal sign of Parkinson's disease.
- Deep brain stimulation
- Electrical stimulation through implanted electrodes in deep brain targets.
- Optogenetics
- Controlling genetically chosen neurons with light.
- Reinforcement learning
- Learning to act from rewards and punishments.
References
- Albin RL, Young AB, Penney JB. The functional anatomy of basal ganglia disorders. Trends in Neurosciences 1989;12(10):366-375. doi:10.1016/0166-2236(89)90074-X
- Graybiel AM. Habits, rituals, and the evaluative brain. Annual Review of Neuroscience 2008;31:359-387. doi:10.1146/annurev.neuro.29.051605.112851
- Schultz W, Dayan P, Montague PR. A neural substrate of prediction and reward. Science 1997;275(5306):1593-1599. doi:10.1126/science.275.5306.1593
- Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, et al.. Human-level control through deep reinforcement learning. Nature 2015;518(7540):529-533. doi:10.1038/nature14236
- Silver D, Huang A, Maddison CJ, Guez A, Sifre L, van den Driessche G, et al.. Mastering the game of Go with deep neural networks and tree search. Nature 2016;529(7587):484-489. doi:10.1038/nature16961
- Poewe W, Seppi K, Tanner CM, Halliday GM, Brundin P, Volkmann J, et al.. Parkinson disease. Nature Reviews Disease Primers 2017;3:17013. doi:10.1038/nrdp.2017.13
- MacDonald ME (Huntington's Disease Collaborative Research Group). A novel gene containing a trinucleotide repeat that is expanded and unstable on Huntington's disease chromosomes. Cell 1993;72(6):971-983. doi:10.1016/0092-8674(93)90585-E
- Koob GF, Volkow ND. Neurobiology of addiction: a neurocircuitry analysis. The Lancet Psychiatry 2016;3(8):760-773. doi:10.1016/S2215-0366(16)00104-8
- Kravitz AV, Freeze BS, Parker PRL, Kay K, Thwin MT, Deisseroth K, Kreitzer AC. Regulation of parkinsonian motor behaviours by optogenetic control of basal ganglia circuitry. Nature 2010;466(7306):622-626. doi:10.1038/nature09159
- Deuschl G, Schade-Brittinger C, Krack P, Volkmann J, Schäfer H, Bötzel K, et al.. A randomized trial of deep-brain stimulation for Parkinson's disease. New England Journal of Medicine 2006;355(9):896-908. doi:10.1056/NEJMoa060281
- Redgrave P, Prescott TJ, Gurney K. The basal ganglia: a vertebrate solution to the selection problem?. Neuroscience 1999;89(4):1009-1023. doi:10.1016/S0306-4522(98)00319-4
- Dabney W, Kurth-Nelson Z, Uchida N, Starkweather CK, Hassabis D, Munos R, Botvinick M. A distributional code for value in dopamine-based reinforcement learning. Nature 2020;577(7792):671-675. doi:10.1038/s41586-019-1924-6
- Cotzias GC, Van Woert MH, Schiffer LM. Aromatic amino acids and modification of parkinsonism. New England Journal of Medicine 1967;276(7):374-379. doi:10.1056/NEJM196702162760703
- Lozano AM, Lipsman N, Bergman H, Brown P, Chabardes S, Chang JW, et al.. Deep brain stimulation: current challenges and future directions. Nature Reviews Neurology 2019;15(3):148-160. doi:10.1038/s41582-018-0128-2
Related readings
- Neural networks, biological and artificial
How brains and machines learn from experience, the maths both share, and why AI and neuroscience keep borrowing from each other.
Intermediate
- The motor system
How the brain turns intention into movement, from the redrawn motor homunculus to the laws of reaching and the technology that restores it.
Intermediate
- Brain–computer interfaces
How implants like BrainGate and Neuralink read intention from motor cortex, the maths of decoding, and the race to restore speech and movement.
Intermediate
- The limbic system, cingulate and insula
How the brain's inner rim and the hidden insula sense the body, shape feelings, register pain and conflict, and how stimulators and decoders now target mood.
Intermediate
- The frontal lobe and executive control
How the front of the brain holds goals in mind, controls thought and action, weighs the future, and what Phineas Gage, lobotomy and AI reveal about it.
Intermediate
Template anatomy for education. Not patient-specific. Not for clinical decision-making.