Last reviewed on October 2, 2026.
How could something as rich as thought arise from billions of neurons that each do something very simple? Connectionism is cognitive science’s most developed answer. Instead of describing the mind as a computer running rules over symbols, it models cognition as patterns of activity spreading through networks of simple, interconnected units that learn from experience. This guide explains what connectionism is, how connectionist models work, where the approach came from, the classic models that made its case, the long debate with symbolic theories, and how it relates to today’s deep learning and large language models.
Short answer: what is connectionism in cognitive science?
Connectionism is an approach in cognitive science that explains mental processes — perception, memory, language, learning — as the activity of artificial neural networks: large numbers of simple processing units connected by weighted links. Knowledge is stored in the strengths of the connections rather than as explicit rules, representations are typically distributed across many units, and learning consists of gradually adjusting the weights in response to experience. It is often contrasted with the symbolic (“classical”) approach, which models thinking as rule-based manipulation of symbols.
Core Principles of Connectionism
Connectionist models differ widely in detail, but most share four ingredients, set out clearly in the 1986 Parallel Distributed Processing (PDP) volumes:
1. Simple processing units
A network is made of units (or nodes), loosely inspired by neurons. Each unit has an activation level. It receives input from other units, sums it, and passes the result through an activation function — for example, a threshold or a smooth S-shaped curve — to determine its own output. No single unit is intelligent; any intelligence lies in the pattern of interactions.
2. Weighted connections
Units are linked by connections with weights. A positive weight is excitatory, a negative weight inhibitory, and the magnitude sets how strongly one unit influences another. In a connectionist model, the weights are the network’s long-term knowledge. There is no separate memory store where facts or rules are written down.
3. Distributed representations
In many connectionist models, a concept such as “dog” is not stored in one dedicated unit (a “grandmother cell”) but as a pattern of activation across many units, and each unit takes part in representing many concepts. This has useful consequences:
- Automatic generalisation: similar things get similar patterns, so what is learned about one item transfers to similar ones.
- Content-addressable memory: a partial or noisy cue can reactivate the whole pattern, much as a smell can bring back a whole memory.
- Graceful degradation: damaging some units or connections degrades performance gradually instead of deleting specific facts — similar to what is often seen after brain damage.
Not every connectionist model is distributed. Localist models, such as the interactive activation model of word reading and the TRACE model of speech perception, use one unit per letter, phoneme or word, but still rely on parallel interaction between units.
4. Learning rules
Networks learn by changing their weights. The main kinds of learning rule are:
- Hebbian learning: strengthen the connection between units that are active together (“cells that fire together, wire together”). It needs no teacher and finds correlations in the input.
- Error-correcting (supervised) learning: compare the network’s output with a target, and adjust weights to reduce the error. The delta rule does this for a single layer; backpropagation extends it to networks with hidden layers by passing error signals backwards through the network.
- Unsupervised and self-organising learning: networks such as competitive learning and Kohonen’s self-organising maps discover structure in the input without targets.
- Reinforcement learning: adjust weights based on reward signals rather than correct answers.
Because knowledge emerges from many small weight changes across many examples, connectionist learning is gradual and statistical. Regularities in the environment become regularities in the weights. Key terms such as neural network, backpropagation and perceptron are defined in our glossary.
A Short History of Connectionism
- 1943 – McCulloch and Pitts: Warren McCulloch and Walter Pitts published “A logical calculus of the ideas immanent in nervous activity”, modelling neurons as binary threshold units and showing that networks of them could compute logical functions.
- 1949 – Hebb: In The Organization of Behavior, psychologist Donald Hebb proposed that connections between neurons strengthen when they are repeatedly active together, and that learning forms “cell assemblies”. Hebbian learning remains a cornerstone of both neuroscience and connectionism.
- 1958 – Rosenblatt’s perceptron: Frank Rosenblatt introduced the perceptron, a network that learned to classify patterns by adjusting its weights, and the perceptron convergence theorem showed that its learning rule would find a correct set of weights whenever one existed. Bernard Widrow and Ted Hoff’s ADALINE and the closely related delta rule followed in 1960.
- 1969 – Minsky and Papert: In Perceptrons, Marvin Minsky and Seymour Papert proved that single-layer perceptrons cannot compute some simple functions, such as exclusive-or (XOR) or whether a figure is connected. Multilayer networks could in principle solve these problems, but nobody yet had a good way to train them. Funding and interest shifted to symbolic AI.
- 1970s–early 1980s – quiet progress: Work continued on associative memory (James Anderson, Teuvo Kohonen), adaptive resonance theory (Stephen Grossberg) and, in 1982, John Hopfield’s recurrent networks that store memories as stable states. The Boltzmann machine (Ackley, Hinton and Sejnowski, 1985) followed.
- 1986 – the PDP revolution: David Rumelhart, James McClelland and the PDP Research Group published the two-volume Parallel Distributed Processing, which made connectionism a major force in cognitive science. The same year, Rumelhart, Geoffrey Hinton and Ronald Williams popularised backpropagation for training multilayer networks in Nature (the underlying method had earlier roots, including work by Paul Werbos in 1974).
- 1988–1990s – debate and refinement: Critiques from Fodor and Pylyshyn and from Pinker and Prince sharpened the debate with symbolic theories, while connectionist models of reading, memory, development and brain damage multiplied.
- 2006–today – deep learning: Better training methods, large datasets and GPUs made very deep networks practical. AlexNet’s 2012 success in image recognition and the 2017 transformer architecture led to today’s AI systems. Hinton, Yann LeCun and Yoshua Bengio received the 2018 Turing Award, and Hopfield and Hinton shared the 2024 Nobel Prize in Physics for foundational work on neural networks.
Connectionism grew up alongside the symbolic tradition that launched the cognitive revolution; the wider story is told in our history of cognitive science.
Classic Connectionist Models
The past-tense model (Rumelhart & McClelland, 1986)
English-speaking children often go through a U-shaped pattern: they first say “went” correctly, then overregularise to “goed”, then return to “went”. The standard explanation was that children learn a rule (“add -ed”) plus a list of exceptions. Rumelhart and McClelland trained a simple network to map verb stems to past-tense forms and found that it produced both regular and irregular forms, and overregularisation errors, without containing any explicit rule. Steven Pinker and Alan Prince (1988) published a detailed critique, arguing among other things that the U-shape depended on how the training data were presented. The “past-tense debate” became a long-running test case for whether language needs rules, and later connectionist models addressed many of the criticisms.
NETtalk (Sejnowski & Rosenberg, 1987)
NETtalk learned to convert written English text into phonemes that could drive a speech synthesiser. As training progressed, its output moved from babble to increasingly intelligible speech, and analysis of its hidden units showed that it had discovered distinctions such as vowels versus consonants without being told about them. It became a famous demonstration that a network could learn the messy spelling-to-sound rules of English from examples.
Elman’s simple recurrent network (1990)
Language unfolds in time, but early networks processed fixed-size inputs. Jeffrey Elman’s simple recurrent network added “context units” that feed a copy of the previous hidden state back into the network, giving it a form of memory. Trained only to predict the next word in simple sentences, the network’s internal representations came to group words into nouns and verbs, and animate and inanimate nouns, purely from distributional patterns. Elman’s paper “Finding structure in time” is a direct ancestor of today’s next-word-prediction language models.
Other influential models
- Interactive activation model (McClelland & Rumelhart, 1981): explained the word superiority effect — letters are identified better inside words than alone — through feedback from word units to letter units.
- Reading models (Seidenberg & McClelland, 1989; Plaut et al., 1996): single networks that read both regular and irregular words, and when damaged mimicked patterns of acquired dyslexia. See our article on language processing in the brain for the competing dual-route account.
- Complementary learning systems (McClelland, McNaughton & O’Reilly, 1995): a response to catastrophic interference — the finding that networks trained on new information can abruptly overwrite old knowledge (McCloskey & Cohen, 1989). The theory proposes that the hippocampus learns new episodes quickly and gradually teaches the neocortex, offering an explanation of memory consolidation. See memory and learning.
Symbolic vs Connectionist Approaches: The Great Debate
The classical or symbolic view, associated with Allen Newell, Herbert Simon, Jerry Fodor and Zenon Pylyshyn, holds that cognition is computation over structured symbolic representations, like a program manipulating sentences in a “language of thought”. Connectionism challenged this by showing that networks without explicit rules could produce rule-like behaviour.
Fodor and Pylyshyn’s systematicity argument (1988)
In “Connectionism and cognitive architecture: A critical analysis”, Fodor and Pylyshyn argued that human thought is systematic and compositional: anyone who can think “John loves Mary” can also think “Mary loves John”, because thoughts are built from reusable parts combined by structure-sensitive rules. Classical symbol systems guarantee this. Connectionist networks, they argued, do not — a network could learn one sentence without the other. So either connectionism fails as a theory of cognition, or it succeeds only by implementing a classical symbol system, in which case it is a theory of the neural hardware rather than of the mind.
Responses
Connectionists replied in several ways. Paul Smolensky argued that networks are best described at a subsymbolic level and developed tensor product representations that encode structured combinations in vectors. Others argued that human systematicity is less perfect than claimed, and that it can emerge from learning. Recent work has partially vindicated both sides: standard networks often fail tests of systematic generalisation (Lake & Baroni, 2018), yet networks trained with suitable meta-learning procedures can generalise compositionally in human-like ways (Lake & Baroni, 2023). The question has shifted from “can networks be systematic?” to “what training and architecture make them so?”
A useful way to frame the debate is through David Marr’s levels of analysis: symbolic models often describe what is computed, while connectionist models propose how it might be implemented and learned. Many researchers now favour hybrid or neurosymbolic models that combine both.
Symbolic and connectionist approaches compared
| Feature | Symbolic (classical) approach | Connectionist approach |
|---|---|---|
| Basic metaphor | Mind as a computer running programs | Mind as a brain-like network |
| Representation | Discrete symbols and structured expressions | Patterns of activation, usually distributed across many units |
| Where knowledge is stored | Explicit rules and facts in memory | Connection weights |
| Processing | Mostly serial, rule-based manipulation | Massively parallel spread of activation |
| Learning | Often hand-coded or added as new rules | Gradual weight adjustment from examples |
| Handling noise and damage | Brittle: missing a rule can cause failure | Graceful degradation and tolerance of noisy input |
| Strengths | Logic, planning, compositional structure, explicit reasoning | Pattern recognition, similarity-based generalisation, learning statistics from data |
| Weaknesses | Learning, flexibility, grounding symbols in perception | Systematicity, transparency, data hunger, catastrophic interference |
| Classic examples | General Problem Solver, SOAR, ACT-R, expert systems | Perceptron, PDP models, NETtalk, Elman networks |
Connectionism, Deep Learning and Large Language Models
Modern deep learning is the direct descendant of connectionism: it uses the same units, weights and gradient-based learning, scaled up to networks with many layers and billions of parameters trained on huge datasets. Convolutional networks for vision, recurrent networks and, since 2017, transformers for language all grow out of this tradition.
There are important differences of aim, however. Classic connectionist models in cognitive science were deliberately small and built to explain specific findings — a U-shaped learning curve, a pattern of dyslexia, a reaction-time effect. Deep learning in AI is mainly engineered for performance. Large language models are trained on far more text than any human encounters and are not designed to match human development. Even so, the two have converged in interesting ways: deep convolutional networks trained on object recognition predict activity in the primate visual system (Yamins et al., 2014), and language-model activations predict human brain responses to sentences. Cognitive scientists now use these systems both as tools and as candidate models, while testing where they diverge from human minds — for instance in data efficiency, systematic generalisation and grounding in perception and action. For more, see our article on how AI is shaping cognitive science.
Connectionism also connects to other contemporary frameworks. Predictive processing describes the brain as a hierarchical network that learns by reducing prediction error, and embodied cognition pushes networks to be grounded in sensory and motor experience.
Strengths and Limitations of Connectionism
Strengths
- Learning from experience: networks acquire knowledge from examples instead of relying on hand-written rules, which suits domains where rules are hard to state.
- Graded, similarity-based behaviour: they naturally produce typicality effects, partial knowledge and generalisation to new cases, as people do.
- Robustness: graceful degradation and tolerance of noisy input resemble the brain’s resilience.
- Development and damage: they can model how abilities change over learning and how specific patterns of impairment arise after brain injury.
- Neural plausibility: the architecture is closer to the brain than rule systems are, linking cognitive theories to neuroscience.
Limitations and criticisms
- Systematicity and compositionality: networks often struggle to recombine known parts in new ways as reliably as humans do.
- Transparency: it can be hard to say what a trained network has learned or why it gives an answer (the “black box” problem), which also makes models hard to evaluate as explanations.
- Data and training demands: many networks need far more examples than people do, and some, like backpropagation itself, are of debated biological plausibility.
- Catastrophic interference: learning new information can overwrite old knowledge unless special mechanisms are added.
- Too much flexibility: with enough units and the right training regime a network can fit almost anything, so critics ask what a successful simulation actually proves.
Why Connectionism Still Matters for Cognitive Science
Connectionism changed how cognitive scientists think about knowledge, learning and representation. It showed that rule-like behaviour can emerge from statistical learning, that memory can be reconstructive and content-addressable, and that cognitive models can be linked to the brain. Its central question — how much of the mind can be learned by general-purpose networks from experience, and how much needs built-in structure — is more relevant than ever in the age of large language models. Students who want to work in this area typically combine computational modelling and experimental methods; our guide to careers in cognitive science covers where that leads.
Frequently Asked Questions
What is connectionism in simple terms?
Connectionism is the idea that the mind can be explained as the activity of networks of simple, neuron-like units connected by weighted links. Knowledge is stored in the strengths of the connections, and learning means gradually adjusting those strengths through experience rather than following explicit rules.
Who founded connectionism?
There is no single founder. Early foundations came from Warren McCulloch and Walter Pitts (1943), Donald Hebb (1949) and Frank Rosenblatt's perceptron (1958). Modern connectionism in cognitive science was established by David Rumelhart, James McClelland and the PDP Research Group with the Parallel Distributed Processing volumes in 1986.
What is the difference between connectionism and symbolic AI?
Symbolic AI models thinking as rule-based manipulation of discrete symbols, with knowledge stored as explicit rules and facts. Connectionism models thinking as parallel activity in networks, with knowledge stored in connection weights and learned from examples. Symbolic systems excel at logic and structure; connectionist systems excel at learning and pattern recognition.
Is deep learning the same as connectionism?
Deep learning is the direct descendant of connectionism and uses the same principles of units, weights and learning by gradient descent, but on a much larger scale. The main difference is the goal: connectionist models in cognitive science aim to explain human behaviour, while deep learning is mainly engineered for performance on tasks.
What is the main criticism of connectionism?
The best-known criticism, from Jerry Fodor and Zenon Pylyshyn in 1988, is that human thought is systematic and compositional, whereas neural networks do not guarantee this. Other criticisms concern the lack of transparency of trained networks, their need for large amounts of data and catastrophic interference, where new learning overwrites old knowledge.