Matias Luge - https://unsplash.com/photos/a-small-airplane-flying-through-a-cloudy-sky-h0f_MEqdIYM
Table of Contents
Introduction
The Wright brothers’ first flight is, in our view, the archetype of technological innovation that overturns a socially shared dogma. In these types of situations, a somewhat paradoxical phenomenon occurs. The belief that blocked what the community thought could be done, that defined the limits of what was possible within a discipline, suddenly collapses and triggers a cascading chain of events at exponential speed. In these cases, therefore, we move from a situation that had remained static for a long time due to a whole series of beliefs to a succession of changes.
In this article, we have referred to this phenomenon as the “Bannister effect,” a concept born in the sports field. This label serves as a “metaphorical model.” It has been used in sports and motivational coaching but is not an accepted term in epistemology. It is a metaphor to convey a concept in an easily understandable form.
From a socio-epistemological perspective, the Bannister effect has many similarities with Kuhn’s epistemology and his theory of the advancement of science through revolutions. We will also attempt to analyze this aspect. Compared to Kuhn, the Bannister effect has the added characteristic of generating change at a speed never seen before the overthrow of dogma. This represents a difference from the theory of scientific revolutions, in which the period of normal science is typically very long.
We are therefore interested in those cases, in science and technology, in which extremely rapid changes occur after some conceptual dogma has been overcome. The speed with which innovations appear after this critical point is difficult to understand without considering some shared conceptual barrier: a belief. This aspect of sharing implies a community, a social phenomenon. The speed of subsequent changes demonstrates that that belief was not justified. At this point, a critical analysis of the reasons for the belief allows us to understand the conceptual errors that have been committed and allows for a leap forward in our understanding of the discipline.
Of course, not every breakthrough triggers such a cascade; the Bannister effect describes a specific pattern that appears in some historically salient cases.
In our view, Cybenko’s Universal Approximation Theorem has posited another Bannister case in the field of artificial intelligence. What follows is intended to be a serious and critical analysis of this idea. However, the author is not an epistemologist, so the ideas expressed here are naturally open to question. The reader will forgive me if some concepts are not developed with the desired rigor.
The Wright Brothers’ Flight
The 12 Seconds That Changed History
On December 17, 1903, on the deserted beach of Kill Devil Hills (near Kitty Hawk, North Carolina), humanity achieved its first controlled, sustained, and powered flight aboard a heavier-than-air machine.
The feat was accomplished by brothers Orville and Wilbur Wright, two bicycle repairmen and builders from Dayton, Ohio, guided by rigorous scientific and engineering methods.
At 10:35 a.m., with a chill wind of about 25 mph (ca. 40 km/h), Orville Wright climbed aboard the aircraft from the center of the lower wing, lying flat on his stomach. The aircraft rolled along a wooden rail laid on the sand to facilitate takeoff and rose into the air. The flight lasted 12 seconds, covering a distance of 120 feet (ca. 37 m) at about 10 feet (ca. 3 m) above the ground. During the day, the two brothers alternated at the controls for a total of four flights. On the final attempt, Wilbur flew for 59 seconds, covering 260 meters before a gust of wind brought the craft down.

The Aircraft: Wright Flyer I
Unlike the pioneers of the time who sought to build powerful yet undetectable machines, the Wrights understood that the secret of aviation lay in three-axis attitude control (roll, pitch, and yaw). The biplane they used was built of fir and ash wood, covered in heavy cotton fabric (muslin), with a wingspan of 12.3 meters and an approximate weight of 340 kg, including pilot and fuel.
It was designed and built by them, together with mechanic Charlie Taylor: a lightweight 12-horsepower, four-cylinder in-line gasoline engine made of aluminum and cast iron. It consisted of two opposed, two-bladed wooden propellers connected to the engine by bicycle chains. Wing warping was achieved via tie rods connected to a cage for the pilot’s hips (roll), a front “canard” elevator (pitch), and a rear rudder (yaw).
A Bit of History
Although credited with the first human flight, the Wright brothers’ success was possible thanks to a whole series of ideas and previous attempts by other scientists.
One of the key ideas was the fixed wing. Understanding that one shouldn’t imitate birds by flapping one’s wings was a fundamental theoretical breakthrough for all of aviation, and it came about a century before Kitty Hawk.
The original intuition belongs to British engineer Sir George Cayley, considered the “father of aeronautics.” For centuries, from the studies of Leonardo da Vinci until the 19th century, humans had attempted to fly by creating machines with movable wings (ornithopters). In 1799, Cayley realized that two functions needed to be separated:
- Lift (suspension in the air) needed to be provided by a fixed wing made of a curved aerodynamic profile.
- Thrust (forward motion) needed to be provided by a separate propulsion system (such as a propeller or an engine).
In 1804, Cayley flew the first small-scale glider with a fixed wing and rudders, establishing the basic configuration of modern airplanes.
Later, German pioneer Otto Lilienthal demonstrated that a man could soar in the air supported by rigid, fixed wings. Between 1891 and 1896, Lilienthal made over 2,000 flights with rigid, profiled gliders. The Wright brothers extensively studied his lift tables and aerodynamic data.
George Cayley was a true pioneer of practical experimentation. He spent decades testing his aerodynamic insights, moving from laboratories to full-scale testing. Key milestones in his experiments include:
- The Whirling Arm (1804). To understand how different wing shapes interacted with the air, Cayley built a “whirling arm.” It was a scientific apparatus that rapidly spun airfoils in a circle, allowing him to precisely measure lift (upward thrust) and drag at various speeds and angles.
- The First Miniature Glider (1804). After laboratory tests, Cayley built a model glider approximately 1.5 meters long. It had a fixed, kite-shaped wing mounted on a spar, with a movable weight to adjust the center of gravity, and a cruciform tail (with horizontal and vertical rudder) for stability. Hand-launched, the model glided successfully, proving for the first time in history that his fixed-wing theory worked.
- The First “Passenger” Test (1849). At 76, Cayley designed and built a full-size glider. It was an unpowered triplane with spoked wheels. During tests on his Yorkshire estate, the glider was launched down a hill into the wind. According to contemporary accounts, a 10-year-old boy (the son of one of his employees) was put aboard, and the craft managed to lift off the ground for a short distance.
- The Historic Flight of the Coachman (1853). The culmination of his tests came in 1853, at Brompton Dale. Cayley built an even larger and more refined glider (nicknamed the governable parachute). Since he was now too old to fly, he ordered his coachman (often identified historically as John Appleby) to climb aboard. Towed by a team of men down a slope to gain speed, the glider left the ground and glided across a small valley for over 100 meters before landing abruptly. The flight was a scientific success, but somewhat less so for the pilot’s nerves. Legend, confirmed by several witnesses, has it that immediately after crashing without serious injuries, the coachman stood up and said, “Please, Sir George, I wish to tender my resignation. I was hired to drive carriages, not to fly!”
So What Was the Wright Brothers’ True Achievement?
The Wrights didn’t invent the fixed wing or the internal combustion engine, but they did put all the pieces together, solving aviation’s last true puzzle: active flight control.
Until then, the fixed-wing pioneers suffered from a major problem: their machines would crash as soon as they were hit by an unexpected gust of wind, since they had no effective means of controlling the aircraft in the air (Lilienthal himself died in 1896 due to an uncontrollable stall).
The Wrights understood that:
- The fixed wing had to be able to be slightly deformed (wing warping) to allow for lateral tilt (roll).
- The aircraft had to be inherently unstable, that is, continuously guided by the pilot on all three axes of rotation (roll, pitch, and yaw), just like balancing on a bicycle.
The Wright brothers were well aware of Sir George Cayley’s work and studied it extensively even before building their first glider.
In May 1899, when Wilbur Wright decided to seriously address the problem of human flight, he wrote a famous letter to the Smithsonian Institute in Washington, requesting all available texts, essays, and bibliographies on the subject of aeronautics. Among the materials received and studied by the brothers were:
- The publications of James Means‘ Aeronautical Annual (which reprinted Cayley’s historical essays).
- The book Progress in Flying Machines (1894) by Octave Chanute, the engineer who became the Wrights’ mentor, analyzed Cayley’s theoretical contributions and models in detail.
Thanks to Cayley’s writings—particularly his famous 1809–1810 treatise “On Aerial Navigation”—the Wrights found scientific confirmation of the foundations of aerodynamics:
- The separation of lift and thrust: It was not necessary to look for a mechanism that did both (like the wings of birds), but rather a fixed wing shaped to support the weight and a separate engine to propel the craft.
- The analysis of the four forces of flight: weight, lift, drag, and thrust.
The fundamental difference between the Wrights’ work and that of their predecessors lies in the precise definition of what happened on December 17, 1903. Cayley, Lilienthal, and other pioneers performed glider flights (i.e., gliding by exploiting air currents or the thrust of being launched down a hill) or experimented with balloons and dirigibles (lighter-than-air craft). History credits the Wright brothers with the “first flight in history” because their feat at Kitty Hawk simultaneously fulfilled all four fundamental criteria of modern aviation. No one before them had achieved this.
The pioneers before the Wrights discovered the laws of aerodynamics and demonstrated that air could support a wing. The Wright brothers put all the pieces of the puzzle together: they took Cayley’s fixed wing, Lilienthal’s data on wing sections, and added a lightweight engine and, most importantly, a three-axis control system. On December 17, 1903, they didn’t just glide: they took off from level ground, flew, and landed safely.
Wilbur Wright clearly expressed his debt to the past in a famous 1909 quote:
“Sir George Cayley carried the science of flight to a point never before reached and rarely surpassed by his successors.”
— Wilbur Wright, 1909
The 4 Criteria for a First Flight: The FAI Criterion
The FĂ©dĂ©ration AĂ©ronautique Internationale and aviation historians define the Wrights’ feat as the first flight because it was:
- Self-propelled (engined): The aircraft was not launched from a cliff or slope using gravity to glide. It took off under its thrust, produced by an internal combustion engine and propellers.
- Heavier-than-air: It was not a balloon or a dirigible (which float in the air thanks to gases such as hydrogen or hot air), but a rigid machine that generated lift simply by moving through the air.
- Controlled: This was the Wrights’ true masterpiece. The pilot could steer the aircraft, turn left and right, and ascend or descend voluntarily. All previous engine-powered attempts inevitably ended in a crash a few seconds after takeoff because the aircraft was uncontrollable.
- Claimed: The aircraft remained aloft, maintaining or gaining altitude, for a significant time and distance without losing speed, demonstrating continued flight as long as fuel remained.
The work of the Wrights’ predecessors does not meet one or more of these criteria:
- Sir George Cayley (1853): He performed a glide in a glider. The craft had no engine; it simply descended a slope under gravity.
- Clément Ader (1890): His steam-powered Éole rose about 50 meters from the ground at a height of 20 cm, but it was uncontrollable. It had no steering or balance system; it was an uncontrolled “leap into the air.”
- Otto Lilienthal (1891-1896): He was the undisputed king of gliders, but he never used an engine to take off from the plane and controlled the craft by shifting his body weight (as is done today with hang gliders), a system completely inadequate for heavier airplanes.
- Samuel Langley (1903): He designed the state-funded Great Aerodrome with a steam engine. But launched from a catapult mounted on a houseboat/barge in the Potomac River, the machine instantly destroyed itself and fell into the water due to a lack of structural rigidity and an attitude control system.
The Fédération Aéronautique Internationale was founded in 1905, so these criteria were not yet listed in 1903.
Before and After Kitty Hawk: From Unprepared Madmen to Visionaries
I use the word “madmen” in a friendly tone, full of admiration for the work of the Wright brothers. In the scientific and technological fields, the difference between madmen and visionaries often lies simply in the success of an idea.
Before December 17, 1903, there was general skepticism. Among the criticisms leveled at the Wright brothers were:
- Mechanics without degrees: For the scientific and ecclesiastical communities of the time, the idea of ​​powered human flight was considered an unattainable chimera or an act of sacrilegious presumption. The Wright brothers, lacking academic degrees or government funding, were considered mere “eccentric bicycle mechanics.”
- The failure of the giants: A few days earlier, on December 9, 1903, the prototype funded by the Smithsonian Institution and designed by scientist Samuel Langley had crashed miserably into the Potomac River. This event reinforced the public’s and newspapers’ belief that flying was physically impossible.
After December 17, 1903, the world went from mistrust to worldwide triumph. Initially, years of disbelief followed (1903–1907), during which news of the flight failed to spark immediate enthusiasm. The Wright brothers, highly secretive and jealous of their patents, refused to perform public demonstrations without signed contracts. The American press (including the New York Herald) long treated them with suspicion, ironically dubbing them “Flyers or Liars?”
The turning point came in France in 1908. In August, Wilbur traveled to Le Mans and performed daring demonstration flights in closed circuits, performing perfect turns and figure-eight maneuvers in front of engineers and journalists from around the world.
Critics were forced to reconsider. The world realized that the two brothers had truly overcome gravity and solved the laws of aerodynamics. From that moment on, they were celebrated worldwide as the founding fathers of modern aviation.
The Red Baron and the First Airlines
The evolution of aviation in the first three decades of the twentieth century represents one of the most rapid technological leaps in human history. From the Wright brothers’ first 12 seconds on a wood-and-canvas frame, in less than thirty years, tens of thousands of people were transported on airliners.
The first step was taken with the First World War (1914). At that time, airplanes were considered little more than reconnaissance toys or toys for brave pioneers. However, the conflict transformed aviation into a truly high-tech industry.
Manfred von Richthofen (the Red Baron) was the most famous flying ace of the Great War, credited with 80 official aerial victories. His name is linked to the Fokker Dr.I (the triplane), but also to the synchronization mechanism, an invention by the German company Fokker that allowed the machine gun to fire through the propeller blades without destroying them. The war pushed engineers and manufacturers to radically increase engine power, aerodynamics, reliability, and maneuverability.
With the end of the war in 1918, the world suddenly found itself with thousands of surplus military aircraft and thousands of unemployed trained pilots. This was the catalyst for the commercialization of flight.
Initial passenger transport was too expensive and dangerous. It was airmail (Aéropostale in Europe, led by figures such as Antoine de Saint-Exupéry, and US Air Mail in the United States) that made aviation an economically viable commercial activity. Governments paid well to speed up civil and commercial communications.
The first airlines also date from this period:
- 1914 (St. Petersburg–Tampa Airboat Line): The world’s first commercial airline (one seaplane carrying one passenger at a time) was founded in Florida.
- 1919 (KLM, Avianca, DELAG/Lufthansa): The first structured airlines were founded in Europe. The first passenger aircraft were simply converted bombers with seats installed in the hold.
To make the airplane a safe and mass-produced means of transportation (or at least accessible to the middle class), three major innovations were needed, which arrived between the late 1920s and 1930s:
- All-metal structures (1920s): The abandonment of wood and canvas in favor of lightweight metal alloys such as duralumin (e.g., the Junkers JU-52 and Ford Trimotor trimotors) greatly increased safety, structural strength, and load capacity.
- Navigation and weather instruments (1929): With the development of the gyroscope, barometric altimeter, and radio navigation, blind flight (IFR) was born. Aircraft could also fly at night and in fog or bad weather.
- The Douglas DC-3 Revolution (1936): The Douglas DC-3 is the airplane that changed everything. Capable of carrying 21 passengers in comfort, speed, and economy, it was the first aircraft in history to allow airlines to make a profit by carrying only passengers, without having to rely on government subsidies for mail.
The Bannister Effect
Anatomy of a Phenomenon
There’s a question that has always struck me about flight and the Wright Brothers since I read a book on the history of technology many years ago. I was in college at the time.
For centuries, human “heavier-than-air” flight was perceived not as a simple engineering problem to be solved, but as an ontological limit of nature. Despite the ingenuity of figures like Leonardo da Vinci and the insights of 19th-century pioneers, the prevailing belief in the scientific community was that man could never master the skies with a motorized machine.
As recently as 1895, Lord Kelvin, world-renowned physicist and president of Britain’s prestigious Royal Society, stated:
“I have not the smallest molecule of faith in aerial navigation other than ballooning, or of the expectation of good results from any of the trials we hear of.”
— Lord Kelvin, letter to Major Baden Baden-Powell (president of the Aeronautical Society), December 8, 1896.
This psychological and conceptual barrier finally collapsed on December 17, 1903, in Kitty Hawk, North Carolina. What happened in the next 15 years is unparalleled in the history of technology: from the first precarious “leap” of a few meters on the sand, humanity moved on to metal aircraft capable of soaring through the skies at hundreds of kilometers per hour, laying the foundation for the first non-stop transatlantic flight by airplane in 1919.
How was it possible to solve the complex problems of aerodynamics, propulsion, and flight control in just over a decade? The answer lies in a socio-epistemological phenomenon comparable to the so called “Bannister Effect.”
For decades, the world of athletics maintained that running a mile in under four minutes was an insurmountable physiological limit for the human heart. On May 6, 1954, Roger Bannister broke that barrier by running the mile in 3:59.4. The previous record was set nine years earlier: on July 17, 1945, Gunder Hägg ran the mile in 4:1.4.
There’s an anecdote about that day. When the stadium announcer, Norris McWhirter, began announcing the final time, saying, “Result of the mile race… time: three minutes…,” the rest of the announcement (“…fifty-nine and four-tenths of a second”) was almost completely drowned out by the roar of the delirious crowd, aware that the four-minute mark had just been broken for the first time in human history.
In the months and years that followed, dozens of other athletes achieved the same feat: the anatomical limit had never existed; only a mental barrier existed.
In technology, the Wright brothers achieved the same effect. Their true revolution was not the invention of every single component of the aircraft from scratch, but the empirical proof that the achievement was possible.
This is the “Bannister effect,” borrowed from sociology and social psychology to describe the phenomenon that occurs when some innovation unlocks a shared social belief that isn’t objectively justified. Often, a dizzying chain of events then unfolds that demonstrates that the limit didn’t actually exist. This causes a discipline or technology to grow exponentially, attracting funding and talent.
A Sought Record
Roger Bannister’s profile makes his feat even more fascinating for this very reason: he wasn’t a full-time professional athlete but a medical student at St. Mary’s Hospital Medical School in London (who later became a distinguished neurologist).
Bannister applied his physiological knowledge to address the problem of the 4-minute limit:
- Respiration research: During his medical studies, Bannister conducted laboratory experiments on oxygen consumption and the cardiovascular response to extreme effort, sometimes using supplemental oxygen to explore physiological limits.
- Scientific method in training: At a time when runners simply logged miles, Bannister, along with his coach Franz Stampfl, developed an approach based on high-intensity interval running (interval training), designed to optimize the body’s energy efficiency while minimizing training time (he only trained about 45–60 minutes a day during his lunch break in his doctor’s office).
In the 1940s and 1950s, several doctors and scientific articles argued that attempting to run a mile under 4 minutes was life-threatening: it was theorized that the heart could fail, the lung walls could tear, or that the human body lacked the metabolic capacity to clear lactic acid at such speeds. These claims were not rigorously proven.
Bannister critically analyzed these theories and understood that the human body had the biological capacity to do so, provided:
- Distribute the effort perfectly evenly (hence the strategic use of pacemakers or “pacers” like Chris Chataway and Chris Brasher to maintain a constant pace of 60 seconds per lap).
- Overcome the psychological barrier of pain, knowing scientifically that the body would not collapse.
While studying medicine, Bannister realized that the limit was not in the structure of our muscles or our hearts, but in our minds. The capacity to tolerate pain and fatigue is regulated by the brain long before the body reaches its true physical limit.
Years later, as an internationally renowned neurologist (he was also knighted for his contributions to medicine and sport), Bannister continued to publish scientific articles on exercise physiology and the autonomic nervous system.
The Bannister Effect from an Epistemological Perspective: Kuhn’s Revolution
Although not formally considered, the Bannister effect is, in fact, a socio-epistemological phenomenon.
Although the term “Bannister Effect” originates in the fields of sports and psychology, its profound dynamic describes how a community constructs, shares, and modifies its beliefs about what is “true” or “possible.”
Epistemology deals with the nature, origins, and limits of knowledge. From this perspective, regarding Bannister’s record, we have the following stages:
- Construction of the cognitive limit: Before 1954, the belief that the human body could not run a mile in under four minutes was treated as a scientific and biological certainty (until the unfounded hypothesis that the heart would explode under such strain).
- Paradigm disruption: Bannister achieved an athletic feat and provided empirical evidence that disproved the previous theoretical model. In Kuhnian terms (Thomas Kuhn), it is a scientific micro-revolution: the shattering of an epistemic dogma in favor of a new frame of reference.
A phenomenon becomes socio-epistemological when knowledge is no longer confined to the individual but rather changes the cognitive architecture of an entire social group. The salient elements are:
- Epistemic intersubjectivity: Social consensus conditioned individual perceptual limits. The athlete was not only fighting against the clock but also against the cultural convention that declared the feat impossible.
- Cognitive cascade: As soon as Bannister demonstrated the feasibility of the feat, the mental constraint fell for everyone. Just 46 days later, John Landy further lowered the record, and in the following months and years, dozens of runners surpassed the same threshold.
- Collective Self-Efficacy: Drawing on the concepts of sociologist and psychologist Albert Bandura, Bannister’s success reprogrammed the self-efficacy of the entire running community. The shift in collective belief produced a real transformation in the biological/performance capabilities of individuals.
The following table shows the changes that occurred before and after Bannister’s feat:
| Dimension | Before Bannister | After Bannister |
| Limit Status | Biological/Objective Limit (“Reality”) | Social/Psychological construction (“Belief”) |
| Epistemic Status | Demonstrated impossibility | Ascertained and extendable possibility |
| Social Impact | Inhibition of action and performance | Acceleration of new standards and innovation |
The Bannister Effect shows how social knowledge is not just a passive recording of reality, but an active structure that defines the boundaries of action and possibility in areas such as science, technology (e.g., the pioneering impact of electric cars or space travel, the Wright Brothers’ first flight), and social reform.
For a quick visual summary of how overcoming mental limitations generates a ripple effect in human potential, watch the video The Bannister Effect: How Breaking Barriers Unlocks Limitless Potential. This content is particularly relevant because it summarizes Roger Bannister’s story and clearly illustrates how breaking a cognitive barrier transforms collective perception and performance.
The parallel between the Bannister Effect and the structure of scientific revolutions theorized by Thomas Kuhn in 1962 is striking: both describe how a rigid system of beliefs—based on dogmas perceived as absolute—collapses in the face of a practical and conceptual disruption, redefining what a community considers “possible.”
In Kuhn’s view, a scientist works within a paradigm: a theoretical framework, a set of techniques, assumptions, and tools shared by a scientific community. During the period of normal science, researchers did not seek to challenge the paradigm but to solve “puzzles” within it. In sports, before 1954, the “paradigm” of track and field established that human physiology had an insurmountable limit: running a mile in under 4 minutes was biologically impossible (with physiological theories hypothesizing cardiac failure or lung injury). Training at the time was structured within this theoretical boundary.
As time passes, anomalies accumulate in science: empirical findings or phenomena that the prevailing paradigm fails to explain or integrate. Roger Bannister (who, ironically, was a medical student and neurologist) didn’t simply run faster; he approached the problem by questioning physiological dogma. He studied the mechanics of breathing and the efficiency of oxygen consumption, identifying that the barrier wasn’t a biological wall but a critical threshold that could be overcome with the right method. On May 6, 1954, by running the mile in 3:59.4, Bannister generated an undeniable empirical anomaly: the previous dogma proved false.
When a fundamental anomaly can no longer be ignored, the community enters a phase of epistemic crisis. Scientists (or, in our case, athletes and coaches) must reorganize their conceptual framework.
Kuhn compares the paradigm shift to a “Gestalt switch” (a shift in perception): the underlying image doesn’t change, but the mind interprets it completely differently. For example, suddenly, the 4-minute barrier is no longer considered a natural limit but as a “psychological threshold” (a limit of perception).
Once the new paradigm is accepted, the community enters a new phase of stability, but on an entirely different basis. It adapts to the new standard, setting in motion a true “cognitive cascade”:
- John Landy broke the record just 46 days after Bannister with a time of 3:58.0.
- Over the next three years, over 15 athletes broke the same “wall.”
- Today, running a sub-4-minute mile is the minimum standard for any international-level middle-distance runner.
The human body didn’t evolve genetically within a few weeks in 1954; what evolved was the community’s belief structure.
As has been said, the parallels with Kuhn’s revolutions represent an author’s original argument. Therefore, it’s worth making some refinements to this discussion to identify points of contact and differences with Kuhn’s theory:
- In Kuhn, revolutions typically change fundamental concepts, methods, and often even worldviews (e.g., Newton → Einstein, geocentrism → heliocentrism).
- The Bannister case is more of a local micro-revolution: it changes a perceived boundary in a very specific domain (the mile record), but it doesn’t transform basic biology or physiology.
- In Kuhn, anomalies are often experimental data that don’t fit theoretical predictions.
- In the Bannister case, the anomaly is a sports record, which is still an empirical finding, but in a different context than physics or chemistry.
Here is a summary table of Bannister’s feat in light of Kuhnian theory:
| Phase of Kuhn’s Theory | Scientific Model (Kuhn) | Performance Model (Bannister) |
| Normal Science | Problem-solving within established theories. | Training optimized to approach (without breaking) 4 minutes. |
| Dogma / Assumption | Theory accepted as objective truth. | “The human heart cannot withstand a sub-4-minute mile.” |
| Anomaly | Experiment or data that contradicts the theory. | May 6, 1954: Bannister runs 3:59.4. |
| Scientific Revolution | Replacement of the old paradigm with a new one. | The limit shifts from “physical” to “psychological”. |
| New Normal Science | Scientific work based on the new paradigm. | Dozens of athletes run sub-4-minute miles in the following months and years. |
A Bayesian Approach to the Bannister Effect
The Bannister effect can be modeled well in Bayesian terms. Let’s start with the following elements:
- H\equiv “A human can run the mile in < 4 minutes.”
- E\equiv “Bannister ran the mile in 3:59.4.”.
Before the record, each athlete has a prior distribution of “what is the best time realistically achievable by a human.” The belief “< 4 minutes is impossible” translates into a very low probability (practically zero) assigned to the event.
After Bannister’s record, new information arrives (E). Athletes update their belief about H using a form of Bayesian updating:
P(H|E)=\frac{P(E|H)P(E)}{P(E)}P(H) was very low (prior). P(E|H) is high (if it’s possible, it’s plausible that someone will do it sooner or later). Observing E makes P(H|E) skyrocket.
More intuitively:
- Before: The conditional probability of “being able to run under 4 minutes, given that it’s humanly possible” was already high, but the probability of “it’s humanly possible” was almost zero.
- After: The observed event radically changes the estimate of “it’s humanly possible,” and consequently changes the entire chain of decisions (how to train, what goals to set, how much to push yourself in the race).
Wooten‘s article (“Leaps in innovation and the Bannister effect in contests,” 2022) describes exactly this: the breakthrough functions as an information signal that induces a Bayesian updating of what is feasible, and this pushes other participants to attempt bolder approaches.
This is not a formal Bayesian model with calibrated priors and likelihoods, but rather a conceptual analogy: the record acts as a public signal that triggers a radical belief update across the community. It is a public signal that changes the priorities of the entire community, not just of an individual.
Examples of the Bannister Effect in Technology and Science
A “Bannister moment” in the world of technology and science occurs when a feat considered for years an insurmountable barrier (theoretical, economic, or engineering) is overcome for the first time. As soon as feasibility is demonstrated, the entire industry abandons its old skepticism, and what was once “science fiction” quickly becomes the new industry standard, setting off a dizzying chain of events.
Let’s analyze four emblematic cases by exploring the epistemological stages we described previously. The identification of these cases as a “Bannister effect” is my interpretation that can be questioned:
- The reentry and landing of a rocket’s first stage (Falcon 9, 2015):
- The previous dogma: For over 50 years, the aerospace industry treated spacecraft as disposable products. Aerospace giants and government agencies considered propulsive recovery of an orbital rocket’s first stage an undertaking too complex from an engineering standpoint and economically unviable due to the weight of the required propellant.
- Breaking the Limit (2015): On December 21, 2015, SpaceX successfully landed the first stage of the Falcon 9 at Cape Cañaveral.
- The Bannister Effect: The “disposable rocket” dogma instantly collapsed. In the years that followed, all the world’s major aerospace agencies and companies (Blue Origin, Rocket Lab, the Chinese space agency, and even Europe with future projects) redesigned their roadmaps to focus on reusable launchers.
- The Turning Point of Private Markets: The Launch of Dragon (2010/2020):
- The Previous Dogma: Sending humans or supplies into space was the exclusive preserve of states and government agencies (NASA, Roscosmos). The idea that a private company could design, build, and operate a safe spacecraft was considered a pipe dream.
- Breaking the Limit: In 2010, the Dragon spacecraft became the first private vehicle to re-enter orbit; in 2020 (Crew Dragon Demo-2), two astronauts landed on the International Space Station for the first time.
- The Bannister Effect: Its success paved the way for the era of commercial spaceflight, unlocking billions of dollars in private investment and giving rise to an entire ecosystem of space startups (private space stations, commercial lunar landers).
- Deep Blue beats Garry Kasparov in chess (1997):
- The previous dogma: Chess was considered the pinnacle of human intuition and creativity. Grandmasters and computer scientists believed that a computer could never develop the strategic capacity necessary to beat the reigning world champion in a formal challenge.
- Breaking the Limit: In May 1997, the IBM supercomputer Deep Blue defeated Garry Kasparov in a six-game match.
- The Bannister Effect: The victory shifted the frontier of research: the focus of computer science shifted from chess to far more complex challenges (such as the game of Go with AlphaGo or protein structure prediction with AlphaFold), redefining the concept of what computing power can solve.
- The Transformer Architecture and the Explosion of LLMs (2017):
- The previous dogma: Natural language processing (NLP) suffered from severe bottlenecks: the neural networks of the time (RNNs and LSTMs) analyzed text one word at a time, making training on massive volumes of data extremely slow, inefficient, and incapable of handling long contexts.
- Breaking the Limit: In 2017, Google researchers published the paper “Attention Is All You Need,” introducing the Transformer architecture.
- The Bannister Effect: Once parallelization and attention mechanisms were demonstrated to be capable of handling massive amounts of text, the entire AI industry transformed within months. The GPT, Claude, Llama, and Gemini models quickly emerged, transforming AI from an academic niche to a global industrial revolution.
In each of these cases, the main barrier wasn’t just technological but conceptual: once someone pioneered the process, the entire community figured out how to think about the problem, setting off a chain reaction of innovation. The response from the market and the scientific community certainly occurred almost immediately in these cases, but at different speeds for each actor involved.
It Can Be Done! The Brake Due to the Lack of a Precedent
As long as an undertaking is considered impossible, industrial and state capital does not approach it, brilliant minds dedicate their energies to other sectors, and advanced theoretical knowledge remains fragmented and detached from practice.
The moment scientific possibility is demonstrated, philosophical doubt (“Can it be done?”) instantly transforms into an engineering challenge (“How can we do it better, faster, and higher?”).
Contrary to what one might think, as we have already analyzed, the fundamental scientific theory needed to fly existed largely before 1903:
- The foundations of modern aerodynamics and the concept of “lift” had been formalized by Sir George Cayley as early as the early nineteenth century.
- The Navier-Stokes equations of fluid dynamics had been known since the mid-19th century.
- Otto Lilienthal had already compiled accurate aerodynamic tables studying the behavior of wings with gliders in the 1890s.
What was missing was the certainty that converging these theories into a single dynamic system (engine, wing cell, and three-axis control systems) would produce the desired result. The Wrights didn’t create the physics of flight; they provided the empirical spark that validated decades of popular theory.
Once the taboo was broken, technological progress accelerated exponentially. The strategic interest of governments and industrial competition catalyzed unprecedented financial and human resources.
The quantitative and qualitative evolution that occurred between Kitty Hawk’s flight and the end of World War I demonstrates the speed with which engineering responded to the certainty of the outcome:
| Performance Parameter | 1903 (Wright Flyer) | 1918 (Late-Generation Aircraft, e.g., SPAD XIII / Fokker D.VII) |
| Top Speed | ~48 km/h | Over 220 km/h |
| Service Ceiling (Max Altitude) | 3 meters | Over 6,000 meters |
| Engine Power | 12 HP (custom aluminum engine) | 200–300 HP (standard production rotary and inline engines) |
| Materials & Structure | Spruce wood, canvas, precarious metal bracing wires | Metal fuselages (Junkers J 1), advanced airfoils |
| Control Systems | Wing warping (manual wing twisting) | Ailerons, articulated rudders, flight instruments for blind flying |
In just 15 years, aviation moved from the stage of pioneering empiricism to the era of industrial aeronautical engineering. In 1919, just 16 years after the Wright brothers’ first flight, John Alcock and Arthur Brown completed the first nonstop Atlantic crossing aboard a modified Vickers Vimy bomber.
The history of flight demonstrates that the greatest obstacle to technological progress is not always a lack of theoretical knowledge or the scarcity of resources. Often, the main barrier is the lack of a precedent.
The Wright brothers didn’t just build a machine of wood and canvas; they redefined the boundaries of possibility. By demonstrating to the world that the problem of gravity had a technical solution, they unlocked humanity’s scientific potential, unleashing an extraordinary wave of innovation that, within a single generation, took humanity from the sand fields of Kitty Hawk to the edge of the stratosphere.
Often, the greatest barrier to technology is not a lack of knowledge but a lack of precedent. The Wright Brothers didn’t invent all the technologies of flight; they did something greater: they provided proof that the problem had a solution. In the 15 years since, humanity has simply filled in the details of a map whose destination it finally knows.
This mental framework also applies very well to other modern fields (such as the space race in the 1960s, DNA sequencing, or the development of RNA vaccines): as soon as scientific certainty overcomes skepticism, technology advances exponentially.
Cybenko’s Theorem
What Cybenko’s Theorem Says
Cybenko’s Theorem (published by mathematician George Cybenko in 1989) is one of the fundamental theoretical pillars of artificial intelligence and machine learning. It is commonly known as the Universal Approximation Theorem.
In a nutshell, the theorem proves something surprising:
Cybenko’s Theorem: An artificial neural network with a single hidden layer can approximate any continuous function with arbitrary precision.
The theorem asserts something mighty: arbitrary precision. We can set the gap between the approximation and the function to be approximated at will.
This result can be supplemented with any proofs relating to the convergence time of the procedure. While there are no universal and general proofs for every neural architecture, for some it is possible to show that the approximation occurs in finite time, as shown in the following table:
| Architecture | Cybenko Status | Training Algorithm | Time Complexity |
| RVFL / ELM (Fixed random weights) | Yes (Pao 1994, Huang 2006) | Analytical pseudoinverse (closed-form solution) | Instantaneous Finite Time (O(N^3) algebraic operations) |
| Reservoir Computing (ESN/LSM) | Yes (Maass 2002, Jaeger 2004) | Least squares on readout layer | Instantaneous Finite Time (O(N^3) algebraic operations) |
| MLP in NTK regime (Ultra-wide networks) | Yes (Jacot et al. 2018, Arora et al. 2019) | Standard Gradient Descent | Polynomial Finite Time (O(\log(1/\epsilon)) in continuous time) |
| Conventional MLP (Standard Backprop) | Yes (Cybenko 1989, Hornik 1991) | SGD/Adam on non-convex problems | No general guarantee (NP-hard, Blum & Rivest 1992) |
In the case of neural architectures that satisfy both results, that is, for which we have convergence estimations and which at the same time verify Cybenko’s theorem, we have at our disposal formidable algorithms that allow us to obtain approximations with arbitrary precision while at the same time having a guarantee of the result (obtaining the desired approximation in a finite number of steps).
Intuition: Mathematical “Legos”
To understand the theorem without getting lost in the formal details, imagine wanting to reproduce a mountain silhouette or profile using Lego bricks. A neural network actually does just that:
- Each individual neuron in the hidden layer takes an input signal, multiplies it by a weight, adds a bias, and applies a nonlinear function (the so-called activation function, such as the sigmoid).
- This single neuron produces a small “step” or a gentle curve in the final graph.
- If you add together enough of these appropriately adjusted steps, you can reconstruct the shape of any curve or continuous surface, no matter how complex or irregular.

The partial overlapping of the steps produced by each neuron has an effect similar to that of fitting several Lego bricks one on top of the other to adapt our construction to the shape we want to reproduce.
A Slightly More Formal Statement of the Theorem
In his original 1989 formulation (“Approximation by Superpositions of a Sigmoidal Function”), Cybenko formally proved that:
Cybenko’s Universal Approximation Theorem: Let \sigma be a continuous function of sigmoidal type. Given any continuous function f defined on a compact set in \mathbb{R}^n and an arbitrary tolerance \varepsilon > 0, there always exists a linear combination of the form:
G(x) = \sum_{i=1}^{M} \alpha_i \, \sigma(w_i^T x + b_i)
such that \vert{}G(x) - f(x)\vert{} < \varepsilon for all x.
In subsequent years, mathematicians such as Kurt Hornik (1991) extended its validity to almost any nonlinear activation function (including the now widely used ReLU).
If Only One Layer Is Enough, Why Do We Use Deep Learning?
Cybenko’s theorem is a remarkable result, but it presents three major practical pitfalls that explain why modern networks have tens or hundreds of layers (deep networks) instead of just one (shallow networks):
- Existence \neq Construction: The theorem guarantees that the perfect combination of weights w_i and b_i exists, but it doesn’t explain how to find it. The training algorithm (backpropagation) may never converge to those ideal values. As we’ve seen, however, there are procedures that satisfy the theorem and for which convergence can also be guaranteed.
- Neuron Explosion: To approximate high-dimensional or very complex functions with a single layer, the number of neurons required M can grow exponentially. A network with billions of neurons in a single layer requires an astronomical amount of memory and data.
- Depth Efficiency: By stacking multiple layers on top of each other (deep learning), the network can construct hierarchical abstractions. A “deep” network can represent the same complexity as a “wide” network using an infinitesimal fraction of parameters.
The following table summarizes the elements we have described:
| Aspect | What the Cybenko Theorem States |
| What it proves | A single-layer neural network is a universal approximator. |
| Requirements | A finite (but potentially enormous) number of neurons and a non-linear activation function. |
| Practical limitation | It does not guarantee that the network is trainable or computationally efficient. |
How Does a Broken Straight Line Approximate Complex Curves: Hinged Triangles
The ReLU (Rectified Linear Unit) is the most widely used activation function in deep learning today. Its formula is simple:
\sigma(x) = \max(0, x)
Graphically, it’s a broken line: it returns 0 for negative values ​​and passes the value x unchanged for positive inputs.
At first glance, ReLU seems far too simple compared to the sigmoid (which is a smooth curve). The key to understanding how ReLU approximates any continuous function lies in piecewise linear geometry.
If you take two opposing ReLU neurons and add them with the right weights and biases, you can isolate a single peak or “curtain”:
- One neuron “turns on” the upward ramp at a certain point A.
- A second neuron resets or inverts the ramp at a point B.
By combining multiple ReLU neurons, the network constructs triangular peaks of the desired amplitude and position. By adding hundreds or thousands of these small triangles (or hyperplanes in higher dimensions), the network creates a piecewise continuous function.
As the number of neurons (i.e., the number of segments) increases, these straight lines become increasingly shorter, able to follow the profile of any complex curve with the desired precision, just as pixels on a screen form a sharp image if taken in sufficient numbers.

Generalizations of the Theorem
Cybenko’s original theorem (1989) showed that a feedforward network with a single hidden layer and a sigmoid activation function can uniformly approximate any continuous function on a compact set of \mathbb{R}^n, with arbitrarily small error.
Kurt Hornik (with Stinchcombe and White in 1989 and alone in 1991) generalized this result in several fundamental directions. The key result is:
Generalized Cybenko Theorem (Hornik): The universal approximation property does not depend on the specific choice of the sigmoid activation function. It is sufficient that the activation function be nonconstant, bounded, and monotonically increasing.
This shifts the “credit” for universality from the sigmoid form to the multilayer/feedforward architecture itself. Furthermore, Hornik extends the result to measurable functions (not necessarily continuous), not just on compact sets.
A particularly important contribution: Hornik shows that multilayer feedforward networks can approximate not only a function but also its derivatives (up to a certain order), provided the activation function is sufficiently smooth. This is crucial for applications where the network also needs to approximate the gradient of the target function well.
While Cybenko answered the question “Can a sigmoid network approximate?” Hornik answers a more profound question: “Is it the structure of multilayer feedforward networks—and not the particular choice of nonlinearity—that guarantees universal approximation?” This makes the result much more robust and general, justifying the practical use of different activations (tanh, ReLU in later versions of the theory, etc.) while maintaining the theoretical guarantees.
In 1993, researchers Leshno, Lin, Pinkus, and Schocken proved an even more refined version of the Universal Approximation Theorem:
The Generalized Cybenko Theorem (Leshno, Lin, Pinkus, and Schocken): A feedforward network with a single hidden layer is a universal approximator if and only if the activation function \sigma is not a polynomial.
ReLU vs. Sigmoid
While both provide universal approximation on paper, in practice, ReLU has revolutionized the training of deep networks. A comparative analysis highlighting why ReLU is preferable can be found in the following table:
| Property | Sigmoid ​\sigma(x)=\frac{1}{1+e^{-x}} | ReLU \sigma(x)=\max(0,x) |
| Derivative Computation | Expensive (exponentials) | Instant (0 or 1) |
| Vanishing Gradient Problem | Severe at extreme values (saturates at 0 or 1) | Absent for positive values (fixed gradient of 1) |
| Training Efficiency | Slow in deep networks | Much faster to train via backpropagation |
Cybenko’s Theorem and the Bannister Effect
A Bit of History
In this article, I argue that Cybenko’s theorem can be interpreted as a Bannister effect. In my opinion, this thesis is conceptually sound and largely tenable, provided we carefully distinguish between its historical-epistemological value and what the sources formally document.
When the idea occurred to me, the first thing I did was search for sources where it was explicitly cited. To the best of my knowledge, this specific connection between the Bannister effect and Cybenko’s theorem has not been explored in the academic literature; the present article proposes it as an original interpretative hypothesis (“Bannister Effect” is a term borrowed from psychology and sociology, not a formal category in computer science.) However:
In my opinion, the historical documentation in the field (1986–1989) unequivocally describes precisely the dynamics of the Bannister effect surrounding George Cybenko’s proof.
This dynamic unfolds in the following sequence of events.
First, there’s the “conceptual wall” spanning the period from 1969 to 1986 (1980 really for most authors). In 1969, Marvin Minsky and Seymour Papert published the celebrated book Perceptrons (MIT Press). These authors mathematically demonstrated the limitations of single-layer perceptrons (incapable of solving nonlinear problems like XOR). Although their theorem was limited to simple perceptrons, the scientific community interpreted the result as a doom for the entire discipline of neural networks, which became a “dogma of inadequacy.” This result, together with other factors such as unmet expectations and funding cuts, contributed to the so-called “AI Winter.” For nearly two decades, funding dried up, and the prevailing belief was that neural networks were a conceptual and mathematical dead end.
In 1986, a partial breakthrough occurred. Rumelhart, Hinton, and Williams republished the backpropagation algorithm (Learning Representations by Back-Propagating Errors. Nature), demonstrating that multilayer networks could learn internal representations. However, the fundamental mathematical proof that such networks could approximate any continuous function was still missing. Hinton won the 2024 Nobel Prize in Physics for this and other works, together with John Hopfield.
The breakthrough, as we already know, came in 1989 with George Cybenko’s paper, “Approximation by Superpositions of a Sigmoidal Function,” published in the journal Mathematics of Control, Signals, and Systems. Cybenko mathematically proved the Universal Approximation Theorem: a feedforward neural network with a single hidden layer and a continuous sigmoidal activation function can approximate any continuous function on a compact space with an arbitrary degree of accuracy. Cybenko did exactly what Bannister did: he removed the doubt about its theoretical feasibility. It doesn’t tell you how to train the network efficiently for every single practical problem, but it demonstrates with mathematical rigor that the model’s expressive capacity has no intrinsic limits on the approximation of continuous functions on compact domains, provided enough hidden units are available.
Immediately after Cybenko’s paper, the entire field experienced a sudden acceleration. Hornik, Stinchcombe, and White (1989/1991) published fundamental extensions demonstrating that universal approximation does not depend on the specific sigmoidal function but on the multilayer nature of the network (Multilayer feedforward networks are universal approximators, Neural Networks). The dogma of inadequacy finally collapses: neural networks return to the center of the global scientific agenda, paving the way for the era of deep learning.
To summarize, the following table presents a comparative analysis of the Bannister-Cybenko cases, highlighting the points of contact that support the thesis of a Bannister effect surrounding the Universal Approximation Theorem:
| Element | Bannister Effect (Athletics, 1954) | Cybenko Theorem (AI, 1989) | Reference Source |
| The Wall / Dogma | “The human body cannot sustain a sub-4-minute mile.” | “Neural networks are not theoretically suitable for modeling complex problems.” | Minsky & Papert, Perceptrons (1969) |
| The Reconfiguration | Bannister analyzes physiology and proves feasibility. | Cybenko applies functional analysis (Hahn-Banach Theorem) to prove theoretical suitability. | G. Cybenko, MCSS (1989) |
| The Feat (“Anomaly”) | May 6, 1954: 3:59.4 | 1989: Proof of Universal Approximation | Cybenko Paper (1989) |
| The Cognitive Cascade | Dozens of runners break the barrier shortly after. | Hornik, Leshno, Lin, Pinkus, and Schocken extend the theorem: explosion of neural network publications. | Hornik et al., Neural Networks (1989, 1991) |
Cybenko’s 1989 theorem served for artificial intelligence the same socio-epistemological function that Roger Bannister’s feat served for athletics in 1954 or the Wright brothers’ for aeronautics in 1903: it not only solved a specific problem, but it also demolished a conceptual dogma (triggered by the limitations of Minsky and Papert’s perceptron), recalibrating the scientific community’s self-efficacy and orientation and unlocking a cascade of new research and theoretical extensions.
The First Misunderstanding
With all this evidence in hand, however, the very fact that we arrived at AI Winter is incomprehensible, even for technical reasons: the Perceptron is a linear network, while current networks are nonlinear. Is it possible that a result valid only for the former has been generalized, that no one has verified what happens with nonlinear neural structures?
To understand why, we need to distinguish between what Minsky and Papert demonstrated mathematically and what the scientific community believed they had demonstrated.
Minsky and Papert (1969) formally demonstrated the limitations of single-layer perceptrons (without hidden layers), which, being linear separators, could not solve basic nonlinear problems like XOR. Minsky and Papert understood that by adding hidden layers and nonlinearity, the formal limit was not mathematically applicable, but they expressed skepticism about the theoretical and practical possibility of effectively training multi-layer networks.
At this point, a “dogma” effect (a phenomenon studied by the sociology of science) set in. The scientific community and funders of the time (such as DARPA) unduly extrapolated this result. The prevailing view became, “Neural networks are limited and conceptually flawed architectures.” The limitation of a single-layer linear structure was mistakenly projected onto the entire neural network paradigm.
This is where the importance of the Universal Approximation Theorem comes in. Cybenko worked on feedforward networks by introducing a hidden layer with a nonlinear activation function (specifically, sigmoidal functions such as the logistic) and demonstrated that by combining the weighted sum of linear combinations with a nonlinear activation function in the hidden layer, the network can approximate any continuous function. This caused an epistemological reversal. Cybenko bridged the chasm, leaving little mathematical doubt: introducing nonlinearity into layered models was not only a good empirical heuristic but also guaranteed a theoretically unlimited expressive capacity for any continuous function.
If we analyze this detail, the parallel with the Bannister effect and paradigm shifts becomes even more evident:
| Dimension | Bannister Effect (1954) | Cybenko’s Theorem (1989) |
| The “False” Wall | The belief that the human body could not sustain a sub-4-minute mile (confusing an empirical estimate with a biological limit). | The belief that neural networks were limited by the linear separability problem (confusing the limits of the basic perceptron with those of neural networks in general). |
| The Key Element | Bannister realized that a new exertion/breathing method was needed to unlock the body’s capacity. | Cybenko formally proved that by introducing non-linearity into a hidden layer, the theoretical capacity of the network became universal. |
| Impact on the Community | Once the dogma was shattered, athletes realized the limit was in the perception of the problem. | Having proven universal approximation, researchers realized the limit was not in the theoretical capacity of the network. |
Cybenko’s socio-epistemological impact lies in having remedied the overgeneralization of Minsky and Papert’s limitation. If Minsky and Papert had demonstrated the inadequacy of the single-layer linear perceptron, the community would have concluded that the entire paradigm was a dead end. Cybenko, by introducing and formalizing the role of nonlinearity in the hidden layer, demonstrated that with nonlinearity the limitation collapses entirely, just as Bannister demonstrated that by changing the training conditions, the 4-minute limit did not hold.
Mass Epistemic Oversight
In the history of science and technology, this phenomenon is often referred to as a “mass epistemic oversight” or an authority effect. To understand how it was possible for a community of highly qualified scientists and mathematicians to fall into this “illusion,” we must analyze the sociological and scientific dynamics of the time.
The book Perceptrons (1969) was a mathematically flawless text. Minsky and Papert never wrote the phrase “no nonlinear network will ever work.” What they did was a fatal combination of two elements.
First, they conclusively demonstrated that the single-layer perceptron (linear structure) could not solve basic nonlinear problems (such as XOR or spatial continuity).
Secondly, however, they introduced a pessimistic conjecture, a speculation. In the final chapter (“Scopes and Limitations”), they hypothesized that extension with multiple layers or with nonlinear elements was “sterilely complex.” They argued that even if a multilayer network could theoretically overcome these limitations, there would never be an efficient method to train it (the problem of giving credit to hidden weights).
The scientific community carried out a cognitive simplification: it merged the mathematical demonstration of the first point with the speculative conjecture of the second, transforming an estimate of difficulty into a sentence of absolute impossibility.
It seems impossible that no one disputed this, but the AI ​​environment in 1969 was dominated by three critical socio-epistemological factors.
First, Marvin Minsky and Seymour Papert were no ordinary researchers: they founded MIT’s AI Lab, the most influential and respected figures in the world in computer science. When two figures of that caliber published a 200+ page book full of formal theorems claiming that neural networks were a dead end, most researchers simply stopped their research so as not to risk their careers.
Second, in the 1970s, AI research in the United States was almost entirely dependent on military funding from DARPA. DARPA was looking for immediate and reliable results. When Perceptrons came out, DARPA drastically cut funding for connectionism/neural network research, moving a significant part of the capital towards Symbolic AI (logical rule systems, strongly supported by Minsky). Without funding, it was nearly impossible for researchers to do experiments to disprove the thesis.
Third, the computers to empirically test deep nonlinear networks did not exist at the time. Without powerful computers and without a clear optimization algorithm (backpropagation had not yet been made famous by neural networks), Minsky and Papert’s conjecture appeared empirically true.
Historiographical Confirmations of the First Misunderstanding
This phenomenon of “mass misinterpretation” is widely analyzed and documented in the historiography of computer science.
Mikel Olazaran (1996), in his sociological study “A Sociological Study of the Official History of the Perceptrons Controversy” (published in Social Studies of Science), demonstrates how Minsky and Papert’s book was used as an institutional weapon to delegitimize the connectionist approach in favor of the symbolic one.
Juergen Schmidhuber (deep learning pioneer and creator of LSTM networks) has repeatedly emphasized in his historical writings how the scientific community of the time culpably ignored that nonlinear and multilayer models had already been proposed (for example, by Ivakhnenko and Lapa in 1965) and that the “condemnation” of Minsky and Papert was completely mathematically unjustified.
Pamela McCorduck, in her celebrated book Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence, describes the publication of Perceptrons as a veritable “ideological attack” that paralyzed the field for 15 years. McCorduck, who directly interviewed many of its key figures (McCarthy, Minsky, Newell, Simon, Samuel), constructs a personal and critical history of AI, showing how theoretical choices, disputes, and personalities shaped the field. In this framework, Perceptrons is presented not only as a technical achievement but also as an act with strong ideological implications and a shift in scientific power, which strengthened the symbolic paradigm and marginalized the connectionist approach.
When Cybenko mathematically proved the Universal Approximation Theorem in 1989, he largely unleashed the “cascade” of research we now call deep learning. It took decades to arrive at modern neural networks, but from a mathematical-formal perspective, the path was already paved.
A Second (Paradoxical) Misunderstanding
This section is more speculative than the previous discussion, which we called First Misunderstanding, and is less well-founded in terms of historical references. I offer largely personal reflections on what happened next.
First, I’d like to speak out in favor of Minsky and Papert. We’ve emphasized the same points several times (the reader will forgive me), and from that, one might deduce some blame for these authors. Nothing could be further from my intentions. Minsky and Papert are two giants who have allowed us to advance our understanding of neural mechanisms. Ultimately, is it permissible for someone to write whatever they want in their book, or not? The “blame” lies with everyone else, who have been misled for decades by questions that aren’t strictly scientific (lacking empirical proof), as we’ve analyzed.
However, we shouldn’t be too harsh. We’re human, and mistakes are part of the equation. The important thing is to recognize them and apply the appropriate corrections, as the story we’ve told teaches us. Fortunately, there are always people capable of challenging dogma, as Rumelhart, Hinton, Williams, Cybenko, Hornik, and Leshno have done.
Unfortunately, we humans never learn. As Oscar Wilde said:
“Experience is simply the name we give our mistakes.”
— Oscar Wilde
Or in a more extended version:
“Experience is a wonderful thing. It enables you to recognize a mistake when you make it again.”
Not only have we not learned the lesson, we have even repeated the same mistake! There is a second, sensational socio-epistemological paradox (in my opinion, this name is fully justified) in the history of Artificial Intelligence.
The effect of Cybenko’s Theorem (1989) has generated a “second misunderstanding” (or boomerang effect). If the Bannister Effect had initially unlocked the field by breaking down the wall of pessimism, its superficial interpretation would have created a new one, slowing the transition to deep learning for years.
The scientific community has once again committed a cognitive extrapolation error:
Cybenko proves: “A single layer is sufficient in theory.” \longrightarrow The community understands: “Using multiple layers is useless in practice.”
Thus, the dogma of shallow networks was born: why complicate life with the difficult optimization of multilayer deep networks if a single hidden layer theoretically “can do everything”?
The misunderstanding was caused by forgetting the fundamental warning of mathematical analysis: the difference between theoretical approximation capacity and computational efficiency.
The Universal Approximation Theorem guaranteed the existence of the solution on a single layer, but at the cost of a potentially infinite number of neurons (an infinitely “large” network). It is well documented in the literature why this interpretation was a practical dead end:
- Curse of Dimensionality: To approximate complex and highly nonlinear functions (such as visual recognition), a single layer requires several neurons that grow exponentially regarding the input.
- Depth vs. Width: A deep network (with many layers) can represent the same complex functions using several neurons that grow only polynomially because it reuses the abstractions learned from previous layers (from simple lines to contours to faces).
This has led to an unjustified, unforgivable accusation against Cybenko. However, with the explosion of deep learning in the 2010s (led by Hinton, Bengio, and LeCun), people have begun to look back critically at the “shallow” period of the 1990s and 2000s.
George Cybenko himself had to clarify several times that his work was a theorem of pure mathematical existence, not a guide to the optimal architecture to be engineered.
In their iconic 2015 paper Deep Learning, published in Nature, Yann LeCun, Yoshua Bengio, and Geoffrey Hinton explicitly explain this socio-epistemological dynamic:
As LeCun, Bengio, and Hinton explain in their 2015 review in Nature, although universal approximation theorems show that networks with a single hidden layer can approximate any continuous function, subsequent work has clarified that shallow networks require an exponentially larger number of parameters to represent certain functions that deep networks can encode much more compactly.
Here is a table summarizing the history of AI dogmas, structured chronologically to highlight the alternation between conceptual building blocks, epistemic oversights, and related breakthroughs:
| Period | Event / Authors | The False Dogma / Wall | What the Community Understood | The Socio-Epistemological Impact | The Turning Point (The “Bannister Effect”) |
| 1969 – 1986 | Minsky & Papert (Perceptrons) | Single-layer linear perceptrons cannot solve non-linear problems (e.g., XOR). | “All neural networks, even multi-layer and non-linear ones, are a dead end.” | 1st AI Winter: Funding freeze (DARPA), abandonment of connectionism in favor of symbolic AI. | Cybenko’s Theorem (1989): Formal proof of universal approximation via non-linearity in the hidden layer. |
| 1989 – 2006 | Cybenko’s Theorem (Universal Approximation) | A single non-linear hidden layer is theoretically sufficient to approximate any continuous function. | “Using multiple layers is useless in practice; a shallow network is enough.” | Shallowness Dogma: Abandonment of deep architecture research due to parameter explosion on a single layer. | Deep Learning Revolution (2006+): Hinton, LeCun, and Bengio demonstrate the computational efficiency of deep hierarchical representations. |
| 2006 – Present | Deep Learning / LLMs (Hinton, LeCun, Attention/Transformers) | Deep networks and parallelization solve problems of perceptual and linguistic complexity. | “Deep learning models are capable of scaling indefinitely toward General Intelligence (AGI).” | Current Paradigm: Massive concentration of resources and economies of scale on large-scale models (LLMs). | Future Breaking Point: The potential discovery of hallucination and logical reasoning limits in current autoregressive models. |
If the original Bannister effect demonstrates how the breaking of a dogma unlocks a cascade of progress, the “Cybenko case” demonstrates the other side of the same epistemological coin: a theorem of theoretical possibility (one layer is enough) can be misinterpreted by the community, transforming into a dogma of architectural laziness, which delayed the advent of deep networks until the empirical practice of the early 2000s. While initially unlocking investment and research in a field that before the proof was considered uninteresting, it shifted this lack of interest from the entire discipline to a part of it. The attempt was to pursue the path of least resistance, uncritically and without solid scientific rationale.
Conclusion
In the mid-1950s, it was considered physiologically impossible (a given for many doctors and analysts) for a man to run the mile under 4 minutes. On May 6, 1954, Roger Bannister shattered the world record by running the mile in 3:59.4. An extraordinary feat that places him among the greatest athletes of all time.
Sport offers us many other examples of feats like this. Consider Bob Beamon’s 8.90 at the 1968 Olympics, or Nadia Comaneci’s uneven bars routine at the 1976 Olympics, the perfect routine, which earned her the first 10.0 in history. Both feats were so extraordinary that even at the Olympics, the world’s greatest sporting event and certainly at the cutting edge of technology, no one was prepared to record the measurements: in Comaneci’s case, the display showed 1.00 instead of 10.00 because it only supported three digits (the audience initially thought the athlete had suffered a catastrophic penalty), while for Beamon, the judges had to measure the jump by hand with a metal tape measure (it took 20 minutes) because the gutter at the edge of the track where the laser was positioned had a maximum run of 8.60 (when the athlete realized the measurement he had made he fell to the ground, victim of an attack of hyperventilation and nervous shock, but this is already history).
Sporadically, a prodigy arises who raises the bar beyond what their contemporaries can achieve, setting a limit only achievable by the next generation, if all goes well. This generation grows with the new goal in mind, which becomes “normal” (in the Kuhnian sense) and does not cause conflict, as it has been known “forever”; it becomes the new paradigm. This, plus training methods that evolve to adapt to the new situation, generates overall growth in the discipline.
Bannister’s case, however, is different. He didn’t set a record that would stand for a long time. On the contrary, he demolished a dogma, a belief, which set in motion a cascade of events that would have been difficult to predict in advance. In the following months after the record, dozens of runners began going under 4 minutes, surpassing Bannister and the previous record that had stood for almost a decade: John Landy surpassed Bannister by running the mile in 3:58.0 just 46 days later.
There are other examples in sports similar to Bannister’s. Consider, for example, when Dick Fosbury astonished the world with his “backward” jump (he won gold at Mexico ’68 with 2.24), which sparked an immediate revolution in training methods for high jumpers.
How can such a phenomenon be explained? It seems like others were waiting for someone to lead the way.
Indeed, that’s precisely the case. The “Bannister effect” has been borrowed from sociology and social psychology to frame the phenomenon that occurs when some innovation unlocks a shared social belief that isn’t objectively justified. A dizzying chain of events then unfolds, revealing that the limit didn’t actually exist. This causes a discipline or technology to grow exponentially. If we were to use a bold analogy from physics, we could say that the Bannister effect appears to be a phase transition.
In this article, we illustrated this effect with a well-known case from the history of technology: the Wright brothers’ first flight. This allowed us to list, with a real-world example, the salient features of the Bannister effect when applied to technological innovation.
In my opinion, Cybenko’s Universal Approximation Theorem posited another “Bannister effect” for connectionism. Our analysis justifies its treatment in this way. Furthermore, it has, again in my opinion, an undoubted epistemological basis. The blockade of the discipline before the proof of the theorem, due to a “dogma of inadequacy,” gave rise, together with other factors, to the historic period that today we call “AI Winter.”
Our epistemological analysis of the Bannister effect as a microrevolution in the sense of Kuhn may find readers disagreeing, although I, personally, believe the similarities are undeniable. This aspect, along with the Bayesian analysis done, deserves further exploration.
This text is a divulgative essay; some historical and epistemological interpretations are proposed by the author and do not necessarily represent academic consensus. I have tried to be explicit when the discussion reflects personal opinions not found in the scientific literature.
Bibliography
- Cybenko, G. (1989). Approximation by Superpositions of a Sigmoidal Function. Mathematics of Control, Signals and Systems, 2(4), 303-314.
- Hornik, K., Stinchcombe, M., White, H. (1989). Multilayer Feedforward Networks are Universal Approximators. Neural Networks, 2(5), 359-366.
- Hornik, K. (1991). Approximation Capabilities of Multilayer Feedforward Networks. Neural Networks, 4(2), 251-257. DOI: 10.1016/0893-6080(91)90009-T.
- Leshno, M., Lin, V.Y., Pinkus, A., Schocken, S. (1993). Multilayer Feedforward Networks with a Non-Polynomial Activation Function Can Approximate Any Function. Neural Networks, 6(6), 861-867.
- Minsky, M.L., Papert, S.A. (1969). Perceptrons: An Introduction to Computational Geometry. MIT Press, Cambridge, MA. (Edizione ampliata: 1988, con nuova prefazione di Léon Bottou).
- Rumelhart, D.E., Hinton, G.E., Williams, R.J. (1986). Learning Representations by Back-Propagating Errors. Nature, 323(6088), 533-536. DOI: 10.1038/323533a0.
- LeCun, Y., Bengio, Y., Hinton, G. (2015). Deep Learning. Nature, 521(7553), 436-444. DOI: 10.1038/nature14539.
- Ivakhnenko, A.G., Lapa, V.G. (1966/1967). Cybernetic Predicting Devices—originally published in russian (1965); translated in Joint Publications Research Service (1966); american edition as Cybernetics and Forecasting Techniques, American Elsevier, New York, 1967 (with R.N. McDonough).
- Pao, Y.-H., Park, G.-H., Sobajic, D.J. (1994). Learning and Generalization Characteristics of the Random Vector Functional-Link Net. Neurocomputing, 6(2), 163-180. DOI: 10.1016/0925-2312(94)90053-1.
- Huang, G.-B., Zhu, Q.-Y., Siew, C.-K. (2006). Extreme Learning Machine: Theory and Applications. Neurocomputing, 70(1-3), 489-501.
- Maass, W., Natschläger, T., Markram, H. (2002). Real-Time Computing Without Stable States: A New Framework for Neural Computation Based on Perturbations. Neural Computation, 14(11), 2531-2560. DOI: 10.1162/089976602760407955.
- Jaeger, H., Haas, H. (2004). Harnessing Nonlinearity: Predicting Chaotic Systems and Saving Energy in Wireless Communication. Science, 304(5667), 78-80.
- Jacot, A., Gabriel, F., Hongler, C. (2018). Neural Tangent Kernel: Convergence and Generalization in Neural Networks. Advances in Neural Information Processing Systems 31 (NeurIPS 2018). arXiv:1806.07572.
- Arora, S., Du, S.S., Hu, W., Li, Z., Wang, R. (2019). Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks. Proceedings of ICML 2019. (Il testo cita anche, sotto lo stesso anno, il correlato Arora, S. et al. (2019), On Exact Computation with an Infinitely Wide Neural Net, NeurIPS 2019, arXiv:1904.11955).
- Blum, A.L., Rivest, R.L. (1992). Training a 3-Node Neural Network is NP-Complete. Neural Networks, 5(1), 117-127. (Originally presented at NeurIPS 1988).
- Olazaran, M. (1996). A Sociological Study of the Official History of the Perceptrons Controversy. Social Studies of Science, 26(3), 611-659. DOI: 10.1177/030631296026003005.
- Schmidhuber, J. (2015). Deep Learning in Neural Networks: An Overview. Neural Networks, 61, 85-117 — ma questa attribuzione è ricostruita, non confermata dal testo originale].
- McCorduck, P. Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence. Prima edizione: 1979 (W.H. Freeman); edizione ampliata: 2004 (A K Peters/CRC Press, Natick/Boca Raton).
- Kuhn, T.S. (1962). The Structure of Scientific Revolutions. University of Chicago Press. [titolo ricostruito: l’articolo cita solo “Kuhn, 1962” senza titolo esplicito].
- Wooten, J.O. (2022). Leaps in Innovation and the Bannister Effect in Contests. Production and Operations Management, 31(6), 2646-2663. DOI: 10.1111/poms.13707.
- Bandura, A. — concetto di autoefficacia collettiva citato senza riferimento bibliografico specifico nel testo. [Il lavoro classico di riferimento sarebbe: Bandura, A. (1997). Self-Efficacy: The Exercise of Control. W.H. Freeman — ma non è esplicitato nell’articolo].
- Campbell, M., Hoane, A.J. Jr., Hsu, F.-H. (2002). Deep Blue. Artificial Intelligence, 134(1-2), 57-83. DOI: 10.1016/S0004-3702(01)00129-1.
- Hsu, F.-H. (2002). Behind Deep Blue: Building the Computer that Defeated the World Chess Champion. Princeton University Press. ISBN 0691118183.
