Sitemap

Brain vs AI: Neuroscience, Math & Philosophy

13 min readMay 3, 2025

--

When the deep state players want to sell you a product, they often create a perception that the product will make the life easy for you and is there to merely serve you. But, in the case of AI, it is more about making you feel dumber and lower your self-worth (and most importantly your effort value in job market). The idea is to get the maximum out of you, yet make you feel lesser.

The myth that is Artificial Intelligence, is ironically exposed by the foundation upon which it has been built.

What is Intelligence?

Intelligence is a fundamental entity, and perception is merely a result.
Intelligence is an art of decision making. Choosing from options is statistical.

Intelligence is characterized by judgement, discrimination, knowledge, ascertainment, will, virtue and detachment.

As per Sankhya Literature, Prakriti, the primordial material matrix of physical universe, when churned by ‘gunas’, the first evolute emerging from it is Intelligence or ‘buddhi’. These ‘gunas’ are catalysts in evolution of material universe, when there is interaction between ‘Prakriti’ and ‘Purusa’. And it is due to specific intermixture and proportionality of gunas, living beings exhibit different mindsets and psychological dispositions.

Etymology of “Intelligence”

The word “intelligence” has its roots in Latin and evolved through several languages before reaching modern English. The Latin word “intelligentia” means “understanding,” “discernment,” or “the faculty of perception”. The Latin “intelligentia” passed into Old French as “intelligence” and only by the 14th century, intelligence entered English through Anglo-French and Middle English. But with time, its meaning expanded to include “information (e.g., “intelligence reports”)” and “superior cognitive ability (e.g., reasoning, learning, problem-solving, etc.)”.

So, what does this shift about the perspective of one single word tell us?

Nature, Time and Reality

What has been is what will be, and what has been done is what will be done; and there is nothing new under the sun”. — The book of Ecclesiastes (chapter 1, verse 9, Revised Standard Version).

Nature is both periodic and perpetual. However, a moment never repeats itself. Here is a paradox to trouble your mind. There is an old Greek saying, which is attributed to Heracleitus of Ephesus, “Upon those who step into the same rivers, different and ever different waters flow down”.

In a periodic system, to predict the future we need a model which truly reflects the past. However, causality among world events is linked to our perception of time. According to Walter J. Freeman, time exists in the material world as a measure of motion, applicable to both living and non-living entities yet he also argues that causality is a construct of human intentionality. Predictions and relationships are constructed by successful ordering of the events according to elapsed subjective time. As time progresses, precision decreases. Let us consider the illustration where brain exhibits “logical illusion”:

“You are driving on a highway, and you suddenly spot a dog crossing the road. You instantly press the brake pedals and the car halts.”

If you will be asked to mentally reconstruct what happened earlier, you (considering you are not a neuroscientist) will most likely say that you noticed a dog (cause) and realized your car will hit the animal fatally and hence applied brakes (effect). However, in a laboratory setting, the explanation would be as follows: a dog appeared on the road (first event), driver applied the brake (second event) and the driver recognized the animal as dog (third event).

So, what happened exactly?

Our brain reacts to an unexpected event (less than half a second) faster than consciously recognizes an object (longer than half a second). Conscious recognition involves dense and distributed neural circuit, where information is broadcast across multiple brain regions. Within 200–300 ms, the brain may extract features like “four legs” or “brown color” in the inferotemporal cortex. However, full conscious recognition (“It is a dog, and it will be dangerous to hit it”) requires integrating these features with memory (hippocampus), context (prefrontal cortex), and attention (parietal cortex), which takes over 500 ms. The synchronization takes time, as it requires coordinating millions of neurons across the cortex. During this time, unconscious reflexes (e.g., braking) can occur, as they rely on faster, simpler circuits.

However, during mental reconstruction of the event, the brain integrates sensory, motor, and cognitive information into a coherent percept, as if it is trying to compensate for the delay. So, what do we conclude from above illustration?

“Our time reconstruction is a consequence of an accumulation of past experience rather than a truthful representation of real time”. — Gyorgy Buzsaki.

Brain and its Complexity

“Life is a sum of all our choices.” — Albert Camus.

Brain is a complex machinery. It is perpetually active, even in the absence of environmental and body-derived stimuli. Here, when we say, ‘perpetually active’, it means that most of the brain’s activity is generated from within and external stimulus only causes a minor perturbation from its robust state. When we add the complex adjective to a system, this usually applies to all the constituents the system is made of and in the case of brain they would be neurons.

Synaptic Transmission (source: www.simplypsychology.org)
# Overview of the Nervous System and Brain Function

The brain is composed of billions of nerve cells called **neurons**. Neurons are interconnected through **synapses**, which are specialized junctions that allow communication between neurons. These connections form a complex network that enables the brain to process information and perform various functions.

## Sensory Input

Our senses (such as **sight**, **hearing**, **touch**, **taste**, and **smell**) receive information from the external environment. This information is transmitted as **sensory signals** to the brain for further processing.

## Neural Communication

Neurons process and transmit information through a combination of **electrical impulses** and **chemical signals**:

- When a neuron receives a signal, an electrical impulse called an **action potential** travels along the neuron's **axon**.
- At the synapses, **chemical neurotransmitters** are released to transmit the signal from one neuron to another.
- Neurotransmitters from the **presynaptic neuron** bind to receptors on the **postsynaptic neuron**, allowing signal transmission.

## Signal Integration and Cognitive Functions

Neurons integrate signals from multiple sources, making complex decisions based on the combined input. This process of integration enables the brain to perform various **cognitive functions**, including:

- Memory
- Learning
- Reasoning
- Emotion

## Structure of the Nervous System

The nervous system is divided into two main parts:

- **Central Nervous System (CNS)**: Includes the brain and spinal cord.
- **Peripheral Nervous System (PNS)**: Connects the CNS to the rest of the body and includes:
- **Sensory neurons** (carry information to the CNS)
- **Motor neurons** (carry commands from the CNS)

## Motor Output and Feedback Loop

After processing and integrating information, the brain sends **motor commands** through the nervous system to control body actions:

- **Motor neurons** transmit commands to **muscles** and **glands**, enabling:
- Movement
- Response to stimuli
- Other bodily functions

This entire system operates as a **feedback loop**:

1. **Sensory input** triggers
2. **Neural processing**, which leads to
3. **Motor output** and action
4. The results are **sensed again**, providing feedback that influences future responses.

Information can be constantly exchanged across boundaries, as evident in complex systems. In case of brain, this is done through the dense neural circuits. The evolution of such a system. A complex system such as brain, is characterized by nonlinear differential equations, because they exchange energy with the surrounding environment (e.g., via sensory input and metabolic processes) and loose energy over time. It is the energy dissipation that drives a dynamic process like neuronal communication. We call these systems as Dissipative Systems, as suggested by the Belgian-American chemist Ilya Prigogine.

Nonlinearity arises because neuronal interactions are not additive. A small change in input (e.g., slight variations in light or brain state) can lead to disproportionate changes in activity due to feedback loops, inhibition, and excitation. Past experiences modify synaptic connections (via Hebbian learning or long-term potentiation). The system itself is chaotic. As Gyorgy Buzsaki subtly puts it in his book The Rhythms of the Brain, “The spatiotemporal trajectory of neuronal activity depends not only on the constellation of light impinging on the retina but also on the perceiver’s brain state and past experience with similar physical inputs”.

What Buzsaki suggests is that even if the same stimulus is presented multiple times, the neuronal trajectory is unique due to the brain’s nonlinear dynamics. He further suggests that complexity can be perceived as non-linearity and from non-linear equations unexpected solutions emerge. The reason being that it is not easy to predict the dynamic behavior of a complex system or even it can not be deduced from the behavior of lower-level entities (constituents the system is made of), as they are themselves complex in nature and independent at many levels.

But independence of constituents at many levels does not mean that the behavior of a system is defined by the sum of local interactions. The emerging behavior of a complex system at one point of time is determined from the manifold interactions of various constituents of the system, and even a slight change can significantly change the resultant behavior of the system.

What is the “modern age AI”?

“Any man could, if he were so inclined, be the sculptor of his own brain.” — Santiago R.y Cajal

The modern age intelligence, often termed as “Artificial Intelligence” has been built upon the Perceptron model, proposed by Frank Rosenblatt in 1958.

Press enter or click to view image in full size
Perceptron Mathematical Model

A perceptron is a fundamental building block of artificial neural networks and serves as a simple model of a biological neuron. The foundational principle of a neural network mirrors the interconnectedness and information processing observed in human neurons. Instead of electrical impulses, each neuron in an artificial neural network stores a numerical value representing its strength and capacity to transmit information to neighbouring neurons. The architecture of a neural network is organized into layers. The initial layer receives input data, while the final layer produces output or predictions. Neurons in each layer are often fully connected to neurons in the previous and subsequent layers. Information flows sequentially from the input layer through intermediate layers to the output layer.

Press enter or click to view image in full size
A multi-layer perceptron based Neural Network

A perceptron receives multiple binary inputs, denoted as x1, x2, x3, …, where each can be either 0 or 1 and produces a single binary output. Additionally, there are corresponding weights w1, w2, w3, … associated with each input. The weights represent the importance or significance of each input in influencing the perceptron’s output. The perceptron computes a weighted sum of the inputs and weights. This is done by multiplying each input by its corresponding weight and summing up these weighted values.

The perceptron then compares the weighted sum to a threshold value. The threshold is a parameter of the perceptron that determines the point at which the perceptron will “fire” and produce an output of 1. If the weighted sum is greater than or equal to the threshold, the perceptron outputs 1; otherwise, its output is 0.

Here, w denotes a vector of real valued weights, m is the number of inputs to the perceptron, and b is the bias. In addition to the weights and inputs, the perceptron model includes a bias term. The bias serves as an offset or a threshold that adjusts the decision boundary of the perceptron. It allows the perceptron to learn and generalize patterns that may not pass through the origin (0,0) in the input space. When combined with the weighted sum of inputs, the bias influences the point at which the perceptron transitions from producing output 0 to output 1. A positive bias makes it easier for the perceptron to output 1, while a negative bias makes it harder.

Neural networks rely heavily on linear algebra for their core operations. The input to each layer is typically represented as a vector or matrix, which is multiplied by a weight matrix to produce a weighted sum. For a single neuron, this is expressed as:

Press enter or click to view image in full size

In deep learning, higher-dimensional tensors generalize these matrix operations across layers and batches. While the weighted sum in each layer is linear, neural networks introduce non-linearity through activation functions applied to the linear combination (z). These non-linear activations allow neural networks to model complex, non-linear relationships in data, enabling them to solve problems like image recognition or natural language processing, which linear models (e.g., linear regression) cannot handle effectively. In traditional feedforward neural networks (e.g., multilayer perceptrons), information propagates sequentially from the input layer through hidden layers to the output layer. Each layer processes the output of the previous layer:

Press enter or click to view image in full size

All that glitters is not Gold

If we look at the mathematical discussion presented above, it will appear that Neural Networks aka AI cover all the grounds by implying non-linear activation functions on top of linear algebra operations. But we skipped the most significant part for a reason.

What is the fuel of AI?

Data.

An AI model heavily relies on data provided to it during its training stage to form patterns and generalize behaviors. Say, we feed the network with two simple 3x3 images with following patterns: Pattern 1: and Pattern 2:

Press enter or click to view image in full size

And then we present the network with an incomplete version of Pattern 1:

The network will use stored patterns to recreate/complete the image. Let us understand how this works.

A node (also called neurons) in the network is a binary unit, which means each neuron can only be in one of the two states, active or inactive. Each pair of neurons in a fully connected network is connected by a weight (represented by weight between ith and jth neuron.), which represents how their states are correlated across different patterns. These correlations are captured through a learning rule (like Hebbian learning) and stored in the weight matrix.

To put it in simply, the weights are adjusted according to the correlation between the states of neurons in the patterns. So, if neurons ith and jth tend to be in the same state in many patterns, the weight between them will be positive, indicating that these neurons should support each other’s states. However, if they tend to be in opposite states, the weight will be negative, an indication to adopt opposite states.

The weights across the network are symmetric and there are no self-connections. The idea here is to use the values of these weights to determine how the neurons collectively behave, and by adjusting these weights, we can store patterns in the network. Each pattern (e.g., text, image, etc.) is encoded by adjusting the weights in such a way that the neurons in that pattern have strong correlations, making the pattern an attractor state of the network, a state that the network naturally falls into when given similar or incomplete information. The weights are adjusted so that the pattern represents a low-energy configuration.

The idea here is that when the network is given an input that is close to the stored pattern (even if some neurons are missing or incorrect), the network will naturally gravitate towards the full pattern during recall. The dynamics of the system will push it toward this low-energy state, effectively reconstructing the stored pattern.

In a nutshell, when we present a partial or noisy version of a stored pattern, the weights (which store the correlations between neurons) guide the neurons back to the correct configuration. Neurons that are supposed to be positive in the full pattern will get positive inputs from their connected neighbours, and those that are supposed to be negative will get negative inputs. The stored weights act like a memory of the patterns.

The system is designed in such a way that each pattern corresponds to a minimum in the network’s energy landscape. A high energy state corresponds to a random, noisy, or incomplete pattern where the neurons’ activations are far from one of the stored patterns. On the contrary, a low energy state corresponds to a stable state, or the attractor state discussed earlier.

When an incomplete pattern is presented, the dynamics of the network pull the neurons into one of these minima, restoring the original pattern. A similar analogy can be thought of in terms of a ball rolling down into a valley in a landscape. As the neurons update their states, the network slides down towards the nearest valley. Once it reaches the bottom, it stays there. Here, the bottom corresponds to a low energy state (or one of the stored patterns).

What’s the point?

Remember, when we discussed about Brain being a complex system and its behavior being represented by non-linear neural networks?

We explicitly mentioned that “most of the brain’s activity is generated from within and external stimulus only causes a minor perturbation from its robust state”.

The point is, even if you ask an AI agent to generate data, it will generate data from the patterns it has observed so far and perform different permutation and combination on top of it. It does not have the ability of creation.

Brain does not perform Singular Value Decomposition on the data matrix available, which results in loss of data in lower dimension, but it merely works on the rule of association. It processes data in form of neural oscillation ranges, which help it perform various cognition activities. A better analogy would be a Fourier Transform. It breaks down complex sensory, cognitive, and emotional signals into frequency components (e.g., neural oscillations in delta, theta, alpha, beta, gamma bands). Relevant signals are amplified, and irrelevant ones are suppressed. And do take a note, a slight deviation in just one signal, may change the resulting behavior quite significantly.

On the contrary, you can ask the AI agent to do some tasks by giving a thorough instruction list and it may perform the task more effectively than a human, but to get the task effectively done, the human instructor must have the knowledge of what has to be done in a detailed manner. An AI agent can do generalization by doing a linear transformation from a higher dimension to lower dimension and try to fit a pattern, but it won’t be able to create something out of the generalized lower dimension. That is the reason why LLMs keep extending their context window in every update so that more and more instructions can be explicitly fed to it. Do not believe me?

Play with the available billion parameters model in the market and ask them to perform coding tasks in the third level of your program and you will understand the point.

You can either realize your self-worth and put an effort to restrain the push or just succumb under pressure and forget who you are and, in the process, loose everything.

--

--

Reeshabh Choudhary
Reeshabh Choudhary

Written by Reeshabh Choudhary

Software Architect and Developer | Author : Objects, Data & AI.