In short, after after training AI on an extraordinarily amount of human cognitive output, we are now facing the possibility that our ability to train by working on hard problems will be slowly stripped away at least in some domains.
It’s like someone offers to build mag lev gym weights. It’s very cool that I can now lift the 500 pound weight with a finger. But what will I do when there’s no power and 500 pounds to lift?
Of course, cognition isn’t a single outcome problem like weight lifting. But we build cognition not wholly unlike how we build muscle: one needs resistance. Otherwise I’m not at all confident we “learn” in any depth.
You know, I pooh poohed his theories. Still do in sum. But there’s a there there that’s building.
Damn Deepak Chopra and his ilk of idiots for making any conversation of quantum mechanics and biology tinged with pseudoscience. Hopefully we’ll keep getting experimental evidence as we go that it’s not at all absurd to consider quantum effects in biology.
That said, those effects are going to look nothing like sustained coherence for long periods of time.
How much of this, I wonder, is a function of the fact that our circadian rhythms get less robust with aging. The circadian clock hugely influences learning and memory processes, gating when you can learn and how much, and shaping the storage and recall also.
We know that with age the amplitude of these rhythms can decrease, as can the synchrony between cells.
This kind of mid-management seems to have at least some circadian component, and it would have been great if they looked to see if the effects were equally bad at all times of day, and knew the chronotype of the participants to use as a reference. I’d love to see if every test participant was tested at their cognitive peak, too.
The brain does not have the Von Neumann bottleneck. Unlike most current digital systems, the brain doesn’t have a separate memory registry it needs to pull from.
Engrams, that is, the physical trace of a memory, are not stable through life. They start out in the hippocampus, but as the stimulus recedes in time without reinforcement, it moves away.
No evidence exists though that the memory is encoded in one set of cells. This spatial segregation of memory is the worst hangover from the “brain is a computer” analogy. Even if it is, why in the world would it be like our digital devices which specifically have the Von Neumann bottleneck? In biology, memory and processing are not segregated.
There’s growing evidence the memory is much more distributed over the network, and is recomposed based on salience overlap with a new stimulus.
Another factor to keep in mind is circadian rhythms. There’s growing evidence for how much the memory system and timekeeping system overlap, at a molecular level. Every neuron (and other cell) has an intrinsic clock that ticks at roughly 24 hours, and continues to do so even in total darkness.
When you encode the memory has a lot to say, based on your chronotype, on how and how well you will remember it. Same with learning: there’s a time of day based variation.
Sleep, and dreaming, is when these memories seem to get replayed and critical features and connections are incorporated into the system and its regime, awaiting the right triggers to access a state similar to when the memory formed.
I’m stitching across a lot of different research, and I want to be clear many aspects of this system are not yet fully worked out.
But what we do know points to a system that works with different physical and algorithmic priors, and the dynamics are sharply distinct from current digital computers.
Hijacking your post with a dubious segue because I’m itching to bounce these thoughts off somebody:
I’ve been consuming a lot of talks / writing recently about “enactive” pictures of how our brains function. From what I gather, recent studies have called into question the entire idea of real world concepts being “represented” by an area of the brain at all. While it’s true that atandard fMRI-style snapshots of brain activity are semi-stable over the course of a short experiment, it’s not true over longer timeframes. The response to the same stimulus will change over time. They refer to this as “representational drift” in the literature, and some people are using this to bolster theories of mind that they consider non-representational. They instead emphasize the brain as a kind of dynamical system that learns to “resonate” with the world to pull itself back into homeostasis. The focus shifts away from facts and memories as data, and sees neuronal plasticity more as a mechanism for tuning the brain’s resonant frequencies. This obviously places high importance on the spiking, recurrent nature of actual neurons, as opposed to the neurons-as-functions / back-propagation / ML approach.
My mind’s not made up on how interesting and revolutionary this approach is / isn’t. The distinction seems to be about whether learning is more like “writing to disk” or “tuning a PID controller” - but in either case, the world is leaving a stateful imprint on your brain that will impact how it processes future data. Is that important to understanding how brains work, or is it just semantics?
Some years ago, I heard a retired Lutheran priest ponder, in radio, about the concept of a prayer, and whether, as an edge case, an unborn child could be able to pray, with the child of that age having no understanding on the required concepts. This is apparently one of those questions that people ponder under the umbrella of philosophy of religion. His conclusion was that prayer was about harmonizing one's self with the universe. It's a beautiful thought, (and I think religions, of all kinds, are really good at providing the fertile ground for thoughts like these).
It's akin to how the saying goes, that when we argue, we need to reach the same wavelength where the other one is to reach an understanding.
In communication studies or sociology or linguistics or one of those fields, there's the idea that communication is about making pacts about meanings, and finding the common ground to understand and delimit the message.
In a sense, from that basing, one could argue that understanding can be seen as an act of harmonizing. I think there's something universal in it.
For a lot of people, I think religion provides this basis or promise of harmony with the world, in a simplified way that stipulates what you should or shouldn't do and how things work, and those people find comfort in having these clear "rules"; we want to be able to rely on our understanding of the world, so that we can act within it with confidence, or faith. Ultimately, the "god" that is present in many religions is synonymous with this absolute truth of how the universe works, it's an unattainable standard that we all try to aim for, intuitively and knowledge-wise, or seek wisdom from, and I believe the intention of prayer is to have a way to get closer to it, to tune the self to the god in order to sense and know this absolute truth more clearly. I reckon that conceptually, prayer and meditation are very similar in this sense, but with a focus on different areas of the self.
If "god" is synonymous with the absolute truth of how the universe works then it can't address moral issues at all, since you can't prove an "ought" like "people shouldn't stab people" from an "is" like "stabbing hurts people and may kill them". You need to add extra axioms like "hurting people ought to be minimized".
The universe (and its intrinsic functions) holds no special value for people or trees or life in general. The best we can say is that it seems to have a slight preference for matter over antimatter for some reason.
All our moral codes and social norms we live by are purely human invention, derived from things that happen to work to produce a somewhat functioning society.
> All our moral codes and social norms we live by are purely human invention
I'd say much of it is selected for by evolution. e.g. Most animals don't fight or kill more than they need to since this is not optimal - risk of personal injury and depleting a food resource (there are exceptions of course such as the fox in the henhouse), and this is so universal it seems it has to have a genetic basis.
>All our moral codes and social norms we live by are purely human invention, derived from things that happen to work to produce a somewhat functioning society.
Religious people would disagree, and the evidences are all to see, hear and ponder from their sacred holy books. But the main questions are that which holy books are truly sacred, meaning that purely and utterly God words. Otherwise it's human innovation/exaggeration or something in between (i.e corrupted God words).
Based on these holy books, religious people adhere and follow their prophets teachings and actions to the best of their ability since the prophets are their fellow human beings and not angels send from the heavens.
Ooh thanks for the hijack! It’s so much easier to talk to someone who isn’t stuck in a picture of the brain from the 1980s.
Yes, I fall more towards the camp that a lot of our cognitive models, built from times when we didn’t have the resolution of understanding we have of the capacity of even a single neuron, and before we knew how astrocytes played a role, suffer from being an abstraction describing an abstraction. They are not tethered in the dynamics of the molecules and cells that give rise to the behavior, but rather from an interpretation of observed behavior.
This paper from the field of chronobiogy is one I’d recommend that helpfully contrasts this:
>In circadian research, the models are not proposals regarding the basic architecture of circadian mechanisms; rather, they are used to better understand the functioning of a mechanism whose parts, operations, and organization already have been independently determined. In particular, circadian modelers probe how the
mechanism’s organized parts and operations are orchestrated in real time to produce dynamic phenomena—what we have called dynamic mechanistic explanation.
And what you’re describing, the enactivist description of cognition, (and 4E cognition more broadly as a framework), is one way the neuroscience community is trying to move past these issues.
Few things that give me confidence these are the right track:
1. Circadian rhythms are evolutionarily ancient. Bacteria have em. Plants have em. But different molecular tools shape very different clocks, though the same 24 hour cycle is being tracked.
2. The way these rhythms are generated is not through some central system that broadcasts the information to other regions. Instead, it’s instantiated in every cell in the body, and the behavioral rhythm is due to the synchrony between cells. Resonance absolutely plays a role, and has been well documented. The brains role, via the suprachiasmatic nucleus or SCN, is to orchestrate this synchrony, but it is not the source of the rhythms.
3. This slow rhythm definitely regulates cognition (time of day effects in learning, memory formation, recall etc are well documented), but turns out, the molecular mechanisms by which the clock responds to light hugely overlap with the molecular mechanisms of learning in the synapse, and even more recent work has shown clock proteins are actually in the synapses and synaptic activity affects the clock.
All this points to nested oscillators with cross frequency coupling, and even better, because this is all grounded in actual molecular dynamics, there’s plenty of falsifiability. The phase amplitude links are best established for the faster rhythms, the famous “brain waves”. Highly recommend György Buzsáki‘s work on this:
What’s missing is going down into lower frequency rhythms, and testing how exactly they all couple. We have a lot of the pieces, but no single experimental paradigm that has looked at the full sweep over different times in the same organism. It’s not easy to do, but we’ll get there.
Clock disruption, depending on how you do it, has huge impacts on time perception, cognition, memory, aging AND consciousness. As that data and evidence gets more and more saturated, I hope we see more studies account for chronotype and the internal dynamical state of their test subjects when assessing outcomes.
Obviously I’m biased (also did chronobiology in school), but hopefully I’ve left you curious. Happy to answer more questions all this may have set off.
Thanks for your response, very intriguing! I have some controls background, and there’s something tantalizing about the idea that perhaps we need to be looking at the brain in frequency space, as it were. Are you aware of reservoir computing and do you see it playing a part in this?
Yes I’ve come across reservoir computing. As a neuroscientist, it made me sit up and take notice.
I’d say that I feel there’s homology in language. What reservoir computing says about the efficiency benefits of having a fixed but tunable dynamics to use as an underlying reservoir feels very adjacent to how I intuitively think of brain function.
The key thing from the circadian field you’ll appreciate:
The biological clock is a limit cycle oscillator. You have a bunch of chemical reactions that have negative feedback and some feedforward arms, and together they create a dynamical 24-regime. About 40-60% of the transcriptome of any given cell shows circadian dynamics.
Now this gives you phase, and an internal temporal reference for all your functions. In chronobiology, you call this the organisms subjective time. The system is chemically partitioned not just physically but over time, and behavior results from the dynamical interactions underneath which are concerned with anticipating solar and lunar periodicities in the environment, since those are so very common and determinative to fitness in many niches.
Note the fact that it is subjective time but has an objective description. However, external measurement without the background of the chronotype accounted for will thing of a lot of variance as “noise”.
My own philosophical conclusion has been that this is the source of our confusion with consciousness. We don’t account for the internal causal order of events, which are timed, and with cross frequency coupling and phase-amplitude linkages begging to be worked out with real world data.
> The brain does not have the Von Neumann bottleneck
Obviously not, which is why I didn't say it did!
However, if you want to identify where long-term memories are stored, then that is in the cortex, but it should go without saying that this doesn't make the cortex the storage component of a von-Neumann architecture!
> No evidence exists though that the memory is encoded in one set of cells
I'm not sure what you are trying to say.
Memories are presumably stored as embeddings - a distributed representation, and an episodic memory may well be stored as "chained together" episodic "scenes/chunks" where each chunk recalls the next.
However, a distributed representation isn't the same as a holographic one, and any redundancy may well still be localized within given cortical columns, so I think you may be wrong if you are saying that individual memories/chunks are not confined to one set of cells (some localized neural assembly such as a cortical column).
> Obviously not, which is why I didn't say it did!
But you sneak it into your assumptions on what a memory must be.
> However, if you want to identify where long-term memories are stored
And why do you assume there is a specific “where” for the memory?
> then that is in the cortex
There’s good deal of evidence disproving this in the way you’re stating this. The cortex is involved in sensing, and yes, the sensory information associated with a memory will recruit appropriate cortical cells. This doesn’t mean the memory resides in the cortex. And detailed episodic recall keeps recruiting the hippocampus even for old memories, which is odd if the memory were somehow only in the cortex.
> I'm not sure what you are trying to say.
Let me restate what I’m saying then:
There are four claims bundled together in the way you were describing biological memory: that a specific ensemble is activated when a given memory forms; that it’s the same ensemble that gets activated over time when that memory is retrieved; that it’s spatially compact, like a column in region of the brain; and that this ensemble it’s dedicated to that memory, or similar memories .
The first is well supported. An engram, a network of neurons, is indeed activated when a memory first forms, and gets stabilized due to repeated stimulus. Re-activating these neurons in a different context can make the subject (a mouse) behave as it would if it had contextual signals to evoke said memory.
However:
1. This engram is not in one particular part of the brain. There’s a cortical part that overlaps the sensory regions that were involved. But plenty of other regions are part of the engram
2. There’s turnover, over the course of weeks, when the specific cells involved in the engram drift, while the behavior remains stable.
3. The same synapses participate in many memories.
> Memories are presumably stored as embeddings
No. Let’s consider songbirds, which are an excellent worked out example (in an animal without the complex columnar cortical architecture mammals show, by the way).
What’s learned is a temporal sequence, with neurons in a nucleus in their brains each firing one brief burst at a fixed point in the motif, so the content of the memory is its dynamics rather than any value. The circuit that evaluates the match against the tutor template is the same circuit generating the output being evaluated. Song degrades overnight during sleep replay and recovers the next day, and in seasonal species the song nuclei change size across the year with neurons added and lost while the song persists. There’s no read that leaves the item untouched, no persistent address, and no substrate holding still. “Stored as” imports all three.
And it goes below the neuron or synapse. Hearing a tutor song drives immediate early gene expression that habituates with familiarity, and singing drives large transcriptional changes in the song nuclei that differ by social context for the same motor output. Since transcription runs on minutes to hours and the proteins turn over, any persistent state has to be actively regenerated rather than deposited.
TLDR: the memory isn’t a static store. There’s no persistent “location” for it, distributed or otherwise, though specific locations can be in the chain that’s activated for retrieval/production. Instead, memory, over time, is driven by a dynamical regime that adjusts its dynamics to account for the temporal pattern in the salient stimulus.
Nothing, down to the epigenetic changes in the chromatin of these neurons, can be seen as “the” location of “a” memory, especially over time.
> if you are saying that individual memories/chunks are not confined to one set of cells (some localized neural assembly such as a cortical column).
My theory was that memory had some form of static store, something molecular like methylation. I hadn't considered something dynamic/temporal like you describe.
Would delay line memory be a fair analogy? Information stored in a delay line memory never stays in one place. I think you are saying biological information is "stored" in the amplitude and phase of oscillations of our cells. Is that correct?
I'm still inclined to believe there is some form of static store, which would be required for inherited memories. Things like our (and many other mammals) ability to recognize emotions in others. Somehow our gametes encode what a happy, sad, or scared face looks like. I instincts as inherited memories.
If someone's hippocampus is destroyed, they lose the ability to form new (episodic) memories, and may lose some more recent old ones, but they certainly do not lose older ones. This is basic knowledge.
> and that this ensemble it’s dedicated to that memory, or similar memories
No - that's the exact opposite of what I said. My whole point was that a cortical column is NOT dedicated to a single memory (we'd run out of memory!), but rather acts as an embedding space containing many (sparse) embeddings.
Due to the size of the embedding space and sparsity of individual embeddings, there is little chance of much overlap between embeddings and therefore associative recall is reliable. When there are too many memories stored using the same set of neurons, then there will be non-trivial overlap between embeddings (this is the definition of "too many" / "full") and associative recall becomes unreliable.
Note incidentally that this explanation holds regardless of whether distributed embeddings are stored in a more localized area (or areas - visual, auditory, etc components) or more globally distributed. At the end of the day evolution has equipped us with a right-sized brain, and an individual that outlives the useful life it is adapted for can expect to experience memory failures.
I'm not sure why you bring up bird brains, and specifically bird songs(!), but FWIW it seems that their short term memory likely works similarly to our own in as much as it is based on the hippocampus, with a very strong correlation between bird hippocampus size and memory capacity (ability to memorize 10's of thousands of hidden seed locations in some species). Some birds such as crows certainly have long term memory where I'd guess those may have migrated to their pallium, but we're discussing human memory here (or at least I thought we were).
> If someone's hippocampus is destroyed, they lose the ability to form new (episodic) memories, and may lose some more recent old ones, but they certainly do not lose older ones. This is basic knowledge.
Yes, like basic reading without digging into details. You must have heard about HM, since you’re saying all this. But here’s the facts:
When H.M.’s remote memories were probed carefully, they turned out to be gist-like and semanticized, not vivid re-experiencings of specific events. Here’s the paper:
Only semantic memory of the episodes can be said to be “cortical” (though please note, lack of hippocampus doesn’t mean lack of other brain regions…). Rich recall absolutely does require the hippocampus.
Once again, please try not to “spherical cow” the complexity of the brain to try and fit it into your analogy to digital computing. You will get an underdermined model that will miss the subtleties, and lead you to claims that are poor fits for the reality.
As for why I brought up bird brains… there part of the same evolutionary web. Is there some reason you want them excluded? They’re a well studied model for a fairly complex memory task, using substantially smaller neurons more densely packed in a different architecture than mammals.
In cognitive science, as in computer science I’d imagine, it’s useful to look at the full picture before making strong claims.
The bird case is interesting because the region of interest is a nucleus, rather than cortical columns, and actually has well documented structural variations in size, as well as gene expression, over the seasons, while the memories are forming.
If your model is correct, it needs to account for those facts.
There is 2006 paper "Polychronization: Computation with Spikes" E. Izhikevich. that describes one simulation they have done in silico and it explains how exactly distribution of activity happens and why you can't run out of memory. Basically the small group of neurons, say 10, can represent much larger amount of information say 1000 because they can fire in different orders, that is what they call poly-synchronous activity.
> If your model is correct, it needs to account for those facts.
Let me make it simple for you.
We have a finite number of neurons in our brain, as do birds, and our brain is attempting to store an ever growing number of memories in those. And, no, this is not a digital computer (is your reading comprehension really so bad?).
Nobody, including you, knows exactly where different types of memory are stored, and for my argument it makes no difference. What does make a difference is how they are represented, which I am suggesting is sparse embeddings.
So far, you've been ignoring my actual argument and instead responding to various strawmen of your own making, so it's not clear if you even understand what a sparse embedding is.
If you do understand, then it should be obvious that it makes no difference whether the neurons comprising this embedding space are in the hippocampus, cortex/pallium or anywhere else. Clearly you do NOT understand, since you bring up bird brains (pallium vs cortex) and want to argue about location of storage (hippocampus vs elsewhere) as if it made a difference to MY argument.
My argument (if you care to respond to it, which so far you have not) is that when sparse embeddings have little to no overlap, then associative recall by a similar pattern will work reliably, but when multiple embeddings have too much in common then recall will suffer as multiple embeddings will match.
Hint: if you think this has anything to do with digital computers then you have misunderstood and need to go back and re-read more carefully, or google for any terms you do not understand.
Sliding past the mistakes pointed out, shifting goalposts and trying to recover I see.
Let’s say I’m a complete moron and don’t know what a sparse embedding is.
Pretty please, can you define it for me and then tell me, in detail, where in whatever region of the brain you think this is going on… how is it going on?
Explain how “memories must be stored as embeddings with single multi-neuron assemblies (cortical columns?) storing multiple embeddings as a kind of contents-addressable memory”
You have moved past the cortical column. But still seem to be insisting it’s a bunch of neurons, somewhere… or has that also conveniently changed? Whatever your current position is, please go ahead and explain what components of what cells or otherwise are involved in this process you’re describing.
An embedding space is a (typically) high dimensional space that has enough dimensions such that examples of some type of entity (e.g. faces, words, or thoughts) can be represented as points in that space, positioned such that they are nearby to other entities with which they have things in common.
An entity embedding doesn't need to use all the dimensions of the space it is positioned in - some dimensions may be unused (sometimes represented as a coorrdinate of 0 in that dimension). These are called "sparse" embeddings. For example, an LLM's tokens are represented as embeddings in what is typcially an approximately ~1000 dimensional space, but start out as sparse embeddings just representing a short letter sequence (but then go on to be transformed/augmented with additional information and so become less sparse).
As an example, let's say an embedding space has 10 dimensions, then a couple of sparse embedding examples could be:
[0 0 1 0 0 1 1 0 0 0]
[1 1 0 0 0 0 0 1 0 0]
These two embeddings have no overlap (where both are non-zero), and the more dimensions you have the more likely it is that two random sparse embedding will have little in common.
Embeddings are used in many types of artificial neural networks, not just LLMs, for example face recognition networks, where they are trained such that similar faces (multiple photos of the same person) are close together in the embedding space, and post-training you can then "look up" any arbitrary photo (in the training set or not) by embedding it and seeing what is nearby in the embedding space, which will be similar looking faces.
Presumably real neural networks used embeddings in a similar way, since, for example, it obviously requires many neurons to represent the many differences between different faces, and there is going to be overlap between the neurons used to represent multiple faces (this is not a computer with one storage location for face #1, and a different location for face #2).
A neural network, real or artificial, uses groups of neurons (e.g. a cortical column) to represent an embedding space, with each neuron corresponding to a dimension. A single group of neurons (column) can store multiple embeddings (e.g. faces) represented as different activity patterns (which neurons are firing), and if these are sparse embeddings then the firing patterns corresponding to different memories stored in the same column will have little in common.
Now, I don't know how you believe associative recall is implemented in the brain - how does someone's voice, or half obscured face, recall their entire face, so feel free to imagine it as implemented however you will, but I'd suggest that in an assembly such as a cortical column that when a set of synaptic inputs are triggered the assembly as a whole will learn to reactivate the entire pattern when only part of the original set of synaptic inputs are triggered, and this is the basis of associative recall. There are papers that suggest exactly how this may work given the cortical column microcircuit.
So, with all that said, the suggestion I was making for why (or at least one reason why) memory degrades with age, with memories blending together, is that with a finite quantity of "storage" (cortical columns) you will eventually be storing so many memories (absent a deliberate forgetting mechanism) that there will inevitably be overlap between the sparse embedddings, and this associative recall will therefore not cleanly recall individual memories but rather recall blended memories according to what they have in common.
Obviously some types of memory are at least initially stored in the hippocampus, so no reason to focus on cortical columns, but I expect the use of embeddings is universal.
Ok great, thanks. Now you’ve brought up facial recognition, and that’s actually a great example to show where the analogy breaks, and the idea of a few sparse cells encoding specific faces has been conclusively disproved:
Faces live in a ~50-dimensional continuous space (25 shape axes, 25 appearance axes). They measured about 205 neurons across 2 macaques (human studies have substantiated much of this, some from the same lab), and the key thing is: every neuron participates in every face.
The paper shows faces are embedded, but as points in a dense linear space where neurons are axes, not as sparse activity patterns where neurons are on/off slots.
The mapping between the neuronal activity and the facial structures is invertible. Record these same cells, and their firing pattern can be used to reconstruct the face. Or, if you generate a novel face, you can predict the firing rates of these neurons for it. As far as I understand, this doesn’t work for sparse embeddings.
Some cells carry the shape coordinates and others carry the appearance coordinates, in a heirarchy.
There’s an embedding space, yes. But that space isn’t defined by a network of “on” and “off” neurons. The embedding space is instead constructed by the activity of neurons, and the differences in activity distinguish the faces, using the same set of neurons.
And distance in the ensemble activity of these neurons tracks the distance in face space.
If faces use sparse embeddings, you wouldn’t expect similar faces to evoke similar activity would you? Yet that is exactly what this paper shows, and the same has been shown in the human brain for faces.
There are places where it’s sparse activity of a subset of neurons that maps to specific memories. What you’re describing is what you’d see if you look at how the dentate gyrus (part of the hippocampus) handles your memories in the same location.
But even there, the sheer number of cells makes this combinatorially such a vastly overdetermined system for a lifetime that there’s no capacity limit of the kind you’re describing. Even 1% of these cells lighting up for a specific memory leaves you with so many possible combinations that you’d have to live for a few million years to be in the right scale to at least being to talk about capacity issues.
The brain just isn’t capacity limited by the number of neurons the way your intuition is pointing you.
If you say this has nothing to do with the Von Neumann bottleneck or computational functionalism, fine, but how do you square that with the statement below, which you made further down responding to another post?
> but it's hard to imagine that all of the classical chemistry, let alone quantum, details are important. It's necessarily built out of chemistry, but selection is happening at the level of behavior - presumably depending only on a much higher level set of abstract capabilities (ability to learn, etc), not the exact details of chemistry.
The success of LLMs, a crude prediction mechanism built atop a crude ANN, does tend to support the idea that low level details don't matter. Timing will matter if we want to go beyond LLMs to AI that can learn time-based things and not just sequence order, but how much else will matter remains to be seen!
It’s really odd to see these two paragraphs, because the second actually tells you why your first is wrong.
Simply put, the biochemistry is timed. I urge you to study how temperature compensation of circadian rhythms is achieved. That anticipatory function goes all the way down to the molecular level.
It might go down to the quantum level too. In birds, magnetoception depends on a protein called cryptochrome IV, which uses a singlet born, entangled radical pair of electrons to sense the very weak magnetic field of earth.
Now cryptochrome 4 is bird specific and mammals don’t have it. Other cryptochromes are critical clock molecules. And the whole shebang of these evolved initially to be sensitive to blue light and repair DNA.
Try as you might, you can’t separate out the deep linkages from the molecular to the behavioral in biology.
Trying is perfectly fine for stuff like language models. But if you’re going to build models with internal time, best of luck if you ignore the molecular and the energetic considerations. Time emerges from the ground up, in biology, as in physics. Doubt we’ll get a free ride with computers.
> A neural network, real or artificial, uses groups of neurons (e.g. a cortical column) to represent an embedding space, with each neuron corresponding to a dimension.
You:
> The paper shows faces are embedded, but as points in a dense linear space where neurons are axes, not as sparse activity patterns where neurons are on/off slots.
So you are saying that neurons are axes (aka dimensions), exactly as I just said!
> If faces use sparse embeddings, you wouldn’t expect similar faces to evoke similar activity would you?
Yes, of course you would, because that is precisely how embeddings work, and how you recognize someone even though their head is turned or they are wearing a baseball cap or whatever.
This is the ENTIRE point of emebeddings and why evolution has discovered them as a way of representing things and a way to recall them. You may have seen someone a million times, and yet the sensory patterns your visual cortex is fed are likely different every single time because they are not in the exact same orientation, making the exact same facial expression, with the exact same haircut, etc, etc, etc.
To your brain these are merely similar inputs, similar faces, but there is only so much facial variation between individuals, and if the input is similar along dozens or hundreds of axes of variability (i.e. close in embeddign space) then it is alomst certainly the same individual.
Note that "recall keys" (embeddings) are typically sparse even any stored embedding is not, since the face you are looking at may indeed be turned left or half obscured, and this partial/sparse pattern needs to recall the full one.
How can you be a neuroscientist, or even self-identify as one, if you are not already familiar with things like embeddings, and are making such basic 100% wrong assumptions as "you wouldn’t expect similar faces to evoke similar activity" ?!!!
> Now you’ve brought up facial recognition, and that’s actually a great example to show where the analogy breaks, and the idea of a few sparse cells encoding specific faces has been conclusively disproved
Well, my analogy was comparing computer hash tables collisions to sparse embedding collisions, so what you are discussing now is my suggestion itself (pertaining to embeddings and recall), not the analogy, which is fine!
The study we're discussing was nominally about associative recall, not faces per-se, and specifically about the hippocampus not the cortex (that Macaque face study).
> Memory accuracy for pairing faces with objects and scenes dropped sharply
That said, I wouldn't be so sure that face embeddings are fully dense, even if they are not particularly sparse either, given that not all faces have the same set of features, such as facial hair, glasses, blemishes, etc. OTOH, it's possible, perhaps likely, that similar faces are stored together, in which case they may be more dense.
> If faces use sparse embeddings, you wouldn’t expect similar faces to evoke similar activity would you? Yet that is exactly what this paper shows, and the same has been shown in the human brain for faces.
With embeddings in general, sparse or not, you'd expect individual dimensions/neurons to represent different axis of variability, so you would expect individual neurons to be active for multiple different faces that are similar along that same axis (e.g. eye color). Note that the study you are citing used individual neuron recordings as well as fMRI, but of course we don't currently have the ability to simultaneously record from the hundreds of neurons that are likely being used to embed faces, so I don't think this study has much to say about the degree of sparsity of these embeddings. Obviously IF faces both with and without glasses are stored in the same embedding space (same set of neurons), then one would expect the "glasses neuron" not to be firing for a face without glasses, which would confirm some degree of sparsity.
> The brain just isn’t capacity limited by the number of neurons the way your intuition is pointing you.
It's highly unlikely that our brains are wasteful and have unused capacity - this recalls daft pop-sci articles saying that we only use 10% of our brain ... We know that brains and memory do degrade with age, and the only question is how - maybe the encoding mechanism itself is failing resulting in embeddings that have more overlap than they should (or one could hypothesize a dozen other possble failure modes). Do you have any theory that explains the "aging brains blend memories" study that we're discussing, at the level of detail of the hippocampal patterns they are seeing?
> It’s really odd to see these two paragraphs, because the second actually tells you why your first is wrong.
> Try as you might, you can’t separate out the deep linkages from the molecular to the behavioral in biology.
Of course the linkages are there since our brain is built from chemistry, yet selection pressure is happening at a much higher functional level. The part of my response you are referring to is addressing the question of how much of this molecular level detail needs to be retained in an ARTIFICIAL neuron model sufficient for it support the same phenotype-level functional behavior, and the answer is we just don't know, because nobody has yet tried to do it.
Prior to LLMs a lot of speculation about what is necessary in the brain to learn language, e.g. Chompysky-ian language-organ nonsense, might have sounded logical and compelling, but now we have proof-by-existence that "prediction is all you need". We're going to need to wait until we have built an artificial brain, capable of learning time-based phenomena, and everything else our brain is capable of, to similarly be able to point at something (a future elaboration of an artificial neuron model), and then be able to say that this is the most that is needed.
As far as this specific point - how much of the detail of a real neuron is functionally necessary vs how much of it is just a reflection of how it is built, you could also compare the massive complexity of something like a digital circuit transistor or logic component if you get down in the weeds and look at the specific gate architecture, and how it operates via quantum tunneling etc, or you could instead look at the functional behavior as a circuit component, and realize that none of it actually matters, and that transistors are interchangeable as long as they are functionally equivalent.
None of that changes whether there is a physical capacity, which I think was the larger point? There is no reason to believe distributed memory doesn't suffer from the capacity component of the bottleneck. In fact iirc there was some late 80s/early 90s papers on the memory capacity of NN. Btw I would advise against the absolute statement that there's absolute segregation of memory and processing.
I guess my point is the brain not being "von Neumann" in architecture or digital is not proof that isn't a "computer" of some sort.
> None of that changes whether there is a physical capacity, which I think was the larger point?
Capacity in what sense? Are we saying it’s X MB of data the brain can store? That claim is steeped in assumptions.
On the other hand, no one is claiming the brain has infinite memory or anything. And it’s certainly not a very accurate memory system. I’m arguing against “capacity” being understood as “these specific physical components located here and here we can ID store memories, and can get crowded with too many memories” sense.
This most particularly fails because not all memory is even identical in the brain, whether we mean the physical changes associated, the topology of the information, or how it’s activated.
> There is no reason to believe distributed memory doesn't suffer from the capacity component of the bottleneck.
I didn’t know there was a capacity component to the bottleneck, only a bandwidth one.
All I’m trying to say is that analogy to current typical memory storage systems to explain the brains memory processes is not helpful.
> I guess my point is the brain not being "von Neumann" in architecture or digital is not proof that isn't a "computer" of some sort.
Indeed, since the word computer was first used for humans. But what kind of computer matters enormously. Ising machine? Quantum+classical stack? Reservoir computer? All those frameworks have processes in the brain they can point to as homology.
Which points to a possibility: maybe the brain is multiple types of computers interacting. And the physical realization of these computing architectures aren’t spatially separated but thread through each other in the biochemistry and physical dynamics of cells.
https://www.pnas.org/doi/pdf/10.1073/pnas.79.8.2554 in a very precise sense going back a long time. I don't really know who you're arguing against or refuting? You seem to be touching on some pretty well accepted ideas.
"All I’m trying to say is that analogy to current typical memory storage systems to explain the brains memory processes is not helpful." Again, my impression is that the original author's idea was about capacity in general, then he gave an analogy. The analogy was wrong, but your reply went way beyond his specific analogy to the extreme of discarding memory itself as useful concept. To be clear, I agree that it is distributed and lossy and time-dependent. I agree it is not just a simple read-off of a static chunk with a fixed address.
"Which points to a possibility: maybe the brain is multiple types of computers interacting. And the physical realization of these computing architectures aren’t spatially separated but thread through each other in the biochemistry and physical dynamics of cells."
I would say that is both non-falsifiable and well-accepted.
I’m not sure what you think I’m arguing against but saying that it matches with one of the cognitive models du jour is… odd.
The problem with the hopfield model is it simplifies the brain too much. The base unit is “the neuron”. Ok… but what about the Astrocyte? Mathematically you can write it as a different kind of neuron. Or ignore it. But why, as a biologist, must I buy this model which ignores the third partner of every synapse, which has an entirely distinct physical tiling architecture compared to neurons, and which are at a temporal offset from neurons?
Those facts about the brain are missing from the model from 1982. Which isn’t shocking since we didn’t know all this then.
Are you saying the brain is a Hopfield network, and that’s it? Because later you indicate otherwise. Kinda confused what I’m to make of it.
> Again, my impression is that the original author's idea was about capacity in general, then he gave an analogy. The analogy was wrong, but your reply went way beyond his specific analogy to the extreme of discarding memory itself as useful concept. To be clear, I agree that it is distributed and lossy and time-dependent. I agree it is not just a simple read-off of a static chunk with a fixed address.
Ok, but my argument wasn’t with the OP mentioning capacity, but with their analogy. Am I not allowed to break down that analogy with evidence?
> I would say that is both non-falsifiable and well-accepted.
Why’s it non-falsifiable? If you do find a single computational paradigm that explains all brain dynamics we can measure, then you have falsified the hypothesis that it’s an integration of multiple computational types.
That actual evidence already gives the notion credence doesn’t make it unfalsifiable in principle.
Went through the description. Doubt it’ll interest me. As a rule I’ve stopped giving too much time to models that predate the last decades actual mechanistic facts. They’re fun curiosities, but hard to take seriously anymore. Here especially, the absence of astrocytes in the picture makes it hard to buy they have anything real to say about the mechanics at play. Half the cells of the brain not in the explanatory picture is just too likely to fail.
(note: I’m certain astrocytes are mentioned as support cells, or maybe regulators… but we just know a lot more now due to new techniques that makes downgrading them like that questionable science to me)
> Ok, but my argument wasn’t with the OP mentioning capacity, but with their analogy. Am I not allowed to break down that analogy with evidence?
If someone makes an analogy between aspects of A and B, then there is an implicit assumed shared understanding that A & B are DIFFERENT, and what is being pointed out is that they nonetheless may be considered as having something, typically fairly abstract, in common.
Calling a horse's reins as analogous to a cars steering wheel doesn't mean that the person making the analogy can't tell a horse from a car - that would be a very dumb take. A reasonable criticism might be "it's an imperfect analogy since the reins also act as the brake".
If you disagree with an analogy, then you need to address the analogy, not do the dumb take of "a horse is not a car" ("a brain is not a computer").
So, yes, you're "allowed" to "break down" (criticize) the analogy, but please don't just do the dumb take. So, hash key collisions are not analogous to associative recall and sparse embedding overlaps in what way?
It's more about computational complexity it seems but maybe it cites stuff that might interest you.
It's using computational neural networks that no one has ever believed represent the biology of the brain, but I think I must still disclose the following to save the precious time of the genius solving the problem all by himself:
It doesn't fully account for every biological detail ever documented, so it's probably a meaningless "curiosity".
Oh it probably also doesn't account for every detail discovered since publication so even if had value at time of publication it is not worth reading now.
> Calling a horse's reins as analogous to a cars steering wheel doesn't mean that the person making the analogy can't tell a horse from a car
When a person says “I crashed the car because my steering wheel tore, like reigns tear”, they are overfitting their analogy, and it’s perfectly fine to point out the structural and physical differences that make the analogy useless for the question at hand.
You have dismissed the biology that shows the problems with your analogy as immaterial. And continue to insist it’s the right one for the question at hand. This is a pretty pickle, because no facts can shake you from your certainty that you're right.
You seem in love with your analogy no matter how incorrect it is. And I have no problem with that. But when you put half baked biological claims to support it, I’ll point it out. If that’s too much for you to bear, maybe come up with better analogies?
Where the hell did you get that from?! I mentioned the hash table / embedding space analogy precisely ONCE, in my initial post, and, just to try to avoid people like yourself being triggered by it, I even included "obviously the brain is not a computer".
But you still got triggered by it, still did the dumb take of "a brain is not a computer" (no shit - I just said that), then spent your entire energy on fighting your own strawmen and not once even responding to my actual suggestion.
Now, finally, you are asking "what is a sparse embedding"! Maybe next time don't bother responding to something if you don't even understand what is being talked about.
He hadn't even considered whether capacity was well defined in a distributed system. He has a brand spanking new theory that allows him to jump straight past abstractions and simplifications and "ancient 1980s views"[1] and "fun curiosities" like the basics of dynamics, structure, organization, and composition described in an ancient book (2000s) he couldn't possibly learn anything from because he's a biologist and knows better. Nevermind that it smells an awful lot like [2]. But he doesn't need to know anything about it it or acknowledge it because dummy non-biologists couldn't possibly help someone as smart as him.
[1] https://www.pnas.org/doi/pdf/10.1073/pnas.79.8.2554 a paper he confidentally rejected as meaningless, unaware and not interested in the overwhelmingly likely possibility that his current brand new ideas trace back to it.
[2] "I’m stitching across a lot of different research, and I want to be clear many aspects of this system are not yet fully worked out.
But what we do know points to a system that works with different physical and algorithmic priors, and the dynamics are sharply distinct from current digital computers."
Hmm, I wonder if "we" figured that stuff out from individual and especially collaborative efforts combining abstractions and more theoretical approaches with expertise in biology, not to mention EE, biophysics, physics, and whatever else other fields that don't know about his favorite molecule and thus can be ignored if not shit on?
I mean with respect to your other thread too, where you seem to argue/believe academia has no idea about your brand new dynamical distributed view. I pointed to a really old paper with a self-admitted extremely simple model to point out that one of your specific insights over the field re. memory not being a memory stick is ancient and that your arrogance wrt being certain that your notion of memory is so novel that you weren't even sure there was a definition of capacity for distributed memory.
I intentionally gave older books and papers. I very clearly understand as does everyone in the field that they are not current or complete. But that awareness is precisely my point.
"Are you saying the brain is a Hopfield network, and that’s it? Because later you indicate otherwise. Kinda confused what I’m to make of it." You being confused by that is kind of my point, it shows the projection you're applying to the entire field. I understand models are wrong and iterative. I also understand the parts that survive, survive in the literature and the field at large. For example, the PNAS paper that showed "hey we think memory is distributed, right? Here's a simple model showing how memory might roughly be stored and recalled via these very simple dynamics over this very simple model. That's cool. Let's build on it." Did I make it clear enough yet that it was simple and wrong? Do you thus believe it wasn't valuable? I guess that would match your confident opinion that the whole field is dummies and of course those dummies cited an incorrect model so much.
If you stopped with saying "we know memory isn't a memory stick and here's the well-established evidence", then I'd fully agree with "Ok, but my argument wasn’t with the OP mentioning capacity, but with their analogy." and would've said nothing. Instead your posture was "I know some specifics based on recent findings that aren't obviously incorporated as of now into the current understanding -- which I refuse to read because it doesn't address precisely this thing only I understand(1) -- , so here is my theory that is certainly novel don't dare try to point me to preexisting literature that might help me refine my theory."
To be clear, I want everyone to explore their ideas and think it's great that you doubt things. But I don't like someone claiming superiority over a field they refuse to understand. And if your doubt something maybe you should check first if there's anyone else doubting it. (Hint: not only did everyone doubt it, they knew it was wrong and they're all actively working on using that doubt to improve it.) Of course a theory that predates a discovery doesn't cover that discover. Of course a simple model of memory isn't a complete model of memory, that was in fact the point! The notion of distributed memory and how it could feasibly work according to dynamics is made more precise than ever before to that time in that paper (among others). And of course it's outdated. But it contains a more precise statement and falsifiable statement of "memory is complicated dynamics and those dummies don't know and also that book is stupid because it doesn't cover X(2,3)" your theory espouses. My issue is with the arrogance and certainty you know best while refusing to actual understand what others think. And thus give no credit to. Science is iterative. Someone has an idea, builds a simple model as proof of concept. Someone else builds on that model precisely where it has holes (whose precise shape only present itself after new experiments motivated in part by the obviously wrong and simplified model and other attempts to break the model). Repeat. Built on the shoulders.. etc.
"Why’s it non-falsifiable? If you do find a single computational paradigm that explains all brain dynamics we can measure, then you have falsified the hypothesis that it’s an integration of multiple computational types.". I said non-falsifiable AND well-accepted. Non-falsifiable was the wrong word. What if I instead said trivially falsified and also adds nothing not known except to the layman who posits an analogy they clearly aren't sure of? Its non-falsifiable relative to the current state of the field, which I now know you're not interested in, because it says nothing not already stated more precisely, as of like 40 years ago.
(1) "As a rule I’ve stopped giving too much time to models that predate the last decades actual mechanistic facts. They’re fun curiosities, but hard to take seriously anymore. Here especially, the absence of astrocytes in the picture makes it hard to buy they have anything real to say about the mechanics at play. " In response to a book that has a lot to say about your new and novel dynamical theory.
(2) A hole it couldn't cover because the experiments and techniques didn't exist at that time. But they knew their theory wasn't complete. They were sharing what was learned up to that point especially the parts that seemed to generalize and survive specifics. That is, an incomplete model they knew was incomplete. But which I promise would help advance your journey.
(3) "Instead, memory, over time, is driven by a dynamical regime that adjusts its dynamics to account for the temporal pattern in the salient stimulus." Which is true and everyone in the field agrees with. And everyone agrees current models don't fully account for everything. Hence their continuing efforts to improve them by accounting for new findings.
Perhaps undiscussed in the paper is how the blending of memories may also be better. If evolutionary pressure encourages the structure of brain to optimise for accurate prediction, not memory, then it is quite possible that older individuals may be making better predictions than younger ones, in situations where they don't remember the details...
I'm not sure about better, unless this is the normal mechanism for generalization, but if a form of generalization (not just confusion) is the effect, then it'd be a nice form of graceful degradation, whether selected for or not.
I wonder to what extent evolution selects for longevity (& graceful degradation) past a certain point? I would think that once you are past breeding age it's generally more beneficial for the species if you die off and make way for the next generation, other than having some residual benefit as babysitters for the grandkids and leading the herd to the water source in the next drought.
My understanding is that its just straight up deterioration? A wishful view might be something like a pressure toward generalization(wisdom?) allowing for partial overwriting/grouping of specifics. Like a coarser but more predictive representation.
This is really fascinating. The actual mechanisms of the human mind are distinct from computer systems, yet there are some parallels.
There are some hints that increased memory access times scale with the amount of information the brain has stored vs the more typical narrative that aging decreases the capabilities of the brain.
> Our results indicate that older adults'; performance on cognitive tests reflects the predictable consequences of learning on information-processing, and not cognitive decline. We consider the implications of this for our scientific and cultural understanding of aging.
I wonder the effect of diary writing on the brain, I see always recommended as either part of the self-improvement space or meditation. But reading your comment, I wonder if it allows to strengthen memories, especially if ones reread it after a while. I must experiment, now it's easier than ever with LLMs
There’s a lot that’s good here, but you just can’t be slime molds in a modern capitalist corporation.
A substantial portion of company goals is set by a very small group that is far removed from the day to day work, and the goals are often connected to timelines and financial expectations well before any team member gets into the project.
To have slime mold like behavior, a corporation would have to hire employees, give them time and resources, and general guidelines and goals, but let specific targets and projects bubble up.
This is fundamentally incompatible with a next-quarter profit driven corporate financial structure.
True. This was written with the context of Google which was, to a large extent and up until very recently, a good example of the resilience and magic a slime mold can be.
It is not true today, and as you mentioned "fundamentally incompatible with a next-quarter profit driven.."
This is so so wrong I don’t even know where to begin. There may be no objective definition or measure of consciousness. But it is not a property arising from mutual care. That argument is like one of the many “just so” evolutionary psychology “theories”.
Why or how would you care about another, or even distinguish yourself from another, if you weren’t conscious and had a felt boundary between yourself and the other world?
Only in the last paragraph did it become clear to me what was going on: a Google VP is proclaiming the thing he cares about is conscious, and so of course this is why everyone else must be assigning consciousness! Not convenient at all, no siree. Pure rational science and personal desire and economic incentives just happened to align here.
Here’s how I’m thinking about it (still digging into the details): across genomics, it’s become clear few traits have clear traceability to a few loci in the genome.
Instead, evidence has been growing that epistasis, the nonlinear interaction between genes and other genomic regions, predominates in explanations of most phenotypes.
What this paper does is show where upstream of the genome various combinations of mutations can interact to cause damage during development, thus leading to the phenotype. Rather than correcting a particular mutation, or targeting drugs to their protein products, we may find downstream protein-protein interactions that are strong drivers of the phenotype, and hopefully find ways to prevent/reverse these effects.
For me the image in front of my eyes is not gone. I’m just not going to remember any details about it because my focus is inwards, on what my minds eye is showing me.
Right, but the question is - do you see actual images with this "mind's eye", or is it more of a sensation, a logical understanding of what you're imagining but with no real imagery?
Oh proper images. And if I focus I can increase resolution. I can get to where it’s as good as real life, for most things. But if it’s a book character I’ve never seen, the details are more oddly distributed and not stable over time, except when the author does a really good job with descriptions.
Seeing mental images that are as good as a well done animation, on the other hand, is a lot easier and less draining to hold.
I see actual images, but at least for me its not anywhere near the same as actually looking at something. Its sort of like watching a low fidelity movie for me and I can't really "focus" in on specific details of what I'm visualizing. I think there is definitely a spectrum.
reply