Joscha Bach

speaker
510 appearances 2 recordings 2 series first heard Aug 2023 last heard 23 Jan

Joscha Bach’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.

Appearances

newest first · ▶ plays the moment
For me, it's fascinating that they are so vastly different and yet in some circumstances produce somewhat similar behavior. And the brain is first of all different because it's a self-organizing system where the individual cell is an agent that is communicating with the other agents around it and is always trying to find some solution. And all the structure that pops up is emergent structure.
Right. So one way in which you could try to look at this is that individual neurons probably need to get a reward so they become trainable, which means they have to have inputs that are not affecting the metabolism of the cell directly, but they are messages, semantic messages that tell the cell whether it has done good or bad and in which direction it should shift its behavior.
Once you have such an input, neurons become trainable and you can train them to perform computations by exchanging messages with other neurons. And parts of the signals that they are exchanging and parts of the computation that they're performing are control messages that perform management tasks for other neurons and other cells.
I also suspect that the brain does not stop at the boundary of neurons to other cells, but many adjacent cells will be involved intimately in the functionality of the brain and will be instrumental in distributing rewards and in managing its functionality.
So first of all, there's a different loss function at work when we learn. And to me, it's fascinating that you can build a system that looks at 800 million pictures and captions and correlates them. Because I don't think that a human nervous system could do this.
For us, the world is only learnable because the adjacent frames are related, and we can afford to discard most of that information during learning. We basically take only in stuff that makes us more coherent, not less coherent. And our neural networks are willing to look at data that is not making the neural network coherent at first, but only in the long run.
By doing lots and lots of statistics, eventually patterns become visible and emerge. And our mind seems to be focused on finding the patterns as early as possible.
Yes, it's a slightly different paradigm, and it leads to much faster convergence. So we only need to look at a tiny fraction of the data to become coherent. And of course, we do not have the same richness as our trained models. We will not incorporate the entirety of text in the internet and be able to refer to it and have all this knowledge available and being able to confabulate over it.
Instead, we have a much, much smaller part of it that is more deliberately built. And to me, it would be fascinating to think about how to build such systems. It's not obvious that they would necessarily be more efficient than us on a digital substrate, but I suspect that they might.
So I suspect that the actual AGI that is going to be more interesting is going to use slightly different algorithmic paradigms or sometimes massively different algorithmic paradigms than the current generation of transformer-based learning systems.
My main issue is I think that they're quite ugly and brutalist.
Yes, they are basically brute forcing the problem of thought. And by training this thing with looking at instances where people have thought and then trying to deepfake that. And if you have enough data, the deepfake becomes indistinguishable from the actual phenomenon. And in many circumstances, it's going to be identical.
Yes, that's a very interesting question. I think that these models are clearly making some inference. But if you give them a reasoning task, it's often difficult for the experimenters to figure out whether the reasoning is the result of the emulation of the reasoning strategy that they saw in human written text, or whether it's something that the system was able to infer by itself.
On the other hand, if you think of human reasoning, if you want to become a very good reasoner, you don't do this by just figuring out yourself. You read about reasoning. And the first people who tried to write about reasoning and reflect on it didn't get it right.
Even Aristotle, who thought about this very hard and came up with a theory of how syllogisms work and syllogistic reasoning, has mistakes in his attempt to build something like a formal logic and gets maybe 80% right. And the people that are talking about reasoning professionally today read Tarski and Frege and built on their work. So in many ways, people, when they perform reasoning,
are emulating what other people wrote about reasoning. So it's difficult to really draw this boundary. And when François Chollet says that these models are only interpolating between what they saw and what other people are doing, well, if you give them all the latent dimensions that can be extracted from the internet, what's missing? Maybe there is almost everything there.
And if you're not sufficiently informed by these dimensions and you need more, I think it's not difficult to increase the temperature in the large angles model to the point that is producing stuff that is maybe 90% nonsense and 10% viable and combine this with some prover that is trying to filter out the viable parts from the nonsense in the same way as our own thinking works, right?
When we're very creative, we increase the temperature in our own mind and recreate hypothetical universes and solutions, most of which will not work. And then we test. And we test by building a core that is internally coherent.
And we use reasoning strategies that use some axiomatic consistency by which we can identify those strategies and thoughts and sub-universes that are viable and that can expand our thinking. So if you look at the language models, they have clear limitations right now. One of them is they're not coupled to the world in real time in the way in which our nervous systems are.
So it's difficult for them to observe themselves in the universe and to observe what kind of universe they're in. Second, they don't do real-time learnings. They basically get only trained with algorithms that rely on the data being available in batches. So it can be parallelized and runs efficiently on the network and so on. And real-time learning would be very slow so far and inefficient.
Showing 261–280 of 510 · page 14 of 26 ← Previous Next →