Erhan Giral
speaker
389 appearances
1 recordings
1 series
first heard Aug 2026
last heard 3 Aug
Erhan Giral’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Aug 2026 with 1.
Appearances
But at the core, just like you said, we do rely on LLM's reasoning capabilities to figure things out, to do the orchestration especially.
Yeah, yeah.
So we follow the similar philosophy of divide and conquer, create good encapsulations around capabilities.
So we follow an architecture pattern called mixture of experts.
What this is, is it allows you to... These pre-trained models are trained on piles of piles of publicly available textual data, if you're talking about language models.
And what we found was, actually, we didn't find this, but a meta sort of led the way into this architecture in our adoption.
They described, they had this wonderful paper to talk, where they talked about, you know, how
these models can be further tuned for different purposes while encapsulating that training in a specific set of weights and biases called experts.
So what that is, is you take one of the open weight models
OpenLate model means a model that you can basically download to your computer and run as much as you want based on, of course, depending on the licensing agreements.
So what you do is you take one of these open source models and then you identify a problem
data set and a reward function.
So you basically devise a clever way of signaling the network, hey, when you give me a good response, I'll give you cookies.
If you fail, you'll get the stick.
So you basically
train these models on additional data sets that allow you to focus on that problem.
And then you capture these as what we call experts.
So experts are just like you're saying, they're mini, I'm going to say they're services, but they're encapsulations of that training.
They're the artifact of that training cycle.
And within this mixture of expert architecture, you also specify gate activation values during your training, which means the model then knows when it's prompted with a question, whether to activate that training set or not.
Showing 101–120 of 389 · page 6 of 20
← Previous
Next →