Matei Zaharia
speaker
417 appearances
1 recordings
1 series
first heard Jun 2026
last heard 24 Jun
Matei Zaharia’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Yeah, exactly, yeah.
Oh, yeah, very cool, yeah.
Yeah, so we decided, you know, even though we did launch an open-source model, DBRX, and, you know, we went up to, like, sort of above the LAMR 3.0
scale, we decided that we really want to focus on, there'll be so many people releasing models, and instead of doing the general model where like, you know, a big part of the recipe is just throw in a lot of compute and just scale, we want to focus on like the next step also of, let's say you have the very smart model, how do you make it, you know, useful?
For us, it was a lot about
automating, making it very good at querying data.
That's the first party agents we have called Genie.
It's like a virtual data scientist.
Imagine there's someone who already knows all the stuff in your company inside out and knows all the machine learning libraries, all the data libraries, all the stuff on the web, and you can ask them questions.
That's what we wanted to do first.
that meant let's not focus as much on let's just train some frontier model, but let's build this system using either external models or fine-tuned customized components.
We're still doing quite a bit of model training though, and in fact, we're procuring lots of GPUs and stuff all the time to do it.
There's a few places where we're doing it.
One is there are many high-volume use cases where if you have a specialized model,
It's just so much better than any of the general models you get.
A nice example of that is understanding documents like PDF, Word documents, stuff like that, parsing them.
If you've ever tried to do that, it's frustrating because you send it to like Claude Fable or whatever, it almost gets it, but it gets some things wrong and it's super expensive.
You just burnt a huge amount of tokens plopping in an image into there.
So our team built this document sort of vision model that takes a page and gives you back a nice JSON with all the components.
And it's very competitive.
Showing 341–360 of 417 · page 18 of 21
← Previous
Next →