Sebastian Raschka
speaker
1,024 appearances
1 recordings
1 series
first heard Feb 2026
last heard 1 Feb
Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.
Appearances
Yeah, I think that is right now still a hot topic.
And also big companies like OpenAI, they approached private companies for their proprietary data.
And private companies, they become more and more, let's say, protective of their data because they know, okay, this is going to be my mode in a few years.
And I do think that's like the interesting question where...
If LLMs become more commoditized, and I think a lot of people learn about LLMs, there will be a lot more people able to train LLMs.
Of course, there are infrastructure challenges, but if you think of big industries like pharmaceutical industries, law, finance industries, I do think they at some point will hire people from other frontier labs to build their in-house models on their proprietary data, which will be then again another unlock with pre-training that is currently not there because even if you wanted to, you can't get that
You can't get access to clinical trials most of the time in these types of things.
So I do think scaling in that sense might be still pretty much alive if you also look in domain specific applications, because we are still right now in this year just looking at general purpose LLMs on Chachapiti, Anthropic and so forth.
They are just general purpose.
They're not even, I think, scratching the surface of what an LLM can do if it is really specifically trained and designed for a specific task.
And like Nathan said, also two layers to it.
Someone might buy the book and then train on it, which could be argued fair or not fair.
But then there are literally straight up companies who use pirated books where it's not even compensating the author.
That is, I think, where people got a bit angry about it specifically.
I have a repository called ML Extend that I developed as a student, I don't know, 15 years, 10 years ago.
And it is a reasonably popular library still for certain algorithms, I think, especially like frequent data mining stuff.
And there was recently, I think, two or three people who submitted a lot of PRs in a very short amount of time.
I do think LMs have been involved in submitting these PRs.
Me as the maintainer, there are two things.
First, I'm a bit overwhelmed.
Showing 321–340 of 1,024 · page 17 of 52
← Previous
Next →