Arvid Lundmark
speaker
98 appearances
1 recordings
1 series
first heard Sep 2024
last heard Sep 2024
Arvid Lundmark’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
I write, like, less than in a comment and, like, I maybe write the greater than sign or something like that. And the model is like, yeah, you look sketchy. Like, are you sure you want to do that? But eventually it should be able to catch harder bugs, too.
I mean, there's certainly cool solutions there. There's this new API that is being developed for... It's not in AWS, but, you know, it certainly is. I think it's in PlanetScale. I don't know if PlanetScale was the first one to add it. It's this ability to sort of add branches to a database, which is...
Like if you're working on a feature and you want to test against a broad database, but you don't actually want to test against a broad database, you could sort of add a branch to the database. And the way to do that is to add a branch to the write-ahead log. And there's obviously a lot of technical complexity in doing it correctly. I guess database companies need new things to do.
They have good databases now. And I think TurboBuffer, which is one of the databases we use, is going to add maybe branching to the Red Hat log. And so maybe the AI agents will use branching. They'll test against some branch, and it's sort of going to be a requirement for the database to support branching or something.
Right. Yeah. I feel like everything needs branching. Yeah.
I mean, there's obviously these super clever algorithms to make sure that you don't actually use a lot of space or CPU or whatever.
I have a few friends who are super senior engineers, and one of their lines is like, it's very hard to predict where systems will break when you scale them. You can sort of try to predict in advance, but there's always something weird that's going to happen when you add this extra zero. You thought you thought through everything, but you didn't actually think through everything.
But I think for that particular system, we've... So for concrete details, the thing we do is obviously we upload, we chunk up all of your code and then we send up sort of the code for embedding and we embed the code. And then we store the embeddings in a database, but we don't actually store any of the code. And then there's reasons around making sure that
We don't introduce client bugs because we're very, very paranoid about client bugs. We store much of the details on the server, like everything is sort of encrypted. So one of the technical challenges is always making sure that the local index, the local code base state is the same as the state that is on the server.
And the way sort of technically we ended up doing that is, so for every single file, you can sort of keep this hash. And then for every folder, you can sort of keep a hash, which is the hash of all of its children. And you can sort of recursively do that until the top. And why do something complicated? One thing you could do is you could keep a hash for every file.
Then every minute you could try to download the hashes that are on the server, figure out what are the files that don't exist on the server. Maybe you just created a new file. Maybe you just deleted a file. Maybe you checked out a new branch and try to reconcile the state between the client and the server. But that introduces, like, absolutely ginormous network overhead.
Both on the client side, I mean, nobody really wants us to hammer their Wi-Fi all the time if you're using Cursor. But also, like, I mean, it would introduce, like, ginormous overhead in the database. I mean, it would sort of be reading this... Tens of terabytes database sort of approaching like 20 terabytes or something database like every second. That's just kind of crazy.
You definitely don't want to do that. So what you do, you sort of, you just try to reconcile the single hash, which is at the root of the project. And then if something mismatches, then you go, you find where all the things disagree. Maybe you look at the children and see if the hashes match. And if the hashes don't match, go look at their children and so on.
But you only do that in the scenario where things don't match. And for most people, most of the time, the hashes match.
And I mean, the point of, like, the reason it's gotten hard is just because. Like, the number of people using it and...
You know, if some of your customers have really, really large code bases to the point where, you know, we originally reordered our code base, which is big, but I mean, it's just not the size of some company that's been there for 20 years and sort of has a ginormous number of files. And you sort of want to scale that across programmers.
There's all these details where like building a simple thing is easy, but scaling it to a lot of people, like a lot of companies is obviously a difficult problem. Which is sort of independent of actually, so there's part of this scaling our current solution is also coming up with new ideas that obviously we're working on. But then scaling all of that in the last few weeks, months.
And it's not a problem of like weaker computers. It's just that, for example, if you're some big company, you have big company code base, it's just really hard to process big company code base, even on the beefiest MacBook Pros. So even if it's not even a matter of like, if you're just like,
a student or something I think if you're like the best programmer at a big company you're still going to have a horrible experience if you do everything locally I mean you could you could do edge and sort of scrape by but like again it wouldn't be fun anymore
Don't you want the most capable model? You want Sonnet?
Showing 61–80 of 98 · page 4 of 5
← Previous
Next →