Ray Fernando
speaker
138 appearances
1 recordings
1 series
first heard Jan 2025
last heard Jan 2025
Ray Fernando’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
And that actually involves using this thing called Open Web UI. So I can show you that. So while this is thinking, if it even returns the results or does anything here, we're going to go ahead and pop on over to the other side. So on the other side, we're going to go to here. And so I have an instance of what's called open web UI, and it looks very similar to like a chat GPT.
And to get this set up, I'll probably go through a little bit more details, but I'll just go ahead and show you what this looks like. So in here, I have the model selected. I can go to deep seek. And so what's great is that you can connect to an API provider And I'm using Fireworks AI. So Fireworks AI here is currently hosting DeepSeq model.
And they allow you to use the model just by, you know, using getting an API key and then putting in the exact model string and so forth here. And so from here, if I go to the open web UI, I'm able to select them and say, OK, this is my DeepSeq account. I'm going to go ahead and just paste the exact prompt that we had, paste it in with my transcript and everything here.
And I should be able to get everything out. So let me just double check here that I got everything. So it's still timing out here. Yeah, server is busy. Try again later. So yeah, that's not fun. So what I'm going to go ahead and do is scroll to the top and hit this little copy button and then go over here. Just make sure I put everything in there. I have the whole transcript.
Yeah, so the whole transcript here. And so what's going to happen here is when I hit send, it's going to send this off to Fireworks AI. And what's great about this thing that's actually running in this open web UI is that it's using the API and it's not actually sending the data to China.
So just for as it's doing its thinking here and showing us what's going on, I'm going to go back to overlay this model of our data in our container and kind of show you what this is looking like in the background. So here in TLDraw, we actually have our Mac and PC. And so I used web UI here and I'm actually using the Fireworks API.
So I'm going to the cloud and this cloud is located in North America. So the data basically resides here in North America and it's going to be delivered back to my device. So that's what we're doing. When we were on the DeepSeq website, we were going out to the China region. So that's just a heads up of kind of how that's working in the background.
And so next I'll show you the difference of what's going to happen and like the speed difference with what Grok hosting provider provides. And then a little bit later, we'll get into a little bit of details for how to get the setup. So as you see, it's kind of outputting this stuff here.
Because these models are still so new, these web apps are still adjusting to take a look at the reasoning stuff. And so what I'm going to go ahead and do is hit this little pencil up here to the very top. I'm going to make this a little bit bigger so you can see. And then from here, when you hit this little pencil, it creates a new chat. At the very top, you can select the model dropdown.
I'm going to type in DeepSeek. And the one that I have set up from Grok is called the Distilled Llama 70B. And so this model is actually like a smaller distilled version that they're hosting, but it's incredibly fast. And so if we hit this here, it seems like nearly instant by the time all the stuff kind of starts finishing. So we'll see this model actually going out.
It's doing its thinking and now it's actually providing the response just like that. Super fast. So if we take a look, this this has thought for a few seconds and actually shows us the reasoning that was going on. So this is actually going through my transcript. trying to really understand what was going on with my transcript.
This was an interview with LDJ who spoke a lot about deep seek and technical details of things that, you know, I really couldn't remember. And so then it basically makes this like, you know, very simple blog posts. Um, so we'll see if my other, um, one that was running the larger models finished. And you can see the difference between the two models.
A distillation model is just going to give us like a small little blog post versus the full model that's also running on Fireworks API is actually giving us quite a bit of detail. It's going to take more time, but take a look at what it's doing right now.
so it's above here was all this thinking stuff but this is actually now doing an analysis on my transcript and generating a really nice blog post uh and so it's telling us about the calculations that ldj talked about in the stream uh geopolitical implications of what's going on with the new ai arms race uh also future predictions and we talked about these details in the live stream and it literally picked them up and is now creating a graph from this this is how crazy these models are if you can really think about it
And here's some key takeaways as well. So that's an amazing thing. And I'll be able to share these prompts with you and, you know, so that you can actually run these analysis on your guys's own transcripts as well. Yeah. So here's the SEO enhancements and final thoughts. Yeah.
Yeah. Or a research engineer that you hire to really thoughtfully take a lot of notes, spend a lot of time analyzing and like put together how you would want to report. And it's even more incredible because these instructions can be configured. So if you want like a graph or you want a type of thing,
we can take this prompt and put it into DeepSeq itself to say, can you give me this type of output instead? And it'll do that for us. It's like, how do we improve the prompt? Or what do you want to see from your outputs all the time from your live streams?
And I think the thing that I've seen, this is kind of the biggest breakthrough that's happening is that I'm seeing this also with O1 Pro, by the way. O1 Pro and DeepSeq reasoning models, these reasoning models spend extra time and actually pay attention to your instructions. And so every little detail that they're seeing, they're like, oh yeah, I haven't done that yet.
Okay, let me go ahead and make sure I still do that. And that's something that I super deeply appreciate. And for me, it's worth the extra 200 bucks I pay a month to open AI. But this is really quickly turning my head and like, oh my goodness, did you understand like what just happened here? It's like, I'm a little still taken, I'm still taken away at this output.
Like you're saying, it's very detailed. And it's, to me, I feel like this is totally a game changer. And I think one thing that people aren't really talking about right now is actually this additional rush to understand who can host this in order to host these huge, like, 600 plus billion parameter models, you need all those GPUs. You need services like fireworks. Grok is trying to spin that up.
Showing 21–40 of 138 · page 2 of 7
← Previous
Next →