Ray Fernando

speaker
138 appearances 1 recordings 1 series first heard Jan 2025 last heard Jan 2025

Ray Fernando’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
I don't think app store app store. I'm going to see if I can go here. Apollo. Let's see. Okay. Private local AI.
yes so i have this app on my phone and they allow you to download the models directly just like you would with olama as well but it's it has its own interface which is really nice and so um i wonder if i could share my screen i think i can so on your phone yeah yeah they have a phone mirroring exactly apollo Okay. Oh, I have to lock my phone. Okay, cool.
So I lock it and then it should be able to connect. Okay, cool. Nice. Awesome. So yeah, let me kind of minimize this here and yeah. Okay. Let me just go to a different screen here. Probably one that's less cluttered and do phone. Whoops. Put that over here. Cool. Maybe this will work. I think. Yeah. Yeah. Sweet. Yeah, I actually have the yes.
Another place to get your models apparently is also through open router. And so, yeah, so this is kind of the Apollo app. You're like, OK, cool. I can. Can I start chatting with this? You know, as soon as I play with the thing, a couple of configurations you have to do is you hit this little hamburger menu at the very top left corner and then you hit settings.
So on the on the phone app, you hit settings and it's going to say AI providers settings. And when you click there, you have three different options. Open Router, which is another API provider. And you can also get access to pretty much every model there, which is also very handy. I think they give you some free credits, but then you would put your credits there.
And then you have the local model and then you have custom backends. So with the local model, they actually can tell how much memory you have on your device and they'll actually have a little download button for those models. The ones that are not available with the download button basically means you can't run that on your device because you don't have enough memory to run them.
So these downloads are pretty big, like four gigabytes, and some of them are several gigabytes. So just depending on the space on your phone. So you can actually run the distilled Lama 8-bit MLX version. And I have the distilled Quen version at 7B. So it just depends on your... Oh, that one's actually not compatible. Which one do I have downloaded? So I think on mine, let's see.
The one I have available is the DeepSeek R1 from Apollo. I think I have it from OpenRouter that's running. So let's take a look here. AI providers, OpenRouter. Yeah. So the one that I have set up right now is from OpenRouter. So OpenRouter will show you all the models. You can select DeepSeek R1 from there, which is awesome. So you can have a conversation.
So this just requires me being connected to the internet. We start a new chat. You're like, tell me more about options trading. And so here you're still talking to the model, but you're actually just using open router. And so that's a little bit different than, you know, sending your stuff directly to deep seek. And they should be able to do that.
It's possible that this model is busy or it's currently down. That can happen. So, yeah, that happens. Yeah. While that's going, I think we could even start another new chat. Let's see this model. You can select a different model. So let's see.
There's so many. Yeah. It's like, how do you know which one does it? I feel like you just go off vibes. Like what's, what's my friend telling me? It's yeah. Like what's the real vibes right now? So the vibes right now, obviously R1 is like the real hotness. People are like totally into that right now. Um, and it makes sense cause you know, reasoning, uh, at a much lower costs. So, um, let's see.
Um, there's probably something going wrong with my API key or something. So AI providers, I can select local model to run. You know, I want to see if there's something small here that we can download. So we could do, yeah, this distilled quen, just for speed purposes, we'll just download the gigabyte one. So this is going to download, wow, that's really fast.
the quen model 1.5 b and so that'll run deep seek locally and so basically it's just downloading it directly from i think hugging face and then the model is being loaded on my phone and um this this is actually optimized to run on apple hardware or apple silicon so that's um you know one way that you can kind of take a look at it uh to run this thing and so what's nice yeah if this phone runs out of internet or i need to ask some questions or do some stuff
I will have this R1 reasoning model that's a much smaller version to run on device. And I think that's another good point about AI that's running. And you don't always need the most powerful thing running for every single type of thing. I think it's really important to... understand different use cases, you know, because maybe you don't need that depth of reasoning.
You just need something that's really quick or you just need something that's really good at like gathering lots of information and just telling you some topics or something like that. And that could just be done really quickly. So it's kind of like picking the right tool for the job and experimenting. So we're at a good age today where you can actually get these models and experiment with them.
So now I should be able to select this guy and run it. So let's go ahead and hit done and start a new chat. And then over here, we're going to go ahead and select the model. So here we're going to type in, oh yeah, it already has it at the top. So you see this little icon that signifies that it's running locally. And then we're going to hit cancel. Hit done. Okay, great.
So now we're running with that local model. And I think we're just using a default system prompt about it being Apollo. And you're like, yo, tell me about options trading. And it should basically start to cook. So it's using my phone's power there. And it's now thinking. And so if we click this little dropdown, we'll actually see the reasoning tokens. Wow. Yeah. I have reasoning on my phone.
No internet. Completely running locally. 2025 is insane. Yeah. Yeah. And imagine being able to run this on your watch. Like that'll just be because this is already showing its capability. Like we're doing this input. If you're, if you can make an app that can run on a watch all locally, you know, just think about like the transcription stuff, right? Uh, you have a very, very lightweight model.
You send the audio, you know, from the watch, you know, especially of a loved one, maybe they've fallen or something. It can just turn on the speaker and try to understand the situation and, And then if it listens to paramedics or something about asking questions and they don't really know, maybe the watch can show, hey, there's this app here.
I'm going to show you the emergency card this person has for their medications. Or this is something that's happened in the last, you know, five minutes before this event or something. This is kind of the way that people think about designing apps with these models is trying to think about these use cases. Because now you have really powerful devices just all like on the size of your wrist that
Showing 101–120 of 138 · page 6 of 7 ← Previous Next →