Ray Fernando

speaker
138 appearances 1 recordings 1 series first heard Jan 2025 last heard Jan 2025

Ray Fernando’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
It's really up to you and your content, but I would definitely do two different responses with different temperatures and test those things and see if you see any difference in your output. For me, sometimes I find the temperature zero to be very helpful for very logical reasoning purposes. especially around code.
But it's really up to, it kind of varies and I just kind of want to give you a heads up on that.
You know what would be funny? To basically have like a spinner where you can actually flick it yourself and you kind of see it land on something and then just like hit go.
Because sometimes you really don't care, right? You're just like, I just want to spin the bottle and see what happens. Like... Totally. It's this kind of YOLO mode, kind of. Yeah. Yeah.
Because I think, like you say, there's huge opportunities in the AI space to be playful. And I think that's what's interesting is you have these intelligence of the models. And then now you have to have people who build interfaces to interface with them. And there are a lot of companies who are trying to do that. And, you know, you can get very far with just some prompting as we're seeing here.
And then we're trying this exercise here is to try different models. So. if you think about it, Ollama is sort of the gateway to all these different types of models that you can try out and see if it even works for your use case. And this web UI is actually a really nice user interface to keep track of that. It's safe locally on your machine. You can go back to them at any time.
There's additional options at the bottom here, which is really nice. So you can actually have this read out loud to you. So if you're a person suffering maybe from dyslexia or you actually prefer audio, you can have that for you. This will give you some information there. You can continue the response. Sometimes if you have too much information, it still needs to continue going.
So you hit the continue and it'll just continue on or regenerate the responses. So that's kind of some of the basics there. So yeah. Um, so this is the output of this model and I'm fairly impressed for being a 7 billion parameter model at running locally on my machine, uh, that it took that entire transcript and did this analysis type of thing. That's I'd say is pretty close to the bigger model.
And, um, and in terms of details, it's not as detailed as the other one, if we kind of take a look. So like, The previous, this is one that came out with before, you know, with this nice big blog post type of thing. So it's pretty good and it's running, you know, locally. I can run this on the plane as far as that. So yeah, so to get started, basically, again, it's just open web UI.
There is a getting started. It's literally a couple steps to run. Make sure you have Docker installed there. And then Ollama is going to show you all the different models. So if you go to the models, you'll see kind of stuff that's popular and trending right now. And that'll kind of get you some of that as well as far as getting started. There is, you know, we also talked about Fireworks AI.
So that's Fireworks. It's a good resource for you to, you know, go take a look and put that model in. So like if you want to put that model into your Ollama, you would kind of do the same thing here. So go to user and then you go to admin panel. And then you would go to settings up here. And then from the settings, you're going to go ahead and hit connections.
And so what you'll do is go ahead and hit the little plus connection. And so you have to put in the base URL and you'll also have to put in the API key. So the base URL here for fireworks is this specifically here. It says api.fireworks.ai slash inference slash v1. In the example documents, you'll see slash chat slash completions and things.
You don't need those because that's part of the OpenAI framework is that you just put everything up to v1 and then you'll generate an API key from that model. over in Fireworks, so AI. So if you go to the model here in Fireworks and you go to your name, and then if you go to API keys, once you go to API keys here, you just hit create API key and that'll pop up.
And that's the key that you want to put in there. Similar to Grok Cloud, you just go ahead and hit create API key. So once you go to console.grok.com, There's an API key section here. And then you'll want to hit create API key. And that'll pop up a dialog with those API keys. And so that endpoint will look something like this over here.
So that'll be, if we hit configure, api.grok.com slash openai slash v1. And then you put your key in there. And you don't have to do anything with these IDs. These will be pooled directly from that endpoint. So whatever models you have available will be there. And so now when you hit the plus sign, you'll see like this nice list of models from fireworks. So there'll be the fireworks one.
So account slash fireworks. You can play with any one of those. And then the other ones that are just with the normal name are from Grok. So they have those as available for there. So you can you can play with a lot of these models, which is nice and compare them. And then the ones at the bottom are the ones from Olamo.
And it's a little show like, you know, the colon latest is kind of how you can tell. And if you hover over them, you'll see like some additional information over the parameter count, what quantization level it is. So Q4 means it's quantized to four bits. And that also has a play in its intelligence. Obviously, the higher level of quantization, you know, means more memory. So it's like 32 bit.
16, uh, all the way down. Um, so the, like the, the lower the number, the like not less intelligence, but you may not get the output that you want is expected. So that's kind of part of that process. It's a lot of different things here, but I think, uh, the most important thing is just, um, yeah. How, how do you host this locally, how to start playing around with it?
Um, and that's kind of like a really good primer to get started for doing these models and stuff. Yeah.
Yeah. There is an app called Apollo. Have you heard of that? Apollo.
Showing 81–100 of 138 · page 5 of 7 ← Previous Next →