Excellent video, my friend! I've been working with both local models, like Qwen3.5 35B-AB that you mentioned, as well as free open-source, open-weight cloud-based models, like GLM5.3-Flash, Kimi-K3, and Nemotron-3-Ultra in Hermes Agent on my Lenovo Legion 5 gaming laptop, though my main limitation with local models on this system is that I only have 8GB of VRAM, so I need to do CPU/RAM offloading to fit the larger models, and since I'm always doing a lot on my laptop, it's somewhat challenging. I only dove into all this in the spring, and I've come a very long way since then, and I'm just getting started! Local sovereign AI stacks are the way to go! 😁🙏💚✨🤙
Posted via HiveSuite
You can still run some decent small models on that 8GB of VRAM if you are offloading the context to system memory/swap. Check out Bonsai 27b. You can get the 1bit that will run on your machine and it's actually pretty decent, I was actually shocked. I can run the 2 bit on mine. But it also does in some cases make sense to run the cloud models too. Like I will use Grok for like current event news analysis because it's models are retrained basically daily and it has the X access which helps with market sentiment and breaking news info and such. But I don't pay for it, lol.
Yes, that's definitely true, I was just frying my brain for a while testing hordes of build and run flags for llama.cpp and several of its forks across multiple different models, including the 1-bit Bonzai, though I couldn't get its tok/s to a usable range. I was also using Gemma-4 26B-A4B QAT, which was also pretty good overall, but Qwen is still generally better. I still haven't worked with the closed-source, closed-weight models yet, though a coding friend gave me access to his Claude Code account, which I still have to try out. Honestly, I've been so busy using the solid Chinese open-weight models with Hermes, that I haven't really had the time, or even the need to use it yet. X is one of the absolute best sources for real news, and also learning about agentic-AI tech, among lots of other domains, and I would love to be able to make use of all that information. You don't pay to use Grok? You're very fortunate, my friend! 😁🙏💚✨🤙
I haven't tested it enough yet, but Ling3.0-tiny seems like a great fit for 8B VRAM... There's also Gemma4-12B that could fit in there without offloading, but from my testing, it's not as good as the MoE models.
Yeah, you need to be very direct with the small dense models like that. They are really decent when given a task they don't have to think to much about.
I have heard some good things about the Ling models, though I haven't tried them much myself yet. I could only really run the smaller dense models on my GPU, but the MOEs, that can be partially offloaded, work fairly well. The first local model that I used for agentic work in Hermes was Gemma-4 E4B, which for its size was pretty good, though it definitely couldn't compare to its larger Gemma-4 26B-A4B sibling. Thanks a lot for the info, my friend, I appreciate it. 😁🙏💚✨🤙