You may find the following video useful to understand what various hardware can run when it comes to local AI.
Source: YouTube
You may find the following video useful to understand what various hardware can run when it comes to local AI.
Source: YouTube
I understand a bit better now. I bought my PC this year, but the high cost of memory limited what I got. I also need to work out what I would actually do with these tools. Generating images, video and music doesn't really interest me, but I might use them for coding.
How much VRAM does your system have, my friend? Using llama.cpp, or one of its forks like ik_llama.cpp, llama-cpp-turboquant, or buun-llama-cpp, it's feasible to run much larger MOE (Mixture of Experts) models than would normally be expected, since part of the experts are kept in RAM/CPU. Models like Qwen3.5 35B-A3B or Gemma-4 26B-A44B are very doable even with 6-8GB of VRAM. And as to coding, that's what I'm doing. I've been working on an autonomous liquidity-pool agent, to assist me in my asset-managing tasks. 😁🙏💚✨🤙
I have an 8GB RTX5060. I got it to help with things like video editing. I will look at this stuff when I find the time.