A 125-billion-parameter mixture-of-experts model, running fully local — on a single AMD Radeon AI PRO R9700. No NVIDIA. No cloud API. Just ROCm, llama.cpp, and one GPU. In this video I walk through the entire setup.
#opensource #llm #llamacpp #qwen3 #r9700 #amd
👉 LLM router AGIEverywhere.com (have more than 10 AI endpoints)
👉ⓢⓤⓑⓢⓒⓡⓘⓑⓔ
👉 !! try Wan Video online at https://agireact.com/wan-t2v !!
#opensource #llm #llamacpp #qwen3 #r9700 #amd
👉 LLM router AGIEverywhere.com (have more than 10 AI endpoints)
👉ⓢⓤⓑⓢⓒⓡⓘⓑⓔ
👉 !! try Wan Video online at https://agireact.com/wan-t2v !!
3090 vs 5090 running Qwen3.8-27B: https://youtu.be/agDPSAO8gZQ
If you found this useful:
👍 Like if the results surprised you
🔔 Subscribe for more local AI benchmarks and hardware deep-dives
💬 Drop your setup in the comments — curious what you’re running models on
🖥️ Hardware Tested:
– AMD R9700 (32GB VRAM)
🤖 Models Benchmarked:
– Qwen3.8-Flash-Next (Q3)
For Gemma4 model comparison, see https://youtu.be/VYc47oqBnqI
Please join the discord server at https://discord.gg/SgmBydQ2Mn where you developed free chatgpt bot and stable diffusion bot!
If you would like to support me, here is …
![]()