Rendered at 15:19:59 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
mongrelion 19 hours ago [-]
I have gone down this route but I'm afraid that 16K of context is nothing for agentic coding tasks. Even if you use a smarter model directly from an inference provider to plan all the work and then use your own 16GB GPU to execute the plan, managing those 16k tokens becomes unusable at some point even with a minimalistic harness like pi.
For other agentic tasks it's nice, and more than just the budget, it's the data sovereignty that you gain (imagine analysing tax forms, contracts, other legal documents, etc.) in my opinion.
cyanydeez 2 hours ago [-]
Qwen3.8-Flash-Next+strix halo+halogen server = TAM
Im convinced AI is local now. Halogen loads QFN in 30 GB at 4bits. Its token gen is slightly faster than readable. You can follow along and leave it alone based on its thinking traces.
It gets 256k context but using dynamic context pruning, i got it to loop between 64k and 128k to maximize its speed vs prefill.
Its real. If the memory cartel is busted, the moat is busted.
joshstrange 57 minutes ago [-]
While I'm sure the goalposts will move, if I could get Opus 4.8+ intelligence with 100tps+ and 1M context window I think I'd be happy to switch, mostly, to local inference.
I cannot wait until we get "AI Appliance"-type things we can just plug in and use (aka a server). I know there are a few examples of this but few are plug-and-play currently. I'm thinking more "Here is your hardware Opus 5.5, it costs $X,XXX, it can do 100tps and X parallel streams with a 1M context window".
Maybe I'll change my tune when that's available due to the frontier being, potentially, still 6mo+ ahead with faster inference and larger context windows but I do pine for local inference for at least the "Execute the plan a bigger/better/smarter model wrote for you and the same model will check your work". Currently I'm trialling DS Flash 4.1 as my "worker" model and seeing positive results with it being wrapped by Opus 5.5 High.
kristianp 19 hours ago [-]
> llama.cpp serves it over localhost at speeds that stop being a complaint after a few minutes of use.
"That stop being a complaint"? That's a strange construction.
I wonder if qwen 4 will make these smaller cards more viable by allowing the ngram storage to be hosted on CPU RAM.
mrandish 20 hours ago [-]
Thanks for writing and sharing. Since I also have a 4070 Ti Super, I'm always interested in hearing people's experiences with local models that fit 'middle-ground' hardware. I kind of feel left out because most posts I see seem to either be about clever ways of making older, lower-end cards usable, leveraging mega-CPU RAM (128GB) or how aweseome high-end GPUs like 4090/5090 can be.
wafflemaker 18 hours ago [-]
It's two different things to be honest.
Local model can be a local assistant, take care of your todos, calendar, private stuff in general. Things you shouldn't share with OpenAI or other Big Tech.
Big Tech models for anything else.
MisterMunchkin 17 hours ago [-]
I use the free versions on mobile, and openrouter on desktop. It’s way cheaper than paying a subscription, and I can always use the latest models.
gentile 12 hours ago [-]
OP could be doing a lot better with a MoE model (uses some RAM as well).
I'm doing Qwen 3.6 35BA3B-NVFP4 (a larger ~20GB model) in 8GB VRAM and 20 GB RAM (I also get more than 16k context, and more t/s).
Doxin 3 hours ago [-]
I've got qwen3.6-something-or-other (also around 20GB) in 12GB VRAM with a 64k context, and I'm decently sure I could go to a 128k context without losing too much speed. it's doing around 40 tokens per second, though the initial prompt processing is a bit of a bottleneck.
MoE models are kind of bonkers.
reddit_clone 19 hours ago [-]
Ok. What can I do with an M4 MBP, with 48G memory?
pllbnk 18 hours ago [-]
I am getting ~10 tok/s in M5 Pro with 48 GB with Qwen 3.8 27B. It's workable but very annoying. Also, it's really noisy. On my RTX 5900 I am getting ~200 tok/s and that makes the model truly usable and useful. Though the limit on context size is a huge issue for more serious work.
smcleod 19 hours ago [-]
If you were only paying $20 for LLMs previously then you were getting very low usage / little done.
braggerxyz 20 hours ago [-]
20$/month vs. 1200$ gpu. That's a lot of months you can pay the subscription
Kirby64 20 hours ago [-]
A $1200 GPU you already own, typically, is how this is being seen. Maybe you already have a gaming PC, so this is 0 extra cost to you. All it takes is the power to run the GPU which would be minimal extra vs. the $20/mo cost.
nuancebydefault 19 hours ago [-]
Where I live electricity costs 30 to 40 cents per kWh, I think it's easy to go over 20 dollars per month. Also I find 20 dollars is really a small amount of money to spend on AI. AI fays itself back fast if it helps you with serious stuff.
20 hours ago [-]
bravetraveler 20 hours ago [-]
A lot of quotas, reset timers, changed terms, and so on too through the subscription. I prefer to buy and use my tools, not adopt them as part of my lifestyle.
mrandish 19 hours ago [-]
I bought my 4070 Ti Super in that magic ~month right in between the Crypto-Craze and the AI Bubble when good cards were pretty commonly available for around MSRP. I even found a small deal the day I was ready to buy, so paid $790 with 2-day delivery included.
Throughout the entire crypto-craze I'd been squeeking by gaming on an OC'd 1080 Ti, so when prices finally fell I was more than ready. Since I was getting interested in maybe playing around with AI soon, I spent up from my budget of around $550-$600 and am very glad I did since it's now clear I'll be squeeking by on this card for more years than I'd planned, just like the 1080 Ti.
bsoqk 19 hours ago [-]
Not to mention the price of electricity.
primeagent001 15 hours ago [-]
[flagged]
AIblemblio 7 hours ago [-]
Running QWen costs so much locally, that a $20 subscription costs the same with better quality, significant better but less privacy.
kittikitti 20 hours ago [-]
This is a great article! Thank you for sharing. I also recommend opencode that has built in features for local model hosting.
WarmWash 20 hours ago [-]
"I spent $1100 on a graphics card so I can save $20/mo running a mid/low intelligence model at 33tk/s with 16k context"
simianwords 8 hours ago [-]
I strongly feel that there’s a certain ideological bend towards self hosting and that it clouds judgement.
Dude I can run Astra on my 20 dollar subscription. The leverage with that vs whatever Chinese model you can run in one GPU is not even comparable.
I almost always reach for the most performant model for any of my tasks because my return almost always pays off.
Im convinced AI is local now. Halogen loads QFN in 30 GB at 4bits. Its token gen is slightly faster than readable. You can follow along and leave it alone based on its thinking traces.
It gets 256k context but using dynamic context pruning, i got it to loop between 64k and 128k to maximize its speed vs prefill.
Its real. If the memory cartel is busted, the moat is busted.
I cannot wait until we get "AI Appliance"-type things we can just plug in and use (aka a server). I know there are a few examples of this but few are plug-and-play currently. I'm thinking more "Here is your hardware Opus 5.5, it costs $X,XXX, it can do 100tps and X parallel streams with a 1M context window".
Maybe I'll change my tune when that's available due to the frontier being, potentially, still 6mo+ ahead with faster inference and larger context windows but I do pine for local inference for at least the "Execute the plan a bigger/better/smarter model wrote for you and the same model will check your work". Currently I'm trialling DS Flash 4.1 as my "worker" model and seeing positive results with it being wrapped by Opus 5.5 High.
"That stop being a complaint"? That's a strange construction.
I wonder if qwen 4 will make these smaller cards more viable by allowing the ngram storage to be hosted on CPU RAM.
Local model can be a local assistant, take care of your todos, calendar, private stuff in general. Things you shouldn't share with OpenAI or other Big Tech. Big Tech models for anything else.
I'm doing Qwen 3.6 35BA3B-NVFP4 (a larger ~20GB model) in 8GB VRAM and 20 GB RAM (I also get more than 16k context, and more t/s).
MoE models are kind of bonkers.
Throughout the entire crypto-craze I'd been squeeking by gaming on an OC'd 1080 Ti, so when prices finally fell I was more than ready. Since I was getting interested in maybe playing around with AI soon, I spent up from my budget of around $550-$600 and am very glad I did since it's now clear I'll be squeeking by on this card for more years than I'd planned, just like the 1080 Ti.
Dude I can run Astra on my 20 dollar subscription. The leverage with that vs whatever Chinese model you can run in one GPU is not even comparable.
I almost always reach for the most performant model for any of my tasks because my return almost always pays off.