Rendered at 11:17:05 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
numpad0 14 hours ago [-]
> The cards had some kind of metal cable guide fin on the back
That's to slot into the front PCI brackets for full length cards. There's "full" size for PCI cards and most GPUs are considered "half" length cards. This is separate to Low Profile, so there are four possible slot compatibility configurations, namely HHHL, HHFL, FHFL, and FHHL. Real workstations(including Mac Pro) has slots cut in the case or a bracket to accept the card or that metal piece(which means a standardized solution to GPU sagging had existed even before PCIe was created).
Also, on fans, it's not like blower fans are quieter at all, but in case anyone encounters a situation where turning up fans seem to be only creating noises without moving air or cooling yhe cards, it might be worth remembering that radial fans have higher static pressure and are better at forcing air through.
edit: ps: people recreating this might want to know what case this is. The vast majority of ATX cases, even unnecessarily big ones, only has 7 PCI slots. Cases that can take four dual-slot cards is rare.
MrDrMcCoy 4 hours ago [-]
I'm not aware of any motherboards, even server boards, that have more then 7 slots. Ditto for rackmount chassis. I'm sure they exist. If any do that are standard [E-]ATX, I'd love a link.
numpad0 1 hours ago [-]
The 4th GPU on the 7th slot occupy the I/O openings 7 and 8. There are few _cases_ with that extra opening to support the quad GPU use case, rare as the ones with the front support cutouts. The case in the article has that slot but model number isn't mentioned, and some people might want to know the make/model of that one.
I'm not aware of an 8-slot ATX board either, but there are custom shaped boards for GPU servers. Also in crypto space - btw I don't understand why those shops with inventories of that certain gold colored NVIDIA crypto cards aren't just shipping the whole complete chassis with appropriate price tags at this point, or just have that chassis listed for estimate/RFQ price. There should be some demands in a single node GLM5.3 box.
psd1 2 hours ago [-]
You don't need 8 mobo shots: 4 in use, 3 interstitial. {Although my experience with retail mobos is that your only get two x16.)
The case, otoh, sounds like a dremel job
Aurornis 18 hours ago [-]
> I got these 10,000 RPM fans because I wanted to make sure I was moving enough air, and they were the same price as slower fans. They are really loud! I wanted the motherboard to control their speed based on the GPU temperature, and this didn't work at first, so I just wore ear protection during initial setup.
I think everyone remembers the first time they went from "I think I can tolerate some server fans. How loud can they be?" to "I had no idea a small 12V fan could be this loud"
nazgulsenpai 17 hours ago [-]
The first time I powered on a Dell Poweredge R420 without the lid on
eichin 15 hours ago [-]
mmm the early Dell 2U servers (2003 or so?) didn't have very good linux support, so the fans always ran at full speed. I think the happiest I've ever seen people about a kernel upgrade was (2.5 era?) when it got fan support for that hardware and finally it was still loud at boot, but once the kernel came up it switched to thermally managed and slowed way down :-)
m463 8 hours ago [-]
this is when you start thinking... well if a 40mm fan is jet-engine-loud, and an 80mm fan is just leaf-blower loud, and 140mm still spins too fast... maybe I can use a box fan, cardboard, duct tape...
(personally, I have wondered why people didn't do pc system cooling with a car radiator with huge slow fan, a giant reservoir like 5-gal homer bucket with lid or a 55gal drum, and maybe an submersible aquarium pump or two.)
dereke 5 hours ago [-]
In the late 90s I had this. Drilled a hole in my bedroom floor to put the pipes under the house (my dad didn’t seem to mind!). Found a local fabricator that let me use his mill to make a water cooling block. Good times.
jjav 3 hours ago [-]
> How loud can they be?
LOL.. for home use I go either fanless (very low power), or 4U chassis even if I don't need the space, just to fit large quiet fans.
1U or 2U at home with fans is not what you want.
comandillos 17 hours ago [-]
I ended up buying 2 DGX Sparks interconnected over QSFP with the intention of getting rid (as much as I could) of any cloud-based AI provider. I'm running DS4 Flash 0731 on it and some OCR models, using Oh My Pi and OpenWebUI as my main ways to interface with the agent... and from someone that has been using Claude for a long time, I can definitely say I don't need it anymore. Not for the stuff I'm doing.
nater5000 16 hours ago [-]
I mean, this is cool and all, but 2 DGX Sparks is like $10k, right? That's a huge upfront cost to run DS4 Flash 0731 which is definitely not comparable to higher-end Claude models. It's extra egregious when you consider how much it'd cost to run DS4 Flash 0731 via an API (how long would it take to run up a $10k bill doing what you're doing?).
>and from someone that has been using Claude for a long time, I can definitely say I don't need it anymore. Not for the stuff I'm doing.
I'm curious what this implies. It suggests that you were using Claude prior to this setup, so it doesn't seem like you were limited by security or local features? I can understand that someone not wanting to send their data to Anthropic (etc.) might be willing to pay for this kind of setup to accomplish that, but what was your motivation for doing this?
If it matters, if I had enough money that I could blow $10k on 2 DGX Sparks without it being a significant cost I probably would for the hell of it, so I'm not digging on you if this is ultimately what this comes down to. But, as an investment, this doesn't seem like a very good deal.
comandillos 16 hours ago [-]
Absolutely agreed, I don't see this as an investment leading to cost-savings in the future, at all.
I see this rather as an investment in improving my capabilities and knowledge of this technology, letting me play with a 'GPU cluster', vLLM and other technologies that otherwise would require me to rent GPUs on the cloud.
jrm4 15 hours ago [-]
I cannot stress enough how good of an instinct I think this is; many of my long term nerd decisions over the last 20 years or so have "paid off" in ways that I literally don't know how to put a money value on. Not priceless, but freeing me from so many headaches.
Offhand, these include "preferring text and my own machines for storing information that is useful to me" over "clouds" and e.g. Word docs. Also, having my own domain and email (which I pay for).
I'm thinking of, e.g. the guy that blogged for years and years and then Google just yoinked it and it was gone. Younger me was more probably more obnoxious about it, like, serves you right -- a thing I would NEVER say now -- but, still, people like me really were right all along about this sort of thing, and I daresay it would be good if everyone followed us.
pluralmonad 15 hours ago [-]
It amuses me to no end that colleagues have started using markdown for notes and documents. Thanks AI.
roarcher 6 hours ago [-]
I've been sending markdown documents to my non-technical coworkers for years in the hopes that they would come to appreciate the simplicity, but they insist on futzing around with Word documents and their fragile layouts. Maybe AI will finally convert everyone to the Church of Plain Text.
belval 16 hours ago [-]
Sure but the GPUs in the article are from 2021 and cost ~500-600$ a pop. The OP bought 4 so that's 2k, let's say another 500$ for the rest of the chassis and you stand at $2.5k for a monster server that will burn electricity and warm your house.
Compare that to a $5k DGX than consumes much less and has the same amount of VRAM and there is a real question as whether this is worth doing at all (well aside from the cool factor).
kQq9oHeAz6wLLS 12 hours ago [-]
> Compare that to a $5k DGX
Times 2, because OP said they had two of them. That's considerably more.
fwipsy 10 hours ago [-]
Each spark has as much VRAM as the server built in the article.
eek2121 13 hours ago [-]
Everyone focuses on the frontier models, and they act like everything else is useless. The gap between say, Qwen 3.8 and the frontier models is not as large as you assume, and things on the local front have made significant gains in the past year. That is WHY Anthropic, OpenAI, etc. are worried. They NEED customers to pay top dime for top performance, but if you can spend 95% less money for 95% of the performance of the latest, bleeding edge model from Anthropic, you know which one you'd pick. Most folks don't need that extra 5-10%, especially since the cheaper models will catch up anyway.
fwipsy 10 hours ago [-]
Are you sure that AI hasn't just saturated your personal benchmark? Maybe you're just not asking it to do things which showcase its full capability. Open-source models will catch up for any given use case, but the frontier is interesting because of the possibility it will keep opening up new applications.
esseph 13 hours ago [-]
> But, as an investment, this doesn't seem like a very good deal.
Same thing for garage mechanics with a "fun car".
TacticalCoder 14 hours ago [-]
> That's a huge upfront cost to run DS4 Flash 0731 which is definitely not comparable to higher-end Claude models.
So higher-end Claude models (I've got a subscription btw) do work fine right?
What makes you think that in six months he cannot swap DS4 Flash 0731 for another model that could be equivalent to todays' top OpenAI/Anthropic models?
Or are you going to explain in six months that, after all, the top models from Anthropic from today are unfit for use?
rsolva 15 hours ago [-]
I love this! Hacky recycling of old computer parts, home-cooked fan controllers and a website that is just plain HTML with a stylesheet that is so short I don't even have to scroll to read it all. Just write part 2 already, I want to know what you can squeeze out of those AMD cards <3
fabiensanglard 17 hours ago [-]
Quite courageous to go with AMD. I am very curious to see Chapter 2 and how he solves the software stability that used to plague AMD AI applications.
SwellJoe 16 hours ago [-]
"used to" in your sentence already answers your curiosity.
ROCm runs most things just fine without any ceremony or difficulty. I have a couple of the same cards as covered in the article, bought before they got expensive (I'd recommend current Radeon AI Pro 9700 cards over the V620 now, though, as they have increased in price since I bought mine), and they're at the long end of the supported chart for ROCm...nearly EOL. But, they currently work great with llama.cpp and current ROCm 7.14. They're pretty fast and stable. I also have a Strix Halo 128GB, and while the Strix Halo can run bigger models, the dedicated GPUs are quite a bit faster. And, since the best models you can run at home are probably Qwen 3.6 27B and Gemma 4 31B, and those both fit comfortably on dual 32GB GPUs, I find I use the desktop more often than the Strix Halo.
Anyway, there's very little reason to spend 3x or more for the Nvidia ecosystem these days. The only exception is the Strix Halo vs the Nvidia GB10 platform. Strix Halo was a no-brainer when it was half the price, but it's risen in price to be almost the same price as an Asus GX10. In that one instance, I think the Nvidia based Asus is a better choice. The Strix Halo is almost too slow to make use of 128GB. Some MoE models are comfortable, like Laguna S2.1, so it's probable that there will someday be an MoE model that is better then Qwen 3.6 27B or Gemma 4 31B that won't run on 64GB but will run on 128GB. I don't know of one, yet, though.
AMD also used to focus almost entirely on their data center line and not pay any attention to their lower end stuff for more advanced AI features, but that's been changing, and ROCm support has broadened to include almost everything AMD ships now, including the smaller embedded stuff. I think they've realized that as long as the only way to develop AI for AMD was to have a quarter million dollars worth of hardware, they would always fall behind a platform that can be developed for and tested on consumer hardware.
Muromec 16 hours ago [-]
>I'd recommend current Radeon AI Pro 9700 cards over the V620 now
That's about 4x the price of v620, which is 450 eurobucks. Can't imagine buying 4 of those at 4x the price.
SwellJoe 15 hours ago [-]
Ah, the price has gone up on the AI Pro 9700, again, I see. So, yeah, it's a worse deal now. But, the V620 is not really a good deal, either, at $450 on eBay. It's roughly half the performance of the newer card. Usable but not blazing. And, you probably don't want to deal with cooling them; if you already have a 3D printer, you can print the shrouds for a few bucks and buy the fans for a few more bucks. But, if you don't have a 3D printer, the shrouds you can buy on eBay are for the tiny fans that have to run at ludicrous speed and have to be so fucking loud it should be illegal. I went through three different shroud and fan combos before finding a set that worked in my case with this cards (about $100 worth of experimentation, probably).
I should say, though, that I actually don't recommend buying anything right now. I wrote up my setup and made recommendations (and the main recommendation was "don't"). https://swelljoe.com/post/how-i-run-local-llms/
The V620 at $450 or the Radeon AI Pro 9700 at ~$1400 are a good deal compared to everything else right now, but buying tokens from DeepSeek is a better deal. DeepSeek V4 Flash 0731 is better than anything you can host locally, they'll serve it to you at blistering fast speeds for pennies a day, and with 1 million token context. You could host a 2-bit quantization of it on a Strix Halo or four of these V620s, and it would run at a crawl on either one. The Strix Halo gets 9-13 t/s. Four V620s would, I guess, get two-three times that. Which is still too slow for comfortable interactive agentic use, and much slower than getting it from DeepSeek, and at two bits there is measurable loss. You're paying a lot more for self-hosted and you're getting worse models.
Also, if you want four cards, you need a server-class motherboard and CPU and RAM. More money.
It's all just a bad investment. Self-hosting is a bad idea if you don't already have the hardware, unless and until memory and GPU prices come down. Even at the prices I paid (before RAMpocalypse really kicked off; $2k for the Strix Halo, ~$350 for each V620), I wouldn't recommend it if you don't have a strong urge to tinker with hardware and it'll probably never pay for itself vs. buying inference from DeepSeek directly.
Edit: The same shrouds for all the old Instinct cards also work for the V620, as they have the same dimensions and screw and cable layout. I tried several and ended up with this one: https://www.thingiverse.com/thing:7296707
matusnovak 14 hours ago [-]
I totally agree. I have a very similar setup as OP. I have 4x AMD Instinct Mi50. They are 32GB each with 1TB/s memory bandwidth. I got very lucky and bought them at 220 euro a piece back when prices were sane. Now they cost almost tripple of that. For the server I bought a used X99 xeon that has enough PCIE lanes for all four GPUs, with ASUS X99 WS motherboard. The CPU was only ~22 usd from Aliexpress. The build made some small sense when prices were good. I just did it for fun. Now it makes no sense at all.
Muromec 15 hours ago [-]
I agree with your overall and also use DS4. Economically it makes no sense to run locally (electricity alone would be more expensive than what I pay our brothers in communist faith). But as a matter of fact I do have a tinkering urge and a 3d printer, so I better stay away from this stuff.
cj00 15 hours ago [-]
For shits and giggles I had Claude build out a full voice cloning pipeline that runs completely locally on my Steam Deck. I’ve got Gemma e4b, Qwen and Chatterbox all running on the AMD.
esseph 13 hours ago [-]
lemon-server is amazing and gets you the right models for your hardware, along with a model router
Uhh.... I have a box with two AMD Radeon AI PRO R9700 32GB.
It was the most trouble-free AI setup that I have done. The driver worked right out of the box with the stock Fedora kernel, ROCM can be installed from the regular repo, and most AI software has ROCM builds by now.
_ache_ 12 hours ago [-]
What can you do with "only" 64G of VRAM that a 32G can't?
Also, the R9700 are so loud!
cyberax 12 hours ago [-]
Recently: running Laguna-S with reasonable speed. Before that, I was able to run multiple Qwens or a Qwen and several copies of a smaller model for subagents.
The cards that I have are not loud at all. The cards are capped at 300W each, so that's really not that much heat to move.
hk1337 17 hours ago [-]
Our definition of "Box of Scraps" is wildly different.
matusnovak 14 hours ago [-]
I did a very similar setup with 4x AMD Instinct Mi50. For the air intake I used a single Noctua Industrial PPC 14cm fan, with a 3d printed duct. In the end the cards don't produce that much heat. A single 14cm fan at low speeds was enough to keep them cool at 62 degrees max. A large fan is also more quiet. Also, removing the graphene pad betwee the die and the heatsink, and replacing it with a good paste, helped bring down the temperature by few degrees.
NamlchakKhandro 10 hours ago [-]
a heat stack chimney would make it silent
rcarmo 16 hours ago [-]
Those are some pretty meaty scraps, way above what you’d have on several people’s shelves…
petesergeant 17 hours ago [-]
> Around this time I also realized that I didn't have Ethernet in the garage, so I started cutting holes in the drywall at 11:00 PM.
Ahh, one of those projects
ryandrake 17 hours ago [-]
We've all been there. I don't think I've ever owned a home where I didn't have to cut holes in the drywall to run cables.
Muromec 16 hours ago [-]
The only one place where I need ethernet and don't have it is right on the opposite wall from the electric closet that has the fiber. That's also exactly where I have my desk.
Yes, I was the person drawing the floor plans and forgot the electric socket there too.
monksy 6 hours ago [-]
Shame the guy doesn't have an RSS feed? I want to read the next one.
I was recommended the x299. I didn't go with that and I've been struggling with the AM5.
_def 17 hours ago [-]
> You'll start depending on it and then it'll get taken away from you.
You can not get around this, conceptually. Local inference will keep depend on updated models for quite a while. Partly because they will contain outdated training data, partly because of demands for the improved models. And it's still not clear where this will lead us. We're still in the rosey phase where people get lured in.
ryandrake 17 hours ago [-]
But, we know 100% that cloud-provided anything can and will get nerfed, broken, removed, or in some other way rug-pulled. It already happens all the time, and companies are getting more and more aggressive/stingy about what counts as "yours" and what counts as "bought" and what counts as "acceptable use".
Yes, bringing everything local still means you need to download things, maybe over and over. But once it is on your machine, nobody can yoink it from you just because they want more money or they don't like what you're doing with it.
ericd 10 hours ago [-]
Hermes injects the date/time into the stream, and you tell it to web search for anything newer than the training cutoff date. Works like a charm with deepseek flash.
zdragnar 16 hours ago [-]
Is there any hope that something like unsloth studio will let people retrain models continually so that, when the day comes that there are no good open models being released, something like the final generation of qwen whatever can continue being relevant into the future?
zeeveener 15 hours ago [-]
It would likely fall on the harness (and possibly local datasets) to encourage that the model reach out for up to date information instead of relying on it's own internal "knowledge".
We already see a lot of that with `web_search` tooling, so I imagine it would just become more essential to have tools like that.
zdragnar 14 hours ago [-]
Maybe this is just my lack of experience thinking, but that seems really unlikely to scale without retraining. The context will balloon with updated syntax for languages and APIs for libraries and so forth. Imagine if, for example, something on the scale of custom elements / web components were introduced in this post-open world. You couldn't fit all the information needed in a local model's context to tell it how to write a new custom element of any real complexity, especially if it needed to interact with the new post-release web transport specification to interact with a new language's client for a new database.
esseph 13 hours ago [-]
You can just make archives of the web and language specs, best practices, etc and RAG it via a pipeline every month or so.
edgyquant 16 hours ago [-]
Updated models isn’t the same thing as hosting them somewhere else
hypfer 16 hours ago [-]
We will try anyway. No need to try to FUD people into passive acceptance of the cloud.
Muromec 16 hours ago [-]
One piece of v620 costs about 450 eurobucks on ebay right now. Weird to see a card with no HDMI output at all.
monksy 6 hours ago [-]
I got mine below $400.
16 hours ago [-]
TacticalCoder 14 hours ago [-]
> I printed this out of carbon fiber ASA, but probably boring old PLA would have worked just fine.
I doubt it: PLA begins to soften at as low as 55 degrees Celsius. If there's weight on it, it's worse: the piece shall quickly deform and become unfit for its purpose.
This rig looks like it means business: I think PLA would fail.
If, like me, you cannot print ASA (say because you've got a printer that is not closed and that won't heat enough), then PETG-CF (PETG reinforced with some carbon fiber: it's got better heat deflection than plain PETG) is a safer bet than PLA for parts to put in PCs/servers/rack. Moreover PETG-CF do look really good.
SchemaLoad 12 hours ago [-]
The fan shrouds likely would have been fine. The hot spot on electronics is generally very localised to the chips and drops off sharply as you get further from them. I've printed a 10" home server rack out of PLA which is warm and under constant load, and it's been running for around a year now with no deformation.
A good rule of thumb is if touching it would cause burns, it's too hot for PLA, otherwise it's likely ok.
That's to slot into the front PCI brackets for full length cards. There's "full" size for PCI cards and most GPUs are considered "half" length cards. This is separate to Low Profile, so there are four possible slot compatibility configurations, namely HHHL, HHFL, FHFL, and FHHL. Real workstations(including Mac Pro) has slots cut in the case or a bracket to accept the card or that metal piece(which means a standardized solution to GPU sagging had existed even before PCIe was created).
Also, on fans, it's not like blower fans are quieter at all, but in case anyone encounters a situation where turning up fans seem to be only creating noises without moving air or cooling yhe cards, it might be worth remembering that radial fans have higher static pressure and are better at forcing air through.
edit: ps: people recreating this might want to know what case this is. The vast majority of ATX cases, even unnecessarily big ones, only has 7 PCI slots. Cases that can take four dual-slot cards is rare.
I'm not aware of an 8-slot ATX board either, but there are custom shaped boards for GPU servers. Also in crypto space - btw I don't understand why those shops with inventories of that certain gold colored NVIDIA crypto cards aren't just shipping the whole complete chassis with appropriate price tags at this point, or just have that chassis listed for estimate/RFQ price. There should be some demands in a single node GLM5.3 box.
The case, otoh, sounds like a dremel job
I think everyone remembers the first time they went from "I think I can tolerate some server fans. How loud can they be?" to "I had no idea a small 12V fan could be this loud"
(personally, I have wondered why people didn't do pc system cooling with a car radiator with huge slow fan, a giant reservoir like 5-gal homer bucket with lid or a 55gal drum, and maybe an submersible aquarium pump or two.)
LOL.. for home use I go either fanless (very low power), or 4U chassis even if I don't need the space, just to fit large quiet fans.
1U or 2U at home with fans is not what you want.
>and from someone that has been using Claude for a long time, I can definitely say I don't need it anymore. Not for the stuff I'm doing.
I'm curious what this implies. It suggests that you were using Claude prior to this setup, so it doesn't seem like you were limited by security or local features? I can understand that someone not wanting to send their data to Anthropic (etc.) might be willing to pay for this kind of setup to accomplish that, but what was your motivation for doing this?
If it matters, if I had enough money that I could blow $10k on 2 DGX Sparks without it being a significant cost I probably would for the hell of it, so I'm not digging on you if this is ultimately what this comes down to. But, as an investment, this doesn't seem like a very good deal.
I see this rather as an investment in improving my capabilities and knowledge of this technology, letting me play with a 'GPU cluster', vLLM and other technologies that otherwise would require me to rent GPUs on the cloud.
Offhand, these include "preferring text and my own machines for storing information that is useful to me" over "clouds" and e.g. Word docs. Also, having my own domain and email (which I pay for).
I'm thinking of, e.g. the guy that blogged for years and years and then Google just yoinked it and it was gone. Younger me was more probably more obnoxious about it, like, serves you right -- a thing I would NEVER say now -- but, still, people like me really were right all along about this sort of thing, and I daresay it would be good if everyone followed us.
Compare that to a $5k DGX than consumes much less and has the same amount of VRAM and there is a real question as whether this is worth doing at all (well aside from the cool factor).
Times 2, because OP said they had two of them. That's considerably more.
Same thing for garage mechanics with a "fun car".
So higher-end Claude models (I've got a subscription btw) do work fine right?
What makes you think that in six months he cannot swap DS4 Flash 0731 for another model that could be equivalent to todays' top OpenAI/Anthropic models?
Or are you going to explain in six months that, after all, the top models from Anthropic from today are unfit for use?
ROCm runs most things just fine without any ceremony or difficulty. I have a couple of the same cards as covered in the article, bought before they got expensive (I'd recommend current Radeon AI Pro 9700 cards over the V620 now, though, as they have increased in price since I bought mine), and they're at the long end of the supported chart for ROCm...nearly EOL. But, they currently work great with llama.cpp and current ROCm 7.14. They're pretty fast and stable. I also have a Strix Halo 128GB, and while the Strix Halo can run bigger models, the dedicated GPUs are quite a bit faster. And, since the best models you can run at home are probably Qwen 3.6 27B and Gemma 4 31B, and those both fit comfortably on dual 32GB GPUs, I find I use the desktop more often than the Strix Halo.
Anyway, there's very little reason to spend 3x or more for the Nvidia ecosystem these days. The only exception is the Strix Halo vs the Nvidia GB10 platform. Strix Halo was a no-brainer when it was half the price, but it's risen in price to be almost the same price as an Asus GX10. In that one instance, I think the Nvidia based Asus is a better choice. The Strix Halo is almost too slow to make use of 128GB. Some MoE models are comfortable, like Laguna S2.1, so it's probable that there will someday be an MoE model that is better then Qwen 3.6 27B or Gemma 4 31B that won't run on 64GB but will run on 128GB. I don't know of one, yet, though.
AMD also used to focus almost entirely on their data center line and not pay any attention to their lower end stuff for more advanced AI features, but that's been changing, and ROCm support has broadened to include almost everything AMD ships now, including the smaller embedded stuff. I think they've realized that as long as the only way to develop AI for AMD was to have a quarter million dollars worth of hardware, they would always fall behind a platform that can be developed for and tested on consumer hardware.
That's about 4x the price of v620, which is 450 eurobucks. Can't imagine buying 4 of those at 4x the price.
I should say, though, that I actually don't recommend buying anything right now. I wrote up my setup and made recommendations (and the main recommendation was "don't"). https://swelljoe.com/post/how-i-run-local-llms/
The V620 at $450 or the Radeon AI Pro 9700 at ~$1400 are a good deal compared to everything else right now, but buying tokens from DeepSeek is a better deal. DeepSeek V4 Flash 0731 is better than anything you can host locally, they'll serve it to you at blistering fast speeds for pennies a day, and with 1 million token context. You could host a 2-bit quantization of it on a Strix Halo or four of these V620s, and it would run at a crawl on either one. The Strix Halo gets 9-13 t/s. Four V620s would, I guess, get two-three times that. Which is still too slow for comfortable interactive agentic use, and much slower than getting it from DeepSeek, and at two bits there is measurable loss. You're paying a lot more for self-hosted and you're getting worse models.
Also, if you want four cards, you need a server-class motherboard and CPU and RAM. More money.
It's all just a bad investment. Self-hosting is a bad idea if you don't already have the hardware, unless and until memory and GPU prices come down. Even at the prices I paid (before RAMpocalypse really kicked off; $2k for the Strix Halo, ~$350 for each V620), I wouldn't recommend it if you don't have a strong urge to tinker with hardware and it'll probably never pay for itself vs. buying inference from DeepSeek directly.
Edit: The same shrouds for all the old Instinct cards also work for the V620, as they have the same dimensions and screw and cable layout. I tried several and ended up with this one: https://www.thingiverse.com/thing:7296707
Add openwebui and/or pi and away you go.
https://lemonade-server.ai/
It was the most trouble-free AI setup that I have done. The driver worked right out of the box with the stock Fedora kernel, ROCM can be installed from the regular repo, and most AI software has ROCM builds by now.
The cards that I have are not loud at all. The cards are capped at 300W each, so that's really not that much heat to move.
Ahh, one of those projects
Yes, I was the person drawing the floor plans and forgot the electric socket there too.
I was recommended the x299. I didn't go with that and I've been struggling with the AM5.
You can not get around this, conceptually. Local inference will keep depend on updated models for quite a while. Partly because they will contain outdated training data, partly because of demands for the improved models. And it's still not clear where this will lead us. We're still in the rosey phase where people get lured in.
Yes, bringing everything local still means you need to download things, maybe over and over. But once it is on your machine, nobody can yoink it from you just because they want more money or they don't like what you're doing with it.
We already see a lot of that with `web_search` tooling, so I imagine it would just become more essential to have tools like that.
I doubt it: PLA begins to soften at as low as 55 degrees Celsius. If there's weight on it, it's worse: the piece shall quickly deform and become unfit for its purpose.
This rig looks like it means business: I think PLA would fail.
If, like me, you cannot print ASA (say because you've got a printer that is not closed and that won't heat enough), then PETG-CF (PETG reinforced with some carbon fiber: it's got better heat deflection than plain PETG) is a safer bet than PLA for parts to put in PCs/servers/rack. Moreover PETG-CF do look really good.
A good rule of thumb is if touching it would cause burns, it's too hot for PLA, otherwise it's likely ok.