AI means it has npu, Max+ is marking the memory channel, PRO is a normal label for chips that have extra security baked in, it has been this way since forever.
I typed my (masked) email into the input, pressed tab and enter, it redirected me to some cloudflare captcha marketing site. I thought at first it didn't want the masked email, but then it worked when I clicked with my mouse. Apparently there are 3 invisible focus targets in between!
Dear people who create websites, these things are important, they should work!
I have a framework desktop w/ 128GB that I bought last Christmas and if I’m looking at it right it costs $2000 (CAD) more now because of the RAM shortage (and in any event is apparently out of stock). Would love to have 192 GB but I’m not sure I can justify buying another.
Does anyone have a sense of how this might progress, e.g if I can get a 256 or 512 GB in a year if I wait. In any case I’m jealous this exists and I don’t have one.
One last thing, I assume this isn’t exclusive and there will be other builds with this same config same as current Strix Halo?
They are estimating a 25% shortfall in supply remaining in 2030, even with Chinese companies ramping up DDR5 supply.
It's not looking great - the RAM producers need more of the same machines that other semi-conductor manufactures need and the suppliers of those seem unable to increase production.
Are you generally happy with it? What do you mainly use it for?
I'm kinda tempted by the frameworks because they are somewhat energy efficient and compact (which I both value highly!) but my concern is that the GPU is on the low end for 1440p gaming (while being barely price-competitive with a self-built Ryzen9950 + 9070XT combination).
what I read in I think the direct AMD press, it's essentially like a 10% upgrade to compute but mostly ram.
I agree, this is the direction AMD was going in to larger attached memory and they got hit by the memory cartel pricing; They likely were going to hit 256GB instead of this weird glitch in the sizing.
So, yes, of course they're going to hit higher memory sizes; but since they models are meant for laptops and to get the speed you want for inference, they're soldered, you have no real options.
Framework in particular might have a high value on ebay as I assume this will be a drop in replacement for their existing motherboard.
I understand their reasoning but it’s still a shame Framework opted for a proprietary motherboard instead of mini-ITX with a socket CPU and GPU. I value repairability over a little bit of extra performance.
The whole point is that these APUs present a pretty unique value proposition (lowish performance GPUs with massive amounts of RAM attached), and framework is afaik the only vendor shipping them on standard mini-ITX boards
Framework is going for a computer OEM market where they need to make margin, hence the soldered-on components and relatively limited options. People in the pre-built computer market overwhelmingly don't care about replaceable components, because anyone else will just buy components and build a PC, for both the greater market of options AND overall lower price than buying pre-built (if they don't overspec).
You literally can't make this machine without soldered components. You don't understand this computer's design goals. It's been 1.5 years since they announced the Framework Desktop. I don't get how people are still confused about this.
Tl;dr: think hard about what you’d use this much RAM for in a desktop setting and do some research about how people like the 395 for your use case. Not all use cases work well.
Anyone who thinks they are going to serve some 100+ GB LLM locally, remember that memory bandwidth becomes a key limitation for large models. While you might be able to load a model, token generation can be very slow. MoE models like Qwen3.5-122B-A10B work decently fast, but dense models of a decent size are slow and you won’t want to use them.
I’ve got a 395 system, and found that I’m quite happy with Qwen3.6-35B-A3B, generating at around 50 t/s, but the dense 27B model is 20-25 t/s and that’s the lower limit I’m willing to tolerate. So a 70b dense model is just not going to happen. That means you can’t really use that much RAM.
A reasonable use case is to have multiple smaller models loaded- you can have an image generation model loaded along with the text model. Or you can use this computer for development simultaneously with serving LLMs. Those ideas work okay. But trying to load up a single giant model is going to test your patience.
Im less interested in the memory than the memory bandwidth. The current system with 128GB can load pretty big models, but its meaningless unless you want to wait 40 minutes per prompt.
I've had best results with Qwen3.6-35B-A3B, which uses 40GB of memory, but only uses 3 billion parameters per token which helps with throughput.
Until memory bandwidth significantly improves I just can't see myself wanting to use all that memory. Unless it's just to keep a wide variety of models in memory.
I run Qwen3-Coder-Next-UD-Q4_K_XL and other than the initial wait to initialise context (which takes less than 2 minutes) subsequent prompts return in less than a minute, usually less than 30s.
If your performance is significantly slower then you are probably doing it in CPU - there was some fiddling required to get it to use GPU (I use llama.cpp)
I have an NVIDIA Spark thingy (the ASUS one) and it's the same problem there, though the prefill side is superior to the Ryzen ones.
But I actually think 128GB is too little. There are some compelling models that are above what can fit in that at reasonable quants (e.g. DeepSeek V4 Flash) but could if the system was 256GB.
If RAM prices weren't so f*cked I think we'd be seeing 256GB and even 512GB unified memory systems becoming quite common. As it is I think it will be 10 years before >128GB becomes feasible on a regular consumer machine for normal people again.
You can run DeepSeek Flash just fine on 128GB using lower quants. Antirez' DwarfStar supports it really well. The bigger the model, it seems, the less it is affected by high quantization. I know some people are working on quantizing GLM 5.2 so it can fit into 128GB (using additional tricks to select only some experts, etc). There's a lot going on.
I use exactly the setup you describe; it can handle 1-3 agents working if the agent work has IO delays. with MTP models, prefill is quite fast.
Not sure what damage you have, but no one waits 40 minutes per prompt or even a minute; The A3B model loads within 10 seconds and a simple response in opencode is maybe at most a minute, then catches up in 3-5second bursts depending on IO.
I put dynamic context pruning into opencode and tweaked it for 45k-85k context before it shrinks; this lets me get into 500k token sizes and fixed on medium sized github repos.
There's a chance this comment is a skill error:
1. USe llamacpp with a MTP model
2. Use reasoning-budget and reasoning-message
3. Tailor your agent to use the reasoning-message to use subagents and dyanmic compaction.
---
Now you have a reasonable coding agent for cloning, building and extending any github project I've seen so far.
Woof. What's that going to cost? And will we actually be able to buy one, because the 128GB model has been sold out pretty much since the LLM craze started...
Memory bandwidth certainly is the bottleneck on my 395+ 128GB. It's been a great server. Think I'll wait another year maybe hopefully the prices come down and there's one more gen of hardware and local llm gets even better. Excited for this!
Mmm.. AI395 at 3.5k ( and out of stock mind ). And I still want to give it a shot when it is out. I suppose I just explained to myself the ridiculous jump in prices. The demand is crazy.
the thing about ECC memory is that it doesn't support reporting of ECC errors, so it ends up just hiding the fact that your RAM has started to go bad. this is good if you replace RAM every n years, but not if you intend to use the RAM until it starts degrading.
Ryzen AI Max Plus Pro 495
How does a CEO read this and not immediately fire their entire marketing team?
I know right. Where’s the «Elite»?
Because it was the CEOs idea?
More numbers and letters and phrases = more excite! much customer!
Agreed, what the heck were they thinking, it's obvious it should've been Ryzen AI Max Plus Pro Ultra 495.
AI means it has npu, Max+ is marking the memory channel, PRO is a normal label for chips that have extra security baked in, it has been this way since forever.
I don't see anything wrong with the name.
You literally answered what is wrong with the name.
I typed my (masked) email into the input, pressed tab and enter, it redirected me to some cloudflare captcha marketing site. I thought at first it didn't want the masked email, but then it worked when I clicked with my mouse. Apparently there are 3 invisible focus targets in between!
Dear people who create websites, these things are important, they should work!
You shouldn't need to tab first, right? Enter on an input field should submit the form unless they go out of their way to break it?
I have a framework desktop w/ 128GB that I bought last Christmas and if I’m looking at it right it costs $2000 (CAD) more now because of the RAM shortage (and in any event is apparently out of stock). Would love to have 192 GB but I’m not sure I can justify buying another.
Does anyone have a sense of how this might progress, e.g if I can get a 256 or 512 GB in a year if I wait. In any case I’m jealous this exists and I don’t have one.
One last thing, I assume this isn’t exclusive and there will be other builds with this same config same as current Strix Halo?
On RAM supplies, the best estimates I've seen are https://www.tomshardware.com/pc-components/dram/cxmt-close-t...
They are estimating a 25% shortfall in supply remaining in 2030, even with Chinese companies ramping up DDR5 supply.
It's not looking great - the RAM producers need more of the same machines that other semi-conductor manufactures need and the suppliers of those seem unable to increase production.
Are you generally happy with it? What do you mainly use it for?
I'm kinda tempted by the frameworks because they are somewhat energy efficient and compact (which I both value highly!) but my concern is that the GPU is on the low end for 1440p gaming (while being barely price-competitive with a self-built Ryzen9950 + 9070XT combination).
what I read in I think the direct AMD press, it's essentially like a 10% upgrade to compute but mostly ram.
I agree, this is the direction AMD was going in to larger attached memory and they got hit by the memory cartel pricing; They likely were going to hit 256GB instead of this weird glitch in the sizing.
So, yes, of course they're going to hit higher memory sizes; but since they models are meant for laptops and to get the speed you want for inference, they're soldered, you have no real options.
Framework in particular might have a high value on ebay as I assume this will be a drop in replacement for their existing motherboard.
I understand their reasoning but it’s still a shame Framework opted for a proprietary motherboard instead of mini-ITX with a socket CPU and GPU. I value repairability over a little bit of extra performance.
Based on the product spec sheet[1] it appears that the Framework Desktop motherboard is mini-ITX, and that the chassis is[2] as well.
It’s fully possible that I’ve missed something, but it doesn’t appear to be proprietary to me.
[1]: https://frame.work/desktop?tab=specs
[2]: https://frame.work/products/desktop-case
What do you mean?
The whole point is that these APUs present a pretty unique value proposition (lowish performance GPUs with massive amounts of RAM attached), and framework is afaik the only vendor shipping them on standard mini-ITX boards
What’s stopping you from building that yourself instead of having to go through Framework?
Framework is going for a computer OEM market where they need to make margin, hence the soldered-on components and relatively limited options. People in the pre-built computer market overwhelmingly don't care about replaceable components, because anyone else will just buy components and build a PC, for both the greater market of options AND overall lower price than buying pre-built (if they don't overspec).
You literally can't make this machine without soldered components. You don't understand this computer's design goals. It's been 1.5 years since they announced the Framework Desktop. I don't get how people are still confused about this.
Here's me wanting to drop thousands on this so I can save a few bucks a day by not having to call cloud llms for simple distillation tasks. Lol.
It’s a fantastic dev box i got the framework desktop and its awesome as a developer desktop
Tl;dr: think hard about what you’d use this much RAM for in a desktop setting and do some research about how people like the 395 for your use case. Not all use cases work well.
Anyone who thinks they are going to serve some 100+ GB LLM locally, remember that memory bandwidth becomes a key limitation for large models. While you might be able to load a model, token generation can be very slow. MoE models like Qwen3.5-122B-A10B work decently fast, but dense models of a decent size are slow and you won’t want to use them.
I’ve got a 395 system, and found that I’m quite happy with Qwen3.6-35B-A3B, generating at around 50 t/s, but the dense 27B model is 20-25 t/s and that’s the lower limit I’m willing to tolerate. So a 70b dense model is just not going to happen. That means you can’t really use that much RAM.
A reasonable use case is to have multiple smaller models loaded- you can have an image generation model loaded along with the text model. Or you can use this computer for development simultaneously with serving LLMs. Those ideas work okay. But trying to load up a single giant model is going to test your patience.
Im less interested in the memory than the memory bandwidth. The current system with 128GB can load pretty big models, but its meaningless unless you want to wait 40 minutes per prompt.
I've had best results with Qwen3.6-35B-A3B, which uses 40GB of memory, but only uses 3 billion parameters per token which helps with throughput.
Until memory bandwidth significantly improves I just can't see myself wanting to use all that memory. Unless it's just to keep a wide variety of models in memory.
I run Qwen3-Coder-Next-UD-Q4_K_XL and other than the initial wait to initialise context (which takes less than 2 minutes) subsequent prompts return in less than a minute, usually less than 30s.
If your performance is significantly slower then you are probably doing it in CPU - there was some fiddling required to get it to use GPU (I use llama.cpp)
Try poolside that came out yesterday https://news.ycombinator.com/item?id=49004937
or Qwen 3.5 122B A10B, both use more memory and still have experts sized for decent speed at the 395’s memory bandwidth at 4bit quantization
If you are waiting 40 minutes for a prompt on a Ryzen max+ you have no idea what you are doing.
I have an NVIDIA Spark thingy (the ASUS one) and it's the same problem there, though the prefill side is superior to the Ryzen ones.
But I actually think 128GB is too little. There are some compelling models that are above what can fit in that at reasonable quants (e.g. DeepSeek V4 Flash) but could if the system was 256GB.
If RAM prices weren't so f*cked I think we'd be seeing 256GB and even 512GB unified memory systems becoming quite common. As it is I think it will be 10 years before >128GB becomes feasible on a regular consumer machine for normal people again.
You can run DeepSeek Flash just fine on 128GB using lower quants. Antirez' DwarfStar supports it really well. The bigger the model, it seems, the less it is affected by high quantization. I know some people are working on quantizing GLM 5.2 so it can fit into 128GB (using additional tricks to select only some experts, etc). There's a lot going on.
I use exactly the setup you describe; it can handle 1-3 agents working if the agent work has IO delays. with MTP models, prefill is quite fast.
Not sure what damage you have, but no one waits 40 minutes per prompt or even a minute; The A3B model loads within 10 seconds and a simple response in opencode is maybe at most a minute, then catches up in 3-5second bursts depending on IO.
I put dynamic context pruning into opencode and tweaked it for 45k-85k context before it shrinks; this lets me get into 500k token sizes and fixed on medium sized github repos.
There's a chance this comment is a skill error:
1. USe llamacpp with a MTP model
2. Use reasoning-budget and reasoning-message
3. Tailor your agent to use the reasoning-message to use subagents and dyanmic compaction.
---
Now you have a reasonable coding agent for cloning, building and extending any github project I've seen so far.
Are you on Strix Halo? If so, which llama.cpp backend are you using? ROCm on Strix Halo still seems incompatible with MTP models.
Woof. What's that going to cost? And will we actually be able to buy one, because the 128GB model has been sold out pretty much since the LLM craze started...
It's competition is M5 Ultra, not cheap.
If you have to ask you don’t want to know
Memory bandwidth certainly is the bottleneck on my 395+ 128GB. It's been a great server. Think I'll wait another year maybe hopefully the prices come down and there's one more gen of hardware and local llm gets even better. Excited for this!
Mmm.. AI395 at 3.5k ( and out of stock mind ). And I still want to give it a shot when it is out. I suppose I just explained to myself the ridiculous jump in prices. The demand is crazy.
So... Can it sleep with echo mem >/sys/power/state?
Abd wake up without crashing?
Non-ECC memory.
And into the trash it goes.
the thing about ECC memory is that it doesn't support reporting of ECC errors, so it ends up just hiding the fact that your RAM has started to go bad. this is good if you replace RAM every n years, but not if you intend to use the RAM until it starts degrading.
> Seize the means of computation
Feels like they blatantly ripped that off the title of a Cory Doctorow book [0].
[0]: https://en.wikipedia.org/wiki/The_Internet_Con
That’s a common expression.
No blog post, no order, no preorder; I think the only things about it are this sentence and a half:
> 192GB coming soon
> The most powerful Framework Desktop yet is coming soon with an AMD Ryzen™ AI Max+ PRO 495 processor and 192GB of LPDDR5X memory.