You could find these under $100 yesterday on AliExpress with some referral + coupon stacking (rakuten + coupon + paypal pay-in-4 discount). The cheapest sellers often only put out a limited number of items each day, so you have to check back every few hours.
I have one of these I haven't gotten around to futzing with, but have two friends who bought two (each) and have been playing with them for local inference.
Notably we did some research and my one friend has had some luck using the M.2 slot with a PCIe adapter to a ConnectX 3 card and then using that to do multinode coordination (using RDMA and a llama.cpp patch) and has been getting almost 20 tok/s decode out of Qwen3.8 27b dividing it up between two BC250s and a coordinating PC. Prefill is weak though.
A lot of futzing around so not worth it if you don't enjoy that kind of thing, but remarkably cheaper than buying GPUs or a DGX Spark. (I have a Spark so I really shouldn't be bothering... but....)
I've yet to really run mine since I dont have a great solution for the heat -- I hear it pulls 60w straight so it would be mega hot I presume.
Most just 3D print (or buy) a manifold for a standard PC 12v fan.
I assume duct-tape and cardboard would also work if things are dire. =3
Is your heat sink entirely unmodified? Opening up the fins should make it a lot easier to cool.
Hard to find these for under $300 these days.
You could find these under $100 yesterday on AliExpress with some referral + coupon stacking (rakuten + coupon + paypal pay-in-4 discount). The cheapest sellers often only put out a limited number of items each day, so you have to check back every few hours.
I have one of these I haven't gotten around to futzing with, but have two friends who bought two (each) and have been playing with them for local inference.
Notably we did some research and my one friend has had some luck using the M.2 slot with a PCIe adapter to a ConnectX 3 card and then using that to do multinode coordination (using RDMA and a llama.cpp patch) and has been getting almost 20 tok/s decode out of Qwen3.8 27b dividing it up between two BC250s and a coordinating PC. Prefill is weak though.
A lot of futzing around so not worth it if you don't enjoy that kind of thing, but remarkably cheaper than buying GPUs or a DGX Spark. (I have a Spark so I really shouldn't be bothering... but....)
Fun.