48 points | by mike-the-brain 2 hours ago
5 comments
Amazing tech
> An agent writes in an afternoon what a chatbot writes in a month
But can you just.. not.
Your tech is so good, it speaks for itself. Don't ruin that.
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
Great news, has made low memory bandwidth model usage so much nicer.
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816
llama.cpp PR https://github.com/ggml-org/llama.cpp/pull/27342
Amazing tech
> An agent writes in an afternoon what a chatbot writes in a month
But can you just.. not.
Your tech is so good, it speaks for itself. Don't ruin that.
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
Great news, has made low memory bandwidth model usage so much nicer.
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816
llama.cpp PR https://github.com/ggml-org/llama.cpp/pull/27342