Yup. I posted about the huggingface countdown: https://news.ycombinator.com/item?id=49272534 but seems like both the huggingface page, and the modelscope.cn URL that commenters found, were quickly taken down. Not sure if the submission being flagged is also the doing of whoever took down the huggingface/modelscope.cn pages.
Edit: for anyone curious, I saw the huggingface countdown page for Kimi K3, which gave me a hunch that led me to the countdown page for Qwen3.8-27B. As of ~11am EST today, the countdown page said 2 days and 2 hours to release.
I wonder besides big labs and REALLY big corporations who can run the 2.6TB q8...
I mean, the 1 bit one is 508GB.
Assuming 256k context size as irrelevant at those sizes, I would need 22 AMD 7900XTX (24GB vram) to run this for the 1 bit one, and 113 for the 2.6TB
113 and assumming close to constant load and the gpus taking turns as it does on my machine with two gpus is around 113gpus*113W = 11KW... just to hold and run that one instance... and only god knows the cooling requirements...
I guess just someone on Meta/Googly/Claudy will say... oh nice, let's download it and run it...
How did the unsloth guy did to process a file that size?
> Meta/Googly/Claudy will say... oh nice, let's download it and run it...
You can run it for like <$50/hr. You don't need to be Meta/Googly/Claudy for that. For batch inference it makes sense to run it yourself(high latency, high throughput, predefined workload for few hours or day).
We're all really waiting for the 3.8-27B which is due out in 2 days.
Yup. I posted about the huggingface countdown: https://news.ycombinator.com/item?id=49272534 but seems like both the huggingface page, and the modelscope.cn URL that commenters found, were quickly taken down. Not sure if the submission being flagged is also the doing of whoever took down the huggingface/modelscope.cn pages.
Edit: for anyone curious, I saw the huggingface countdown page for Kimi K3, which gave me a hunch that led me to the countdown page for Qwen3.8-27B. As of ~11am EST today, the countdown page said 2 days and 2 hours to release.
Submitted 3 hours ago with 63 comments so far: https://news.ycombinator.com/item?id=49273478
I wonder besides big labs and REALLY big corporations who can run the 2.6TB q8... I mean, the 1 bit one is 508GB. Assuming 256k context size as irrelevant at those sizes, I would need 22 AMD 7900XTX (24GB vram) to run this for the 1 bit one, and 113 for the 2.6TB
113 and assumming close to constant load and the gpus taking turns as it does on my machine with two gpus is around 113gpus*113W = 11KW... just to hold and run that one instance... and only god knows the cooling requirements...
I guess just someone on Meta/Googly/Claudy will say... oh nice, let's download it and run it...
How did the unsloth guy did to process a file that size?
https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF
Newegg lists h200s ~40k
another 50k for the machine to put your 8 h200s in
Its pricey obviously but the 500-1M price tag is not out of the realm for most large companies.
But nobody that can run this is running it on that hardware, surely?
> Meta/Googly/Claudy will say... oh nice, let's download it and run it...
You can run it for like <$50/hr. You don't need to be Meta/Googly/Claudy for that. For batch inference it makes sense to run it yourself(high latency, high throughput, predefined workload for few hours or day).
I have been testing this for hark.news
this is Qwen3.8 max, right? https://qwen.ai/blog?id=qwen3.8
That's my bet, I doubt they'd have the resources to have trained anything substantially larger
they explicitly say this is qwen3.8 max in the hugging face page