> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.
> Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.
> The model will be released with open weights on October 15.
I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5).
I wonder if other Chinese labs like Kimi/Moonshot will follow suit.
> Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.
Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.
IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.
Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.
I regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.
I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).
I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.
Another possible reason is that the number 4 is considered unlucky in traditional Chinese culture.
Moonshot has already teased K3.1 so not likely
K3.1 would likely be a deeper/longer post-train from K3, so that’d make sense.
It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).
> Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.
Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.
IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.
Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.
For anyone else looking for the pricing: https://platform.stepfun.ai/docs/en/guides/pricing/details#p...
I regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.
> At $0.04/M
Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview.
> $8-12 per day only for cached reads
Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model?
Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview
Yes, I meant the Flash version.
I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).
I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.
How about adding a contested historical facts benchmark?
Huh wonder why they skipped 4?
Sometimes 4 is skipped due to being considered unlucky.
In China and in places influenced by Chinese culture, due to homonymy between "4" and death.
Sounds like death in Chinese.