Beam: Reflection's 501B open-weight model

214 points59 comments3 hours ago
Ariarule

Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

show comments
htrp

> Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

Early access, no weights no tech details, just a sign up here for info

show comments
NorwegianDude

Bigger and still worse than existing free Chinese models that are smaller? Open weight models are nice, but at this point it seems western models are very far behind Chinese ones, despite Chinese companies publishing a lot of their findings. I hope we get more open models and more providers, as being stuck with a model from China or US with no competition is risky.

Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.

show comments
onlyrealcuzzo

This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

Am I missing something?

show comments
drubs

I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!

show comments
michaelkdev

Is it worse than the top open-weight Chinese models? Yes, it is, but at least the West has joined the party, and hopefully they will iterate on this and keep up the pace. The Chinese labs will certainly release new and powerful versions soon, so it's all about relative pace right now.

eaf7e281

> Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.

It's great to see a company that acknowledges it still needs improvement instead of making false claims.

aeetes

the performance chart puts the better open source models behind the fold making it seem like it outperforms them... but it doesn't! all for open source models but this announcement is misleading

show comments
segmondy

Any time a new lab shows up, folks complain about how their models are worse. Really? It would be nice if a new comer comes from no where and beats everyone, but that's rarely the case. The good thing is that other labs/people are figuring out how to build this, and if they keep at it then this is as bad as it gets for them and it would hopefully get better. A new entrant to the market is good for everyone.

TheArcane

If you don't buy into "America good, China bad" narrative, this new entrant & release by Inclusion Ai is a lot more exciting by every measurable metric.

https://github.com/inclusionAI/Ling

hypfer

Someone should name their next model "Workhorse" just for SEO reasons.

It's interesting how the industry converged to this very term, given that very less work is being done by horses since quite a while.

show comments
zopper

Open model that is not yet open or widely accessible via API. Primarily comparing to non-SOTA models like Inkling and GLM 5.2. Included comparison to GLM 5.3 and DeepSeek V4.1 Flash in the table, but not in the charts (I assume they would make them look bad). Also no results from AA Index or Arena.

keeganpoppen

very curious to see more about what kinds of hardware you can run this on and the perf. characteristics… on the face of it, it seems like optimizing for inference speed might(?) be good for running on smaller hardware, but i suppose it could be the other way around and it is actually much resource-hungrier for the number of parameters, etc. …

[deleted]
vcryan

This is like an ad for how great Deepseek V4.1 Flash is.

wg0

Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.

Where do I get the data?

I mean, this many models. They have to start somewhere.

show comments
sharktheone

Am I the only one who thought of the BEAM VM after the first word of thee title?

show comments