Marin 535B - A23B
Özgün başlık: Marin 535B-A23B
A couple weeks ago, we kicked off our capstone run of the year, a mixture of experts (MoE) model with 535B total parameters and 23B active. The plan is for it to pretrain for about 3 months on 18 trillion tokens, including agentic, coding, and scientific data. Consistent with our commitment to open development , the entire run—including architecture, data, predicted loss, etc.—is documented on GitHub .
It's the largest model we've ever trained, and so far, it's going almost boringly well, lining up with our predictions and chugging along with a minimum of infrastructural or numerical fuss. You can follow along in our W B report , on the GitHub run tracker or on our purpose-built dashboard .
It also happens to be just over a year since Marin joined Open Athena, and so it's natural to reflect a bit. In short, it's been a big year. Since joining Open Athena, Marin has grown from one FTE to ten. And in that time, we've built a lot of the machinery we used to wish existed.
We also published some papers and helped incubate two separate efforts building foundation models for biology .
Any one of those accomplishments would have seemed wildly ambitious for Marin a year or two ago. Unsurprisingly, I keep coming back to the 535B run. What I find most remarkable is just how boring it has felt since launch.
You see, our hero runs were traditionally not boring. When Percy and I started Marin back in 2024, our goals, though objectively more modest, still seemed audacious.
The plan was to match Llama 3.1 8B with a fully open source model. Google's TPU Research Cloud generously gave us access to 512 TPU v5e cores, a dramatic jump in our compute resources at the time. Marin 8B succeeded : the model matched or exceeded Llama 3.1 8B on 16 of 19 base model evaluations.
We followed that with Marin 32B, opportunistically launched on a pile of preemptible v5p chips between reservations. We yolo'd the hypers and kicked off the run. No big deal.
But, we soon discovered that, among other things, one FTE is actually not enough to babysit a large training run on cobbled-together infrastructure on top of preemptible compute, particularly once it starts throwing loss spikes . Who knew? After much surgery (notably QK norm) and an eventual retreat to a stable v4-2048, the model somehow limped across the finish line.