Vinci MLE 123B 1.0 is available now on HuggingFace.
Today we're adding a dense 123B to the MLE family too.
First - I know 123B can't be run on consumer hardware.
The weights are around 246GB.
So this is a very big model.
And as an 'own your AI' company,
obviously that's not great.
So why are we making one?
First, distillation.
A strong 123B gives us a much better model to teach the smaller models with.
We can distill from the 123B into the 30B.
And eventually into the 8B too.
The point isn't that everyone should run 123B.
The point is to make the smaller models better.
Second, I think we need to start building bigger models ourselves.
Right now most of the really strong models are basically either:
US + closed source or
China + open weight.
There are great models from both.
But I don't think those should be the only two options.
We're building in Canada,
So I want us to eventually have our own stack too.
To be clear, this isn't our own base model yet.
We're still starting from an existing parent and training on top of it.
There's a huge amount of work between this and pretraining our own 123B from scratch.
It's planned though, and pre-training will happen.
But we have to start somewhere.
For the big models, when they are good, we will find partners to offer them via APIs.
For the small ones,
I want them to keep getting good enough that you can actually own and run them yourself.
Long way to go.
Thanks for following along while we figure it out.