Gensyn #︱📢︱announcements
Introducing open-1b - The first language model with auditable, verifiable training.
open-1b represents a milestone on the path toward verifiable AI, a goal that is absolutely necessary for the future of intelligence.
The most used models are closed and concentrated amongst a few companies. How they were built, what went in and what did not is hidden and unknowable. Their biases are unknown and so unable to be trusted. The owners of those models suggest that they are the only that are responsible enough to be trusted with them.
Instead of having to trust how a model was trained, open-1b comes with its complete pretraining dataset, training and evaluation code, intermediate checkpoints at 100-step intervals, and a canonical state hash for every one of the 80,957 optimizer steps that produced it.
Anyone can load a checkpoint, replay the associated step on their own hardware, hash the result and confirm it with the published fingerprint.
AI verification is fast becoming a critical requirement. Models you can trust are the only defence against models that you cannot. Models that have their entire history on the public record and are able to be replayed.
Determinism means the same machine gives the same answer twice. Reproducibility means a different machine gives the same bits. Existing determinism settings make runs repeatable on the same hardware. That is not verifiable, as it is not reproducible by anyone else.
Our verifiable AI infrastructure allowed that gap to be closed. RepOps - https://www.gensyn.ai/research/verde-a-verification-system-for-machine-learning-over-untrusted-nodes, our library of reproducible operations, and REE - https://www.gensyn.ai/news/ree, our reproducible execution environment, make matrix multiply, normalization and gradient reduction produce identical bits whether it runs on a consumer NVIDIA card, an x86 or ARM CPU, or MacBook.
Auditing the training of open-1b is a collective exercise, and you can take part. Verifying all 80 957 steps alone isn’t practical, but with enough people reviewing enough of it, the whole is verified.
- Download the audit harness
- Pick any step of the run
- Replay it
It runs on NVIDIA GPUs, x86 and ARM CPUs, and natively on Apple Silicon. When your result matches the published hash, your verification is recorded and credited on-chain in the public ledger - https://open1b.gensyn.ai. This ledger assembles the individual checks into a single collective statement: this model was trained exactly as declared.
There is no reward or yield from auditing. The reward is your name on the immutable record of the first ever fully audited training run.
In a future that has achieved democratisation of intelligence, that is an important place in history.
7/
Read more here:
Blog Post: https://www.gensyn.ai/news/introducing-open-1b-auditable-training
Paper: https://open1b.gensyn.ai/open1b-tech-report.pdf
Github: https://github.com/gensyn-ai/open-transformers
Huggingface: https://huggingface.co/collections/Gensyn/open-1b
Auditing open-1b: https://open1b.gensyn.ai
Audit tool: https://github.com/gensyn-ai/pretraining-audit-cli
@everyone
Gensyn | Introducing open-1b: the first model you don’t have to t...
Auditable training is the best defense against the future of AI we’re being warned about. Gensyn has proved it’s possible.
Post #697
2