サクサク読めて、アプリ限定の機能も多数!
トップへ戻る
買ってよかったもの
github.com/JustVugg
A 744B Mixture-of-Experts model activates only ~40B parameters per token — and only ~11 GB of those change from token to token (the routed experts). So: the dense part (attention, shared experts, embeddings — ~17B params) stays resident in RAM at int4 (~9.9 GB); the 19,456 routed experts (75 MoE layers × 256 experts + the MTP head, ~19 MB each at int4) live on disk (~370 GB) and are streamed on de
このページを最初にブックマークしてみませんか?
『github.com』の新着エントリーを見る
j次のブックマーク
k前のブックマーク
lあとで読む
eコメント一覧を開く
oページを開く