Fun_Tangerine_1086B to

LocalLLaMA@poweruser.forumEnglish · 3 years ago

Why is Mistral-7b so capable? Any ideas re: dataset?

1

Why is Mistral-7b so capable? Any ideas re: dataset?

Fun_Tangerine_1086B to

LocalLLaMA@poweruser.forumEnglish · 3 years ago

So Mistral-7b is a pretty impressive 7B param model … but why is it so capable? Do we have any insights into its dataset? Was it trained very far beyond the scaling limit? Any attempts at open reproductions or merges to scale up # of params?

Chat

cleverestxB
link
fedilink
English
arrow-up
1·
3 years ago
Why can we get a 20 - 34b version of this very capable Mistral?