LocalLLaMA@poweruser.forumEnglish · 3 years ago

Quantizing 70b models to 4-bit, how much does performance degrade?

1

Quantizing 70b models to 4-bit, how much does performance degrade?

LocalLLaMA@poweruser.forumEnglish · 3 years ago

The title, pretty much.

I’m wondering whether a 70b model quantized to 4bit would perform better than a 7b/13b/34b model at fp16. Would be great to get some insights from the community.

Chat

a_beautiful_rhindB
link
fedilink
English
arrow-up
1·
3 years ago
70b 4bit will eat those small models for breakfast.