Usign GPT4all, only get 13 tokens. anyway to speed this up? perhaps a custom config of llama.cpp. or some other LLM back end.
model is mistra-orca.
does type of model affect tokens per second?
what is your setup for quants and model type
how do i get fastest tokens for second on m1 16gig
You must log in or register to comment.