Anyone getting more than 13 tokens per second on M1 16g machine

Usign GPT4all, only get 13 tokens. anyway to speed this up? perhaps a custom config of llama.cpp. or some other LLM back end.

model is mistra-orca.

does type of model affect tokens per second?

what is your setup for quants and model type

how do i get fastest tokens for second on m1 16gig

You must log in or # to comment.

Chat