As a beginner, I appreciate that there are metrics for all these LLMs out there so I don’t waste time downloading and trying failures. However, I noticed that the Leaderboard doesn’t exactly reflect reality for me. YES, I DO UNDERSTAND THAT IT DEPENDS ON MY NEEDS.

I mean really basic stuff of how the LLM acts as a coherent agent, can follow instructions and grasp context in any given situation. Which is often lacking in LLMs I am trying so far, like the boards leader for 30B models 01-ai/Yi-34B for example. I guess there is something similar going on like it used to with GPU benchmarks: dirty tricks and over-optimization for the tests.

I am interested in how more experienced people here evaluate an LLM’s fitness. Do you have a battery of questions and instructions you try out first?

  • BlueMetaMindOPB
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 year ago

    Or if you are just playing around, you just write/search for a post on reddit (or various LLM related discords) asking for best model for your task :D

    I made this post as an attempt to collect best practices and ideas.

    use GPT4 to evaluate output of llama.

    That’s always a good option probably but I try to avoid using openAI all together.