Right now it seems we are once again on the cusp of another round of LLM size upgrades. It appears to me that having 24gb VRAM gets you access to a lot of really great models, but 48gb VRAM really opens the door towards the impressive 70B models and allows you to nicely run the 30B models. However, im seeing more and more 100B+ models being created that push the 48 gb VRAM specs down into lower quants if they are able to run the model at all.

this is in my opinion is big, because 48gb is currently the magic number for in my opinion consumer level cards, 2x 3090’s or 2x 4090s. adding an extra 24gb to a build via consumer GPUs turns into a monumental task due to either space in the tower or capabilities of the hardware AND it would put you at 72gb VRAM putting you at the very edge of the recommended VRAM for the 120GB 4KM models.

I genuinely don’t know what i am talking about and i am just rambling, because i am trying to wrap my head around HOW to upgrade my vram to load the larger models without buying a massively overpriced workstation card. should i stuff 4 3090’s into a large tower? settle up 3 4090’s in a rig?

how can the average hobbyist make the jump from 48gb to 72gb+?

is taking the wait and see approach towards nvidia dropping new scalper priced high VRAM cards feasible? Hope and pray for some kind of technical magic that drops the required VRAM while simultaneously keeping quality?

the reason i am stressing about this and asking for advice is because the quality difference between smaller models and 70B models is astronomical. and the difference between the 70B models and the 100+B models is a HUGE jump too. from my testing it seems that the 100B+ models really turn the “humanization” of the LLM up to the next level, leaving the 70B models to sound like…well… AI.

I am very curious to see where this gets to by the end of 2024, but for sure… i won’t be seeing it on a 48gb VRAM set up.

  • nostriluuB
    link
    fedilink
    English
    arrow-up
    1
    ·
    11 months ago

    I have two questions:

    what’s this going to look like in six months, with new Intel, AMD, ARM/RISC UMA, hybrid designs well supported and 7200mt+ DDR5 common?

    Are the high memory models that much better? My impression is you get a lot of reliable utility out of good smaller models, from there it’s diminishing returns.

    I had a honking system with two 3090s, but it felt a bit boondoggle-ish, I sold it and my current plan is to get something like a 4060ti-16gb and also use OpenAI’s API, so I can wait to see what develops, rather than spending it all now while it’s still early days. I can see how someone who is really developing LLMs would want more, but as a “consumer” this seems reasonable.

    Even for the “just get a Mac studio,” it seems like the M3 can use more VRAM and is more optimized, so worth it to wait until the M3 Ultra comes out, unless you can get a bargain bin previous model.