Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> without losing performance

But isn't M2 ULTRA over 20x slower than this thing? ~30 TFlops vs 738.



By losing performance I meant you don’t need to quantize the model a lot since it fits in RAM. My bad for not clarifying it.


According to this [1] article (current top of hn) memory bandwidth is typically the limiting factor, so as long as your batch size isn't huge you probably aren't losing too much in performance.

https://finbarr.ca/how-is-llama-cpp-possible/




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: