Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I would expect that 4060ti to get about 20-25 tokens per second on Mixtral. I can read at roughly 10-15 tokens per second so above that is where I see diminishing returns for a chatbot. Generating whole blog articles might have you sit waiting for a minute or so though.


Thanks, that sounds more than tolerable than "more than an hour"!

I also have the 16GB version, which I assume would be a little bit better.


It depends on the context window, but my 3090 gets ~60/s on smaller windows.


I get 50-60t/s on Mistral 7B on 2080 Ti




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: