Your benchmark is a curious one. I didn't see you included Muse Spark 1.3 contributor even though its price is much lower even than DeepSeek. The low price changes many recommendations completely. And FWIW, DeepSeek retain and train on your data, too.