Your benchmark is a curious one. I didn't see you included Muse Spark 1.3 contributor even though its price is much lower even than DeepSeek. The low price changes many recommendations completely. And FWIW, DeepSeek retain and train on your data, too.
Yeah, it's hard to find a single simple task that all models fail on, in low context length conditions.
Also because models now are actually not that good on knowing things (domain knowledge), as they rely more on web search on tool use. So if I added a question, about some obscure fact, probably the SOTA models would fail it, but in practice they would find it with web search enabled. Not sure how to handle that. This is also why Gemini is on top, it's good enough at coding and instructions following, while having by far best general and domain specific knowledge.
https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...