Try asking it a follow up question not commonly found in the training data: such as "do you think bulgogi would grow on Mars?", and see what kind of gibberish you will get in response. Moreover, the model has no way of self-diagnosing whenever it produces gibberish.
I'd buy that if Meena could infer and reason about her own answers.
Human: Do you like Korean food in general?
Meena: It's okay. I like beef bulgogi, but I'm not a huge fan of kimchi.
Human: Ok what should I shop for ?
Meena : You've got almost everything but you need a pear, the steak and some ginger.
The problem with language models as commonsense is that they are collections of patterns and associations, and that they don't have inference models or solvers - unlike my dog for example!
What do you mean by this? Of course they do, learned from their training data. For example, here is quote from conversation 38 of https://github.com/google-research/google-research/blob/mast...
Human: Do you like Korean food in general? Meena: It's okay. I like beef bulgogi, but I'm not a huge fan of kimchi.
It seems to me Meena "knows" bulgogi and kimchi are Korean foods. Isn't that common sense? If it isn't, what do you mean by "common sense"?