Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.
Also includes outtakes on the ‘reasoning’ models.
Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.
Also includes outtakes on the ‘reasoning’ models.
See, that’s not even an accurate criticism because part of language is meaning. This test is a test of an LLM having enough “intelligence” to understand that you can’t wash your car without your car being at the car wash. If you see the language presented in this test and don’t immediately realize that it would be a problem, then you haven’t understood the language. These are large language models failing at comprehending any language. Because there’s no intelligence there. Because they’re just random word guessers.
randomprobabalistic word guessers.