> Not if an LLM over chat can fool most people they're talking to a human (which it can)
I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing. I really can't understand how this gap persists; but then, there seem to have been at least some people who couldn't sniff out ELIZA, back in the day, too.
>I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing.
That's mostly true for longer LLM output with all the sycophancy / LinkedIn bias thrown in.
Make it casual conversation or comments, and give it instructions on appearing casual, or even better kill the censoring and fixed-prompt (with an open model), and it's orders of magnitude more difficult, unless if you suspect it and try specifically tailored prompts to sniff it.
There's no shortage of people obliviously discussing with AI bots in comment sections.
Ok, you and I can easily spot LLM text. So what? The Turing test has still been passed, as is clear by people falling in love with ChatGPT, not believing something is AI, and by continuously claiming this or that is a bot.
People, many of them at least, cannot make this distinction anymore. You can, I can, but people as a whole are having problems with that.
By this standard, the Turing test was also passed by ELIZA, but nobody serious actually gave it that credit. Aside from which, the understanding of that concept in popular media (both the nature of the test itself, and its supposed significance) has drifted way away from what Turing was saying.
Even this (assuming it's even true) will likely not be true in some near-term future.
>continuously claiming this or that is a bot
I see it as a contemporary form of religious thinking. Like (say) pilgrims seeing blood on a statue of the virgin, plenty of people are now seeing the hand of AI in everything they read. If you want to see something hard enough, it tends to become magically visible.
Many, many people became friends/got romantically entangled with GPT-4o, to the point where OpenAI struggled to replace it due to user backlash.
Most users aren't very critical of the output. They just want a sycophantic ear, and 4o was perfect for that task. It's not _good_ but there is high demand for it.
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing.
That's because they're trained that way. If you trained a modern frontier LLM with the explicit goal of passing the Turing test, it would have no difficulty doing so.
Turn on showdead and look at the killed comments on this thread. (Every thread remotely related to LLMs seems to attract this behaviour.) If it's so easy to get it right, how is there such a high fraction of failures?
I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing. I really can't understand how this gap persists; but then, there seem to have been at least some people who couldn't sniff out ELIZA, back in the day, too.