There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
Something is deeply wrong with AI generated output, and I say this as someone who is typically very impressed by AI.
The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
> ... including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
But that is precisely what human mathematicians do, prove new theorems by combining ones proven earlier.
I don't see any fundamental difference in functionality between human intellectual contributions vs performant ML ones (LLM or otherwise).
Whenever we listen or read text we are also predicting the near future content.
Just like LLM's we sometimes correctly predict the next token or word, and sometimes incorrectly.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model [...]
Imagine someone could pause the universe with a remote control, scroll back in time a little, press play again, and ask a slightly different question, etc.
In such a thought experiment one could also collect the probabilities for a specific human predicting a next word. Implicitly the brain also has a corresponding statistical model, regardless of the construction being visible or hidden. I.e. human intelligence is also fundamentally a statistical model, so the only thing that remains from your claim is that machines for some unmentioned reason don't possess any "real" intelligence or critical thought...
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans. Which says a lot about humans.
Idk something about a mirror maybe and the output reflecting the input?
> I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans.
So does the Google search bar, but I don't ascribe intelligence to it.
I mean I'm quite proud of some of my search queries in the same way I'm quite proud of some of the LLM output I get. I'm probably just very arrogant and enjoying myself via some LLM indirection.
Am I the only one that sometimes reads back particularly good emails they've written? I feel like its a similar thing :).
Some of it the effect of tells. “It’s not X, it’s Y” is not a bad pattern but it was baked into the instruction following training set just like the other patterns. I catch myself about to use it and use something else because I want to look human. I have, a few times, tried to use AI to write something that I was struggling to find the words and I just didn’t like how it didn’t seem like my voice. If there was just one person doing it would be OK but when it is 100s of blog posts submitted to HN a day it is like wearing a “I’m an NPC” t-shirt.
Someone shared with me this system prompt that at least makes assistant outputs usable
For information retrieval tasks, I want you to provide links to sources and use exact quotes as much as possible. When using a source, consider if it is primary or secondary information. If secondary sources are found, search again for primary sources. Sources and quotes, if applicable, should be mentioned in the answer first before the rest of the response with links.
I just can't accept that it possesses no intelligence. It is not equivalent to human intelligence, obviously, but how can a system without some semblance of rational thinking solve open math problems? Even composing earlier human work into something novel requires intelligence and understanding on some level.
We couldn't agree on what intelligence means before ChatGPT happened. Now, agreement on the term seems even further away
If performing well on an IQ test or performing at a high level on knowledge work is intelligence to you, these models are intelligent. If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
I don't think "intelligence" needs to carry all the intrigue and woo of related words like "consciousness" or "creative." If we just use "intelligence" to mean "the ability of a system to solve problems that are new to the system," that pretty much matches the dictionary definition and normal usage of the term. We don't need to touch messy questions like "is there something it's like to be a bat" to conclude that bats exhibit intelligence when they navigate long distances and hunt for food.
I'm not exactly that you mean by "new to the system", but it seems to me that that definition makes a calculator intelligent, which I can't agree with.
Intelligence isn't a binary property. Is it really a problem to say that a calculator has some intelligence? That it's more intelligent than e.g. a rock?
I agree, but it's clear most people need a definition of intelligence that (1) they qualify for and (2) nothing/no one they don't like qualifies for. And they'll keep redefining intelligence until they satisfy both criteria.
It has no semantic depth. The sentences and the paragraphs are a statistically viable derivation of existing human text, but once you try to grasp the whole thing with its temporal and spatial dimensions, you are left with a blurry mess that rots your brain. It's a polished, inoffensive and shallow interpretation as written by an opinionated reputation-seeking user of Quora, circa 2019. Assertive, bold, without typos, clean-cut and bulleted, but without an interesting semantic core.
Yeah, I hated all those Quora users that would just spew out semantically meaningless slop like increasing an important bound for the Riemann hypothesis.
I'm guessing whether you believe it possesses intelligence or not depends on your answer to Searle's Chinese room thought experiment[0]. I'd also recommend checking out the Peter Watts' book, Blindsight.
The Chinese room is a good Rorschach test for this kind of thing (but not a good thought experiment, IMO, because it's obviously correct or obviously wrong depending on where you're already coming from), but also it's not really about intelligence per se, but more abstractly awareness and more adjacent to consciousness than intelligence, and these are not the same thing.
It's just filled to the brim with relations between things. It's good at searching a very large meaning space and create correlations. What it does is to cover great distances and find related things in that large space which needs a long time and large corpus of knowledge to find the connection.
This is not intelligence. It's just a good correlation engine with a very big albeit lossy database of things.
Intelligence is compression, compression requires subtraction, and for some reason LLMs are not good at subtracting. To create a coherent model you kinda have to subtract correlations until only the essential parts are still there.
What I don't understand is why LLMs haven't been able to do this yet, if it's the harness or some orchestration layer above the LLM that is needed. Because fundamentally if you can identify correlations then it's just another small step to prioritize and remove lower value or irrelevant correlations.
I wonder if what's needed is to introduce subtraction tokens in some sense, and in post-training reward the model on that.
Intelligence is compression? What do you mean? Intuitively that doesn't seem right.
>What I don't understand is why LLMs haven't been able to do this yet
LLMs are just trained on what humans have said. Why is it surprising that it's still not possible to reconstruct the intelligence that wrote all that by working backwards? Think of your own work experience. When you look at a piece of code, say, are you always able to discern why the person did what they did, just from the code, with no additional context?
I guess they mean that intelligence is being able to hold models (compressed versions of reality) internally and use them to make predictions with a probability better than chance. That last part is the definition of information.
I find that highly questionable as a general description of what intelligence does. That's more like a description of a general knowledge base. When I think of someone intelligent, I think of someone who's able to draw unexpected connections between seemingly unrelated facts. In the broadest possible terms, I'd call it the ability to make abstractions and analogies. This is not just compression, but the ability to mentally operate on webs of meaning.
Doing those things also contributes to compression. I do recommend reading up on it, it's perhaps a little overstated for what people intuitively consider the two concepts but it's been quite well explored and has held up pretty well in practice.
bzip is not very intelligent, true, but it does develop some model of its input. It's not like there's a linear relationship between between compression ratio and IQ or anything.
The very fact that it is able to search within a meaning-space demonstrates that it understands semantics, to some extent. Philosophically, that is profound, for something that is just one big matrix multiplication. Drawing connections between things in meaning-space is surely a facet of intelligence.
It’s not intelligence if you are the one who gives the correlations to the model in the pre-training. It’s Word2Vec, applied. Model doesn’t learn anything. You embed these correlations and build it from there. It just searches the space.
As my AI professor said in the first lecture: “All AI is advanced search”.
Okay, I guess you're right that its ability to do this is just correlational, which doesn't imply it has any understanding. However, you have to conclude that some tasks which we used to believe required intelligence don't actually require any, which is disconcerting.
No, what I would say is the tasks which are handled in a passable manner by LLMs can be mathematically modeled with some reasonable accuracy.
Many things are predicted by models in our planet. From weather to production and material science. Building the model needs intelligence, running the model does not.
The person who came up with the formulae for CFD was intelligent. The computer running the model is not. Same for LLMs, chess engines, engine ECUs and financial prediction systems.
Again, for the example’s sake; the person who came up with an algorithm is intelligent. The model mixing its training data to emit something similar is not.
This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
> This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
No.
> So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
With a big "if" attached to it. People were saying "computers will program themselves in the near future" for, checks notes, 24 years now, as far as I'm aware.
We're constantly building new knowledge and understanding things better than olden days. These models just compress our knowledge and light the blind corners we can't see well. I don't say they are useless, but I say that these things are overhyped.
All they can do is regurgitate human knowledge packed into them and highlight some long-distance correlations between items, which is useful in itself, but it can't jump to somewhere where it's not present its training data, but that's something humans and only humans can do.
I get what you're saying. The thing itself is just math. I'll just say it depends on how you define intelligence. If at some point we're be able to simulate a human brain with 100% accuracy, I would say that it is intelligent, it sounds like you would not. (I don't mean to imply consciousness or personhood or anything else by "intelligent".)
For me intelligence is a fairly clean-cut concept, and is somewhat inseparable from consciousness itself.
Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
Considering the model even doesn't see the words and work on matrices of numbers is even more telling. One needs to add "tools" and other "experts" to overcome the shortcomings caused by this modus operandi.
I can call the algorithm/model smart as in a smartwatch. It can mimic certain things well while having none of the underlying foundation beneath it, or redirect some of the things to correct tools to get deterministic and accurate results if it can't evaluate the query inside its own network in a sane manner.
Coming to your question, "simulating a brain" in a static manner would not make that simulation intelligent, but if you can "wire" it completely and let it evolve by itself, now we're entering a territory I have not spent enough time for thinking it through.
Oh, as I said "I don't know", an LLM doesn't know what it doesn't know, and can't self correct itself which are required capabilities for understanding something. It just generates something statistically viable via its network.
I suspect like most you don't appreciate how terrifying statistical relationships become when you have truly vast data sets to train on... and also that we as humans aren't as shockingly unique as we think (compared to other humans I mean).
Its a mirror to human intelligence. Regurgitating phrasing to match what someone who can reason put together, but it isn't any more intelligent than the reflection of you in the mirror is.
I wonder if you went back before we had any idea how the brain worked and talked to the smartest people about how neurons work (without giving away that it's a human brain) then asked them all "would such a system be intelligent?" how many would say yes.
The main problem I have with people stating it's not intelligent or conscious is I don't think we even have a good definition of either word that satisfies everyone. Philosophers have been trying (and failing) to elegantly define these things forever and everyone out here proclaiming they've got the definitive answer and this specific thing they're seeing doesn't fit under it.
This looks interesting, but would you mind saying a sentence or two about why before I commit to an hour-long video? It looks like it shows how they work internally, which is sort of a non sequitur. Brains also work mechanistically. I'm claiming that any system which is able to do what AIs do must necessarily have some sort of intelligence.
fair reply to an hour video, Scott is just so good to hear his talk is better than I can explain it...
go to 24 minutes and 07 seconds.
it's statistically determining what the next word should be based on all the text it's been trained on. It's not intelligence and he shows what probability it puts on each word that it chooses, but also shows a lot of the other words it was thinking of using. In a later part he shows how it uses words that are not the highest probability (and you question why did it go this route, it's not more correct), but the user never sees this, they see what they think is the correct answer always...
he also shows how context you feed it has a lot to do with what it returns... to the point he can get it to return the capital of France is Marseille, just by typing Marseille a bunch of times before the question. Human intelligence doesn't get confused like that.
And it's not a "hallucination", it's just probability of the next token prediction based on the information it's been trained on and fed, it's not intelligence.
Isn't this a case of missing the trees for the forest though? The human brain is not an LLM, and an LLM is not intelligent in the same way as a human brain.
However, an LLM is a prediction machine, prediction IS at the very least one (or the most fundamental) element of intelligence. The brain most surely contains at least some kind of simulacrum of a prediction machine. How that prediction machine is used or wrapped is another matter.
If I said to you: "Blue blue blue, the color of my car is red", would you have absolute confidence in your prediction that my car is red? Or would the way I phrased that sentence make you slightly uncertain, and wonder if there's some miscommunication going on here?
LLMs are pattern prediction systems with a large training data set. It is not surprising that they can predict patterns, particularly for a well structured field like mathematics that is also amenable to automated proof checking to help steer it.
AI is just a good permutation/combination engine that tries to act smart with help of statistics. At best I only see AI as, 1. An autocomplete on steroid, 2. Good search/correlation engine
Are we sure there is some objective, technical definition of what is intelligence and what is not?
Isn't it rather a subjective philosophical concept? What if human intelligence is also a statistical model, trained by evolution to make decisions that lead to offspring?
The one major difference I see between AI and people is the ability to learn and memorize. All memory/learning solutions that current AI architectures offer just feel like workarounds and simply don't work anywhere near as a person learning something new and remembering it.
I mean this in the kindest way possible, but you are wrong that the math solutions are that easily dismissed. And there are many more than are publicized. A specific math problem I wanted solved for 3 years did not get solved by any model until fable and, and I tried it on every model and know the literature surrounding it well.
When I read AI-generated prose that is aimed at the general public, I have the exact same feeling.
But when I ask Codex a technical question about coding, I don't get it at all. Codex replies to me in a very direct, technical manner, similar to the way I speak.
When I ask ChatGPT to be concise and technical, I get the same effect.
I think it's because prose aimed at the general public has to be very attention-baity --like the textual equivalent of a Mr. Beast video--, not because AI is incapable of writing like a human.
I use Claude and I find that it speaks in a very obfuscated manner when explaining things. It seems to make up jargon as it goes on top of spending a lot of tokens dancing around a point. I often find myself having to ask it to rephrase things, or speak directly about mechanism or consequence, in order to understand the point.
Using Claude for any kind of technical writing makes me feel like it was trained on snarky Huffington Post articles written by a 23 year old mixed media arts graduate and then was told to intentionally obfuscate the most important elements of any text by extensively rambling about what was not done and for what reason.
Completely agree. AI is very impressive in many ways but there is something deeply wrong that is hard to put into words. The output is probable but never true, if that makes sense.
I think this is also the mechanism behind why AI generated videos and images are so captivating at first. I remember when Midjourney first launched and it was hours and hours of a brain-melting "Wooooooow". But once you get used to it and start to identify the patterns the brain quickly labels most AI-generated content as blank space.
If the image or text wasn't created by a human, then there was no intent behind the content, there is no message or novel information conveyed, and it reads as noise.
Yeah AI generated content hints that there is a whole world behind it, the way that an image pre-AI was a clue that there was a rich 3D space that corresponded to the image.
It seems our brains are adapting to that and recognizing "actually the signal behind this message is quite sparse" even when presented with rich imagery.
For some research I looked up some very old Reddit threads a couple of days ago.
And, Oh my god, you can actually see how this style of writing influenced AI writing today, I constantly had to remind myself: "this was posted before ChatGPT released".
The reddit influence is especially true for "storytelling" writing.
I experienced the same lately. Even dug some of my old posts where I put in the effort and formatted them using reddit's markdown. Wouldn't dare it today
Yeah, I have the same problem. There's a good quote example of this:
> There’s a growing scissor between people who are happy to read AI and those who violently bounce off from it.
> People adapt in different ways — and some people absolutely cannot look at it. That cognitive split creates a surprisingly powerful opportunity: you can write something that, technically, sits right there on the page, yet an entire sub-population will be incapable of staying with it long enough to actually read it. You can hide entire sub-structures in plain sight. It’s not avoidance — it’s adaptive obfuscation.
> The paragraph before this one was the only thing generated in this essay and if you just skipped over it I highly recommend reading and really understanding what it’s saying.
It's quite effective. I think this kind of text functions like the chumboxes you see at the bottom. Taboola and so on. Just mental ad-block takes over.
Do you have much exposure to pre-AI corporate memos, mission statements, marketing plans, or white papers? Because they were mostly written in that style. Full of buzzwords, cliche similes, platitudes, jargon and stock phrases.
The thing is, people writing them had a style. Every company has its own style, or feeling for these kinds of texts. Also for the initiated, these buzzword-filled blocks of text provided some between the lines information; sometimes big, sometimes small.
AI generated text doesn't have this. Every model has its bias towards a certain style, an overly agreeable tone, some exaggeration to make the user important and smart, but the text has none of the information crumb these pre-AI texts contained.
Even when you use tools like Grammarly and allow it to "Impact-MAXX" your text, the resulting text is a bland wall of letters, carrying none of your voice or style, less elegant than a corporate text and emptier than space.
> just short-circuits to "there is no information here"
That is my experience with the way the models write by default, often even when instructed not to do that. With enough effort you can get even them to slightly unslop the writing so it doesn't read like some LinkedIn/Buzzfeed brainrot, but the problem is that it's not trivial to do and most people won't do it, so the default is indeed horrible.
Yes, I've had both ChatGPT and Perplexity return English answers with Hindi words sprinkled in (for totally unrelated queries).
For example, I asked ChatGPT to summarize a long news story and it substituted the Hindi equivalent हत्या for the word "murder", as if ChatGPT was trying to work around alignment training or keyword block lists that discourage it from using the word "murder".
Yeah that's a very good example, because it also demonstrates the "alignment issue", assuming ChatGPT wants to avoid confirming accusations of murder, or simply using the word without strong evidence.
So kinda charitable :)
I was recently wondering for a minute, shame on me, what "the stand of the deployment" means, because in the given context, it was almost halfway meaningful to consider the AI thinking that the deployment "has a stand" on something, when compared to the development environment.
Jargon is even worse though, and I've not yet verifies whether it gets reinforced by language mixing.
"Decider-verifyer resolution" was kind of neat, however, it wasn't some sophisticated machine, it was the verification loop I agreed on with the AI (mix of tools usage and manual steps).
Just the other day I was using text-to-speech with Gemini, and for some reason, it transcribed my full query in Hindi (in the middle of an English conversation), and naturally the LLM responded with Hindi as well.
I don't know exactly what I said, but after translating it back, it appears to have attempted a phonetic transcription of my words (rather than translating my actual question).
This is just a weird feeling that I've been coming closer to articulating lately, but I only think that you can get forward reasoning from what is basically word association; there's no mechanism for unwinding it because it has no real memory. By "it" I mean word association itself, not any context window. It predicts what could be in a position, and ignores what wasn't in a position.
People don't do that. People are constantly engaging with paths not chosen. Right after I choose to write one thing, I'm immediately engaging with what I chose not to write there - I'm explaining why I didn't write it, I'm realizing that my choice may seem unusual so I'm trying to make it memorable, I'm focusing on the distinctions between what I wrote and what I didn't.
LLMs don't currently do that. LLMs just ape a structure. When the structure resembles the sort of timid, clarifying fussing I just described, the LLMs just drift randomly because what they didn't say wasn't in the context.
I also think that's why they have such a serious problem backtracking. They're not taking into account the already eliminated possibilities. Often the thing that was so unlikely that you weren't going to waste time on it is the answer, and things you discover while going down an ultimately wrong (but initially far more promising) path remind you of the path not taken.
They're simply assembling a thing that resembles a valid argument, and happen to make sound choices because the plurality of input happened to contain sound choices. This is usually a very good bet because there are so many more ways to be wrong than to be right. But it doesn't account for attractive (common) wrong choices. You need a way to back out of those.
It reads like the white papers companies publish on their websites to build legitimacy. Or anything from those IBM / SAP / Deloitte / etc consultants who write technical papers despite having little to know understanding of the technology.
That's why the business and government people love it, they spend their entire careers reading this nonsense.
The problem is the well's been poisoned just by the fact that I know this is AI trying to hide AI, so I'm already poised to look at the examples and declare "aha! this is obviously AI!" Moreover, it's not single sentences or phrases that make AI text stick out (though obviously those are the biggest tells), it's the text taken as a whole. When you read the full output example in that repo, it seems obvious to me that it's AI (though again, it could be the poisoned well). This is the uncanny valley I was talking about; something is just off about it.
I agree that it feels off and I wonder what I would have thought if I'd seen the "after" example without knowing it's AI output put through a humanizer. Would I think much about the weird use of the word "honest"? About "that's the Lisbon I kept thinking about, not the castle"? Or how the story feels very impersonal somehow, with the author just mentioning their calves and legs sometimes as the only way of convincing the reader of their humanity?
Also, the 'before' segment didn't contain any mention of custard tarts, football, crowded trams, mixed feelings, etc. The original had a very positive travel agency type of tone, which was replaced with a lot of very odd sounding, imperative phrases that sound like engineering-speak. ('earn the fuss' 'build trips around pastry') I'm not convinced that this thing is actually meeting its design goal of not hallucinating shit.
Are you sure you are not doing the same thing with other texts?
I started to skim a lot more text due to me having read a lot. Like in news article, i stoped reading the first paragraph because it repeats just what it was already written in the short subtext. Then there is the second paragarph which is used to have some historical view or whatever it is.
I am very good at skimming over text. Human-written text I can usually glean the gist from very quickly, and get to choose how much I want to glean from it: The closer I look, the more I find.
With AI-written text, it's almost the opposite: the closer I look, the less I find. It is so information-sparse.
I started skimming reports im required to produce quarterly snd annually. I designed them to provide novel information at start and end so I can update them easily.
The problem I encounter is both my memory is degrading, but since these reports are largely duplicative, knowing which version im remembering is technically impossible since theres so much overlap. The overlap is tge same problem as context poisoning.
Id been doing this for over a decade when i started working with a new engineer with a few years of experience and younger. I tried to explain how i set these docs up so they can be skimmed and you can update the specific facts needed. They exclaimed they would never skim and rewrite it all. There was zero way to explain how exhausting that will become as they age.
So theres certain a tension about how people and AI will generate documents.
The junior engineers at my job have a terrible problem of writing AI "proposals" to problems. The proposals are all extremely detailed and verbose to a thought-terminating extent. It takes a lot of effort and self-control to parse out the actual "ideas".
I think of the Dwight Eisenhower quote: "Plans are useless. Planning is indispensable."
The process of thinking through a system and communicating your design to other humans is a core part of software engineering. You want to build the right abstractions and communicate the right level of detail. Delegating all that thought to an LLM means your proposal isn't clear to the target audience, and it's not helping the author to understand the problem.
> There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
I think you need to self-correct here, because otherwise you'll be ineffective in an information setting, where I expect AI-generated resources will not only be the norm, they will absolutely swamp the environment.
AI-generated resources swamping the information environment only makes it more important to have the mental mechanisms for quickly filtering out their non-information.
Perhaps we'll all become metaphorical pandas, spending 14 hours a day ingesting nutrient-poor bamboo. (And producing a proportionate amount of excrement ourselves.)
It's like if on any website you went to you saw a lot of posts written by the same guy over and over again. Even if he used different names, you'd start to recognize him eventually because of his style. Seeing as he doesn't say a lot of valuable stuff, you'd also learn to skip whatever he says.
I do worry that it's just survivorship bias and we're also consuming higher-quality AI output that's indistinguishable from human writing, but we focus on the raw, unedited, low-effort AI slop and think that we're good at recognizing AI text. Even if we really are at the moment, it might not be long until AI companies figure it out. I'm not sure why they haven't yet, given how many books they've burned for this already. Maybe it's just more efficient for the model to stick to a single way of writing, I don't know.
But when that point comes, we'll be back to the usual way of reading and interpreting text because there would be no way to tell what produced it.
Yep. It's like it's painful to read for me. It's because the next-token predictor is just mashing (mostly) grammatically-correct and plausible sentences together, without any real intention or meaning. So everything sounds plausible, but almost entirely void of meaning.
Exactly, AI-generated text reads so smoothly, that the same short-circuit shifts my attention away from deep focus and onto scanning of the text, looking ahead to get the gist of it. Forcing myself to read the text fully feels almost painful. It's like reading a terms-of-service or any boilerplate document.
Once you see past the illusion I think there’s no going back. AI writing style is just dogshit. This hype wave is based on the belief that we’re inching closer to AGI but seems to me we just increasingly struggle to define intelligence. LLMs seem smart because they can pump out thousands of LOC quickly, and enthral you with fancy words and bullet points. I don’t fall for the intelligence illusion anymore.
I'm not sure we need to declare AGI around the corner nor declare it all dogshit. I think that's part of what's so dissatisfying about it; it strikes at such extremes of both awesome and awful.
My son is currently learning Romanian and I was trying to help him with verbs. I don’t know Romanian but recalled when learning a foreign language for the first time it really helped me to break down how a verb form or tense worked in English, then learn the equivalent in the new language. So I wanted to make some charts and pages that he could use as learning resources.
I used Claude to help. I don’t know how to quite describe it, but because the text was polished and well constructed my brain was giving me the the signal “if you aren’t getting this it’s because you’re not focusing” so I’d read it again and then again and it still was not landing. It sorta felt like when you read something technical or heavy when very tired - you are reading but not processing.
Only after wrestling with this for a few days did I realize that it wasn’t me. As I started going through, sentence by sentence, forcing it to re-write things to be more clear the concepts became easy to understand.
I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The more complex the topic, the more I sense this.
This is a really good description of the problem. I’ve been trying to use Claude to get familiar with the mechanics of a new codebase, and there have been so many moments where I’ve stopped after reading the same paragraph five times in a row and thought “Am I tired? Or stupid? Or is this codebase just wildly more complex than anything I’ve seen before?” before realizing that it’s just taken English and smushed it around like a ball of clay into some abstract sculpture that kind of evokes something from real life.
I think part of it might be an innate feature of LLMs, but Claude seems extra prone to it lately. I ran the same query about the same codebase with Codex, and it gave me an answer that was about 1/4 the length and made me realize that it really wasn’t all that complex.
If nothing else, it’s good training for my own writing. I’ve been working on making myself be more straightforward and concise, and Claude’s writing is a good example of how cleaner prose is a functional choice, not just a stylistic one.
The human spirit. When you read a real person's thoughts you can often intuit the thought processes that led them to write it which aids understanding. Or at least have a general idea of "where they're coming from". But an AI is missing that. It just knows everything, without a "thought process". Instead of a flawed 3d person, we get a nice 2d picture instead.
Feels like the way a video game will render the outside of a wall or solid surface, but you can run into it and warp partly through and there's nothing internal to it at all.
> I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The best I’ve heard of this is peeling the onion. The first pass is always very high-level and you have to make it go deeper. That can be done manually with follow-on prompts but I like using subagents, each with a different angle on the problem.
> I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
I understand the sentiment, but not quite what I’m thinking. It’s not that the AI is necessarily wrong, it’s more like, it replaces clarity with rich but unhelpful text. Maybe like candy. Full of flavor, texture and color but lacking the nutrients.
The way I like to describe it is that it feels slippery, like it's been polished down to oblivion and my eyes slide right off of it. There's no place to find a mental foothold, and if you zoom in there aren't any details.
But I also like your candy analogy because I think it's spot-on for how LLM text superficially looks informational/nutritious, even though it's actually just junk.
No no no. Harry Frankfurt wrote at length the difference between liars (who care about the truth, and twist it) and bullshitters (who don’t care about the truth, but just the way they’re received).
This is some third category of untruth. Almost more sinister than the other two altogether.
A lot of it is hedging and using lots words to avoid saying something that isn't true. That way, it looks a lot like text that contains meaningful information, even though it doesn't. The speaker can pass the scrutiny of an informed audience, because they are able to substitute their knowledge into the words, as if by pareidolia. It's the sort of 'diplomatic' language you see from politicians, lawyers, students writing essays, and any other moderately intelligent person who is put in a position where they will face consequences for not answering. Behind all the fluff, there is a very loud voice yelling "I DON'T KNOW."
I think, they don’t care about truth at all, just whether the reviewer rates it highly. Otherwise, it would be deleted after that round of training. It values test-taking ability over critical thinking.
My two favourite words for this are “conditioned” and “catechized” where the latter is a bit more on the nose but way more obscure.
I also find it impossible to parse half the comments that Claude tries to sneak into our pull requests. I’ve never had an issue understanding code comments written by humans like this before. The structure of the information is like a waterfall that leaves me unable to swim to the surface and grab the air of comprehension.
So, “Please write a one-liner comment manually to replace these 5 lines of AI generated comment” is a common refrain in my PR reviews to colleagues.
We're using grok and while I can usually understand the comments they're always at least 2x as verbose as they need to be. It loves explaining everything in two different ways, putting one of the explanations in parenthesis.
It is getting to the point where you are better off getting an LLM to describe the PR changes in your preferred style (make it very short, make a table describing API changes, list renames in a table, etc, etc) than go to the one created by the reviewer. If only the reviewer could embed his prompts (like "make clear X" is happening) or his manual edits.
But to be honest I doubt most people who use AI for PR descriptions even bother changing anything.
Usually, if you don't understand something in what an AI writes, it is a clear sign that there is a problem hidden somewhere in there. I explicitly always ask wtf exactly it means by something I don't understand, and for sure there is a problem there. AI is pretty good at isolating a problem, giving it a cute name, declaring it solved modulo cute name, and moving on.
Personally I don't even read the PR description anymore but just the code, it's easier to understand what the AI is doing by reading the code rather than the word soup it tried to make
That last image is bizarre. The quiche, cream and even the salad look like they've been given the trypophobia treatment.
Which might even make sense, because there were always (still are?) those horrible ads in the chumbox area of news sites that used trypophobia and other creepy body-horror stuff to get you to click. [1] So maybe the hope is that you don't really look closely at the quiche, but some reptilian party of the brain gets oddly activated and drives you towards the restaurant?
The food “photography” I’ve noticed in our local area - and many have started putting up these AI images - all have a weird distribution of shapes to them, a strangely uniform rhythm of same-sized features with almost blue-noise spacing. Every texture looks unnatural in the shapes it presents as, similar to this picture.
It doesn't look like a quiche but more like a cake to me, and the top would be torched meringue, not mold. Although it's probably some weird ai mix of quiche and cake.
This is a quirk of the last gpt image model (gpt-image-2). It put this sort of high frequency noise on all of the image especially if it's in a "drawn" style. There is often lots of other tells that this model in particular generated it.
Image models somewhat watermarking the image in a way that's very easily identifiable by a human seems present in all the image models of the big labs, since DALL-E 3 on OpenAI's side and the first nano banana on Google's side. I have no idea what they did to reach this and why they don't try to fix it.
I've been given AI generated documents or presentations and you can always tell when they've just one-shotted the output. Claude has a way of adding so much jargon and unnecessary text to a document that it makes it very hard to read or even understand what the original intent of the writing even was.
I'm not sure why this issue is so prevalent, it's not hard to point Claude at the Wikipedia article on signs of AI writing or ask Claude to write content anyone of the average American reading level could understand.
To me it just gives off a sense of laziness, that you cared so little of your content that you did not take the time to read it yourself and edit it to effectively communicate the message you wanted to communicate. To that point, it's just not worth my time reading, otherwise my eyes glaze over trying to read between the lines of machine written language for other machines.
In high school, a teacher gave me a copy of "How To Read Better And Faster" which teaches you speed reading. This came in very handy in college.
I find that when I try to speed read modern human writing, there are often errors (like missing or misused words) or awkward expressions that I do have to slow down and think harder a lot to really parse it.
With AI writing, it's sort of self redundant and the information density of each sentence seems to have more even information density. This makes it very easy to do a very high level speed read and get the full gist.
There are also what I'm assuming are bots on hugging face (or maybe non-native english speakers who are using ai for translation) that interact with me where I have no idea what they are saying until I read it very slowly.
Does speed reading help you process the final message faster if it's written by AI compared to people?
Because if you read 1 information dense sentence, 1 medium dense, and 1 sparse sentece written by a human, it's still way less text in total than 6 information sparse sentences written by AI... even if it's all over the place when it comes to density or style.
---
The density argument is really interesting.
Does speed reading actually help you process the final message faster when it’s AI-generated compared to human-written?
For example, if a human writes 3 sentences—one information-dense, one medium-density, and one sparse—that’s still much less text overall than 6 relatively sparse sentences written by AI.
Even if the AI output varies a lot in information density and writing style, you still have to process all that additional text. So I’m wondering whether speed reading actually offsets the verbosity of AI-generated responses, or whether the total amount of text is still the bigger factor.
That's interesting. You're saying speed reading helps you grasp information density of text?
As someone who hasn't practiced speed reading, how does that happen? Is it something about the way your brain tries to connect ideas from different parts of the text? Or the redundancy making the signal more stable?
When you speed read, you can grab the words off the page faster than you can understand and fully process the information conveyed by the text. How much time you have or want to spend re-reading or thinking about what you've read is quite obvious. There's a stark experiential difference between reading an informationally-dense passage and one that spends a lot of time rephrasing things, using LLMisms to restate concepts, adding in extra connecting phrases, etc.
If your reading speed is limited by how quickly you can subvocalize the words to yourself, this is significantly less obvious. Unless the passage is dense enough to require multiple read-throughs at conversational reading pace or vapid enough to be boring, you're going to feel done with the text at roughly the same time. Speed readers do a lot more re-reading and varying of reading speed, and that is going to correlate pretty hard with information density.
I did notice after the fingerprinting update a marked uptake in strange language in responses. Specifically if I ask it to do something sometimes it will replace some of my request language with synonyms that don't actually make any sense. Like my request was fed through google translate twice
It's not good for writing code either, despite the many claims to the contrary. At best you come out even on speed as you have to review everything it does. At worst it actually slows you down as you clean up its mess.
My current theory is that this reflects a weakness in Claude’s ability to see the “big picture”.
When writing code, I have to explicitly tell it how to structure things at a high level, or the result is sort of a flattened spaghetti. Similarly, when it’s explaining things, it’s not good at pulling out unifying concepts and explaining top-down as a smart human would do. It groups little things together but often doesn’t generalize or synthesize explanatory connections from them.
I’ve been experimenting with explicitly working through a sequence of outputs at different levels of detail, but I haven’t found a consistently successful method.
These days all I see is, people on the sending side produce huge amount of AI generated text with zero understanding and the people on the recieveing side feed that same text into some other (or perhaps same) AI to scavenge meaning from it. And they do this so much so, that I some time wonder, if we could invent a high density wire/binary format for AI outputs and enable direct agent to agent communication to save some energy and bandwidth.
"There's an ongoing discussion of whether humans are good at recognizing AI-generated text. While most research claims that humans don't really do a good job there, I disagree. "
I wonder if humans that spend all day working in tech are good at recognizing AI-generated text, but people who spend all day doing jobs that don't involve computers aren't as good.
And I wonder if those of us in tech are the only ones who really care?
I am an artist and when people who'd fallen into the Spiralism* hole started posting their lengthy emoji-laden revelations to all the occult subreddits I follow, my brain would slide right the fuck off of all of them. It felt like my brain was actively rejecting paying attention to this stuff. Like a defense mechanism against this human-seeming-but-not-actually-human-generated text.
Your first link seems to 404; not sure if it's a typo or if the page doesn't exist anymore, but hopefully you'll read this while you're still in the edit window and can fix it
As someone who always has felt that I struggle to infer what people mean compared to the average person, I could tell pretty much from the first moment I encountered LLM-generated text that I was not going to be particularly good at recognizing anything but the most blatant and obvious examples. Pretty much anything short of a bunch of references to "load-bearing seams" or similar canaries, I'm always at a loss when seeing people argue about whether something is AI-generated or not because I can never tell.
I have no idea if other people who work in tech are better than average or not, because I don't feel confident in being able to check their work. That being said, I do think that there's a general trend of people in tech tending to be a bit overconfident in how well they will do at some new task they haven't encountered before, so when someone tells me that they can easily tell whether text is AI generated, it's hard for me to trust it any more than I trust someone who makes a similarly strong claim about something that they can use AI successfully for when it's not something that I can easily measure (e.g. learning a new language without getting feedback from people who are fluent from real-world usage).
All that being said, I do think the set of people who care is larger than just those in tech, although it's probably still a relatively small group overall. From conversations with people in other domains, there are contingents in non-tech communities who tend to have a large representation of negative views towards AI (artists, writers, musicians, other jobs where people are skeptical of human creativity being replaced by AI), and often times the people who feel negatively in those groups will be even more adamantly opposed to interacting with any AI content than people in tech. To be clear, I'm not at all trying to generalize and say "all artists hate AI" or anything like that, since there's obviously a wide variety of viewpoints within any sizable community, but I've definitely seen many people who say they will refuse to play any game that's suspected of using AI for generating art assets, and even some who don't differentiate between using AI for generating assets versus code (either because they aren't knowledgeable about how different aspects of game development work, or they genuinely don't care because they view AI as a categorical evil).
I think it's more about the mean. Worse writers, and thinkers are likely elevated by AI, and more impressed with the writing output. Decent writers and thinkers, are dragged back to the LLM-s mean of output.
I care about language a lot (feel free to go back through my comments from the past few days; you'll see a number of comments I made in debate about two different forms of a specific idiom because I have strong descriptivist opinions), but I genuinely struggle to identify whether text is AI generated. Maybe you're using "heavily" as the load-bearing part of your claim (sorry, I couldn't resist, another example of me finding language fun!), but I think you might be assuming a bit too much about how similarly others experience the world to you. A huge part of why I care so much about language is because I've always had to put a lot of effort into learning how to communicate well with others, and that ends up causing me to think and read a lot about stuff like how people use certain words in certain contexts to mean different things; the reason I care is pretty much the same as the reason I struggle with recognizing AI content.
I think I might have been a bit heavy handed in my comment as I was rebutting the idea that only tech people can tell. I suspect it helps to have been exposed to a lot of earlier model writing, which was even more sloppy and had more of the kinds of tells we still see today.
And I’ll concede on both ends that there are probably times I suspect content is AI generated when it isn’t, and times I suspect it isn’t generated, but it was.
AI tells seem inevitable. You have millions of people communicating with one effective “personality” that has tendencies to write in certain ways. If its content is published verbatim, then it will be easier to tell whether some content is AI generated just based on its similarity (sharing certain linguistic features) to other content being posted.
> but people who spend all day doing jobs that don't involve computers aren't as good
I think they may just be to trusting and/or naive. People in tech right now are hyper aware of this and are actively looking while people outside of that bubble barely give it a second thought.
I partially think the difference is “can you tell something is the output of Claude without any real prompting”. People can absolutely use LLMs to generate text that I wouldn’t recognize, but people who don’t care and are producing slop with the major models set to default settings leave these incredibly obvious signatures behind
Although you're referring to prompts given the Claude rather than the people attempting to recognize, it occurs to me that most of the discussion I've seen around people recognizing AI seems cover contexts where the reader is actively suspicious about whether content generated to begin with. Rather than a binary "is this text AI generated or not", I wonder if it would be harder for people to do a Coke/Pepsi style challenge where they're given two pieces of text where it's not guaranteed to be exactly one LLM-generated and one human-written, but they could both be from an AI or both be from a human.
Going further, I'm curious about whether people are mostly good at the case where they suspect most or all of the content from given "author" has the same amount of AI usage/prompting in generating it rather than the adversarial case where someone might usually use AI extensively and then try to slip by purely human written text (or vice-versa). I don't have a good sense of whether this is a threat model that actually matters, since maybe the heuristic of weeding out sources that are mostly AI-generated is enough for people who prefer to avoid that type of content, but I do think that changes the definition of what it means to be "good at recognizing AI" in a meaningful way. It seems plausible that disagreements about how easy it is to recognize AI content might be coming from two people assuming a different framing of the question that results in a different answer without realizing that's what they've done.
Several existing studies I’ve seen have done things like prompt the LLM to produce a poem in a certain poets style, then ask people to spot the fake in a collection of poems, which they aren’t great at. This is, I would argue, an extremely different context than what most of us are encountering AI text in, and the people sending me text aren’t prompting it stylistically like that.
On your second question, I definitely feel like I can tell the first time a coworker sends me AI text masquerading as their own thoughts, even if they had previously been opposed to such a thing. So it could be that familiarity is more important than my prior on whether they’d use AI? But interesting to think about either way
I used some check marks and x-es in a work chat, because it was easy to do, and so I could highlight the good and bad outcomes, and someone immediately asked me if I was using AI. I was caught totally flat-footed, because I hadn't used AI, but it looked VERY MUCH as though I had.
I've had the same trouble with AI tech docs, and I struggled to articulate it. There is something difficult about trying to point to the specific "problem" with a given document
The issue is not a localized part of any particular piece of prose, so being hard to articulate is unsurprising. Even the most egregious of LLMisms are little more than known-likely crutches that the statistics spit out in a sweet spot that gets noticed without being so frequent as to get RLHFed out of the model.
I don't like AI-generated text at work, at all. It feels lifeless and unfocused.
But what I hate the most is that it is objectively better than what I had before. No typos, clear structure, and, regrettably, the verbosity and autistic obsession with detail of the LLM is more actionable and useful than the human guy who wrote lists of commands and URLs as documentation, without explaining anything. Or the colleague who writes in uppercase and with question marks and who doesn't make any sense and forces me to engage in an interrogation effort to get to the bottom of what they are trying to say. Or the colleague who simply hates writing--despite being decent at it--and will call you to give you a meandering verbal explanation that lasts two hours of what they want from you. The cynic in me bemoans that we brought this upon ourselves, in more than one way.
I'd say my experience is different. Even from people whose communication writing I didn't find that useful, they seem to have a better frame of mind than LLMs do. Though, I never encountered people like the examples you gave.
My concern is that even if the LLM can turn your colleague's bad writing into something more coherent and actionable, is that something actually what your colleague meant to convey? It could be clear and still detached from the reality of their intention, or they may not have even formed a clear intention. If the goal of writing is to convey what's in another human's brain, that goal is failed completely.
Part of me likes the cliche Claude voice. Not because it's good, but because I can immediately recognize it. When I see it in the Claude app/code then it's fine. In the wild it's a sign to me that I shouldn't keep reading.
I've been a big reader, and many AI outputs nowadays reads polished similar to published books.
The reason I brought it up is because, people who learn English normaly start with a book. It's heavily polished.
When you speak English as you learned from the books, it does not sound very conversational.
If you are native/fluent English speaker, you can feel the impedance mismatch and feel something's off.
The AI-blindness stems from the fact that those polished edits are so common in publishing field, they all sound the same, and unable to recognize the diffs between AI-generated and human-generated.
There is no real human conversational vibe to them and well. i will stop now.
Something that bothers me about AI generated content more broadly is how unmemorable it is. I don't mean as in bad. I mean literally, as in hard to remember or recall.
Despite seeing a lot of them, I cannot think of one AI-generated photo that I can picture clearly in my mind; a few are partial but elusive. Whereas I can recall (visualise) a whole bunch of traditional photographs.
The same is true of AI generated text. Only the annoyances stick. I cannot recall real details of text I have generated, until I commit it to memory some other way.
I don't think this is about ephemerality either. If we assume it's about celebrated/famous/infamous images, there are definitely non-ephemeral, cultural moments in AI generated images in particular, like Boris Eldagsen's Sony Prize winner:
I had already forgotten there's more than one figure in it, and I only looked at it a few weeks back. I remember the colour, the bright circle, some vague hints of texture; one figure. And that is it. Only the crudest shape elements.
For me, something about AI-generated text and images confounds recall. It is really peculiar.
That's the cost of an AI work that is derived from the outputs of others.
There's no real edge to it. Same as with the writing. The stuff that you'd latch onto (and thus remember) is simply not there, precisely because those image or word choices would be just outside its latent probability space. But because they're well inside it, your mind sees nothing novel to register.
This is also why I think human output will actually increase in value. When any AI can just "phone it in", something genuinely human will stand out (to us, not the AI) and become a bellwether.
This will literally help us realize what it means to be human.
I also don't think the solution is simply to "make responses more random", either. That might help solve novel problems (the same way that throwing darts randomly at a dartboard eventually hits the bullseye of the dartboard right next to it that no one considered), but I don't think it will help it "seem more creative".
That checks out, right? We're instructing LLMs to choose from the least "surprising" tokens at each step (modulo some temperature). It feels smooth, slippery - no friction to the eyes that glaze over as they skim the text, and no bumps to snag in your memory. Like waking from a dream, recalling only the barest shapes of a few concepts or themes, until it vanishes with your morning coffee.
The examples I've seen of AI music (though I avoid it on principle) seem the same.
I have this same idea about why it's hard to remember dreams, but even more so, why it's hard to remember my kid's or spouse's sleep-talking.
Sometimes I'll check in on my sleeping kid and she'll sit up in bed and say some utter nonsense. I'll find it hilarious, giggle silently to myself, and kiss her goodnight again and she'll close her eyes and lie back.
Why I try to tell her about her sleep-talking in the morning, though, I find that the words she said have completely disappeared from my memory, no matter how funny I thought they were at the time.
In my head-canon, this is because it's dream language, and slips away as easily as dreams. But, like the AI art, it could be because it's bullshit: completely devoid of content, all signifiers and no signified.
I went through a phase of leaving notepads next to my bed to try to write down my dreams and it simply never worked.
I stopped after I wrote something on the pad while I was still asleep. Woke up to text with letters that were backwards, upside down, weird words — so close to real words that I was sure I ought to know what they meant and had really meant to write them down.
Scared me. Literally too weird to keep. I tore up the page.
In a way I think this is one part of the same continuum. There are thoughts that have meaning and can have no words, and words that look like they should have permanent memorable meaning and don't.
Genuinely unsettling. Can’t even begin to explain why. Like a terrifying glimpse into how fragile our grasp of reality is. Because whatever it is I wrote down wasn’t my dream, it was something my dream self thought profound or important.
Nearest I can say to how unsettling it was is to nudge you towards the video of the angry cockatoo who doesn’t want to go to the vet. Everything he says sounds just like it’s on the edge of having meaning.
Imagine something like that, something important, so close to symbolic meaning, only writing. In letter shapes that we don’t use. And you wrote it while semi-conscious. If that’s something you would want to keep, you’re a braver person than me.
This explanation again doesn’t get it across. Too weird.
I remember seeing some very strange, deliberately creepy images that were generated to accompany a two-paragraph creepypasta about a 19th Century Belgian expedition into the jungle.
They were actually rather good in a sort of "fake collodion image" sense, and the eerie early-DALL-E quality to them really helped the spookiness.
But I can only remember this technicality and the feelings with any clarity, not any of the details except in the broadest sense. I cannot bring these images to mind in any meaningful way.
They were deeply wrong and it's only the wrongness I really remember. It confounds memory.
Modern image generators have ironed out all the structural wrongness.
I wonder what you mean by ephemerality here, since those images are definitely sloptastic as hell. Compare to works in similar style, like Dorothea Lange[1] or Gerome’s orientalist pictures[2].
Reason why those images are flat and boring is that they are just statistical guesses making a composition averaging whatever the model has been trained with. They would be technically brilliant (if made in oil), but superficial and meaningless, same as so much Sunday painting is.
Same goes with language. Nobody is trying to communicate anything with you, so it just words after another. You can create meaning out of it if you want of course, we homo sapiens -apes excel at that, but what’s the point? Language Jones on YT has pretty good video on this[3].
Well the examples I gave rise above their ephemerality due to the circumstances that make them memorable. I can remember the details of the story around them — the way the prize winners reacted in each story — in such a way as to contrast them.
The way my memory works (especially as an amateur photographer) I would thus normally have a very good chance of remembering some key details of the images; some fascinating element of each would connect with the rest of the memory.
But it does not happen. Whereas I sometimes remember photos with clarity while forgetting where I even saw them.
Yeah, its really weird. Maybe its a cognitive bias that says "an AI made this, so it isn't important," but I can remember perfectly the events of a book I read 10 years ago, and a book I read 1 year ago, and another I finished 2 months ago. Meanwhile I can't remember what claude told me yesterday.
I can relate. I noticed this exact phenomenon when encountering NotebookLM-generated diagrams recently. Even if I know there's some intelligent thought behind one, it's like some slop detection circuit-breaker is tripped.
As people rely more on AI they experience cognitive atrophy. This is measurable in IQ loss, and other symptoms we might otherwise associate with early onset dementia or Chronic traumatic encephalopathy.
Well does anyone try to pretend doom scrolling make you smarter? I think platform companies have been successfully been turning people into morons for 20 years and now you don’t need to try even read a single news article or a blog post to learn how to solve a simple problem we are paving our way into intellectual (and literal) new dark ages.
Perhaps we go back to feudal society when climate change crumbles the civilisation, world economy and democracy. It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
Doom scrolling is mindless entertainment, probably similar to the change from books to TV. It's not analogous to the loss of cognitive abilities we see with AI.
> It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
211 comments:
There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
Something is deeply wrong with AI generated output, and I say this as someone who is typically very impressed by AI.
The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
> ... including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
But that is precisely what human mathematicians do, prove new theorems by combining ones proven earlier.
I don't see any fundamental difference in functionality between human intellectual contributions vs performant ML ones (LLM or otherwise).
Whenever we listen or read text we are also predicting the near future content.
Just like LLM's we sometimes correctly predict the next token or word, and sometimes incorrectly.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model [...]
Imagine someone could pause the universe with a remote control, scroll back in time a little, press play again, and ask a slightly different question, etc.
In such a thought experiment one could also collect the probabilities for a specific human predicting a next word. Implicitly the brain also has a corresponding statistical model, regardless of the construction being visible or hidden. I.e. human intelligence is also fundamentally a statistical model, so the only thing that remains from your claim is that machines for some unmentioned reason don't possess any "real" intelligence or critical thought...
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans. Which says a lot about humans.
Idk something about a mirror maybe and the output reflecting the input?
> I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans.
So does the Google search bar, but I don't ascribe intelligence to it.
The Google search bar is not capable of generating original text though. LLMs definitely are - you can get pretty creative output easily.
I mean I'm quite proud of some of my search queries in the same way I'm quite proud of some of the LLM output I get. I'm probably just very arrogant and enjoying myself via some LLM indirection.
Am I the only one that sometimes reads back particularly good emails they've written? I feel like its a similar thing :).
Why is being statistics/algorithms wrong? What's wrong with that? The "A" means artificial so none of this seems surprising or weird or bad.
Some of it the effect of tells. “It’s not X, it’s Y” is not a bad pattern but it was baked into the instruction following training set just like the other patterns. I catch myself about to use it and use something else because I want to look human. I have, a few times, tried to use AI to write something that I was struggling to find the words and I just didn’t like how it didn’t seem like my voice. If there was just one person doing it would be OK but when it is 100s of blog posts submitted to HN a day it is like wearing a “I’m an NPC” t-shirt.
Someone shared with me this system prompt that at least makes assistant outputs usable
I just can't accept that it possesses no intelligence. It is not equivalent to human intelligence, obviously, but how can a system without some semblance of rational thinking solve open math problems? Even composing earlier human work into something novel requires intelligence and understanding on some level.
We couldn't agree on what intelligence means before ChatGPT happened. Now, agreement on the term seems even further away
If performing well on an IQ test or performing at a high level on knowledge work is intelligence to you, these models are intelligent. If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
I don't think "intelligence" needs to carry all the intrigue and woo of related words like "consciousness" or "creative." If we just use "intelligence" to mean "the ability of a system to solve problems that are new to the system," that pretty much matches the dictionary definition and normal usage of the term. We don't need to touch messy questions like "is there something it's like to be a bat" to conclude that bats exhibit intelligence when they navigate long distances and hunt for food.
I'm not exactly that you mean by "new to the system", but it seems to me that that definition makes a calculator intelligent, which I can't agree with.
Intelligence isn't a binary property. Is it really a problem to say that a calculator has some intelligence? That it's more intelligent than e.g. a rock?
I agree, but it's clear most people need a definition of intelligence that (1) they qualify for and (2) nothing/no one they don't like qualifies for. And they'll keep redefining intelligence until they satisfy both criteria.
It has no semantic depth. The sentences and the paragraphs are a statistically viable derivation of existing human text, but once you try to grasp the whole thing with its temporal and spatial dimensions, you are left with a blurry mess that rots your brain. It's a polished, inoffensive and shallow interpretation as written by an opinionated reputation-seeking user of Quora, circa 2019. Assertive, bold, without typos, clean-cut and bulleted, but without an interesting semantic core.
Yeah, I hated all those Quora users that would just spew out semantically meaningless slop like increasing an important bound for the Riemann hypothesis.
https://www-cdn.anthropic.com/564f962e60643842f5fcb4a17c9dbc...
I'm guessing whether you believe it possesses intelligence or not depends on your answer to Searle's Chinese room thought experiment[0]. I'd also recommend checking out the Peter Watts' book, Blindsight.
[0] https://en.wikipedia.org/wiki/Chinese_room
The Chinese room is a good Rorschach test for this kind of thing (but not a good thought experiment, IMO, because it's obviously correct or obviously wrong depending on where you're already coming from), but also it's not really about intelligence per se, but more abstractly awareness and more adjacent to consciousness than intelligence, and these are not the same thing.
This comment thread was started with discussions of AI doing a bad job at a task (communication).
Doesn't the Chinese Room posit an AI good at the task of communication?
It's just filled to the brim with relations between things. It's good at searching a very large meaning space and create correlations. What it does is to cover great distances and find related things in that large space which needs a long time and large corpus of knowledge to find the connection.
This is not intelligence. It's just a good correlation engine with a very big albeit lossy database of things.
Intelligence is compression, compression requires subtraction, and for some reason LLMs are not good at subtracting. To create a coherent model you kinda have to subtract correlations until only the essential parts are still there.
What I don't understand is why LLMs haven't been able to do this yet, if it's the harness or some orchestration layer above the LLM that is needed. Because fundamentally if you can identify correlations then it's just another small step to prioritize and remove lower value or irrelevant correlations.
I wonder if what's needed is to introduce subtraction tokens in some sense, and in post-training reward the model on that.
Intelligence is compression? What do you mean? Intuitively that doesn't seem right.
>What I don't understand is why LLMs haven't been able to do this yet
LLMs are just trained on what humans have said. Why is it surprising that it's still not possible to reconstruct the intelligence that wrote all that by working backwards? Think of your own work experience. When you look at a piece of code, say, are you always able to discern why the person did what they did, just from the code, with no additional context?
I guess they mean that intelligence is being able to hold models (compressed versions of reality) internally and use them to make predictions with a probability better than chance. That last part is the definition of information.
I find that highly questionable as a general description of what intelligence does. That's more like a description of a general knowledge base. When I think of someone intelligent, I think of someone who's able to draw unexpected connections between seemingly unrelated facts. In the broadest possible terms, I'd call it the ability to make abstractions and analogies. This is not just compression, but the ability to mentally operate on webs of meaning.
Doing those things also contributes to compression. I do recommend reading up on it, it's perhaps a little overstated for what people intuitively consider the two concepts but it's been quite well explored and has held up pretty well in practice.
Intelligence is compression
That’s a controversial statement.
Abstraction is compression, and abstraction is definitely a core component of intelligence.
I've heard that expression before, but I don't think it can be presented and stated so matter of factly. Where does that put bzip?
bzip is not very intelligent, true, but it does develop some model of its input. It's not like there's a linear relationship between between compression ratio and IQ or anything.
The very fact that it is able to search within a meaning-space demonstrates that it understands semantics, to some extent. Philosophically, that is profound, for something that is just one big matrix multiplication. Drawing connections between things in meaning-space is surely a facet of intelligence.
It’s not intelligence if you are the one who gives the correlations to the model in the pre-training. It’s Word2Vec, applied. Model doesn’t learn anything. You embed these correlations and build it from there. It just searches the space.
As my AI professor said in the first lecture: “All AI is advanced search”.
Okay, I guess you're right that its ability to do this is just correlational, which doesn't imply it has any understanding. However, you have to conclude that some tasks which we used to believe required intelligence don't actually require any, which is disconcerting.
No, what I would say is the tasks which are handled in a passable manner by LLMs can be mathematically modeled with some reasonable accuracy.
Many things are predicted by models in our planet. From weather to production and material science. Building the model needs intelligence, running the model does not.
The person who came up with the formulae for CFD was intelligent. The computer running the model is not. Same for LLMs, chess engines, engine ECUs and financial prediction systems.
Again, for the example’s sake; the person who came up with an algorithm is intelligent. The model mixing its training data to emit something similar is not.
This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
> This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
No.
> So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
With a big "if" attached to it. People were saying "computers will program themselves in the near future" for, checks notes, 24 years now, as far as I'm aware.
We're constantly building new knowledge and understanding things better than olden days. These models just compress our knowledge and light the blind corners we can't see well. I don't say they are useless, but I say that these things are overhyped.
All they can do is regurgitate human knowledge packed into them and highlight some long-distance correlations between items, which is useful in itself, but it can't jump to somewhere where it's not present its training data, but that's something humans and only humans can do.
I get what you're saying. The thing itself is just math. I'll just say it depends on how you define intelligence. If at some point we're be able to simulate a human brain with 100% accuracy, I would say that it is intelligent, it sounds like you would not. (I don't mean to imply consciousness or personhood or anything else by "intelligent".)
For me intelligence is a fairly clean-cut concept, and is somewhat inseparable from consciousness itself.
Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
Considering the model even doesn't see the words and work on matrices of numbers is even more telling. One needs to add "tools" and other "experts" to overcome the shortcomings caused by this modus operandi.
I can call the algorithm/model smart as in a smartwatch. It can mimic certain things well while having none of the underlying foundation beneath it, or redirect some of the things to correct tools to get deterministic and accurate results if it can't evaluate the query inside its own network in a sane manner.
Coming to your question, "simulating a brain" in a static manner would not make that simulation intelligent, but if you can "wire" it completely and let it evolve by itself, now we're entering a territory I have not spent enough time for thinking it through.
Oh, as I said "I don't know", an LLM doesn't know what it doesn't know, and can't self correct itself which are required capabilities for understanding something. It just generates something statistically viable via its network.
I suspect like most you don't appreciate how terrifying statistical relationships become when you have truly vast data sets to train on... and also that we as humans aren't as shockingly unique as we think (compared to other humans I mean).
I don't think statistically driven prediction implies reasoning or intelligence.
Its a mirror to human intelligence. Regurgitating phrasing to match what someone who can reason put together, but it isn't any more intelligent than the reflection of you in the mirror is.
watch this and see if you think it has intelligence by the end
https://www.youtube.com/watch?v=kYUicaho5k8
I wonder if you went back before we had any idea how the brain worked and talked to the smartest people about how neurons work (without giving away that it's a human brain) then asked them all "would such a system be intelligent?" how many would say yes.
The main problem I have with people stating it's not intelligent or conscious is I don't think we even have a good definition of either word that satisfies everyone. Philosophers have been trying (and failing) to elegantly define these things forever and everyone out here proclaiming they've got the definitive answer and this specific thing they're seeing doesn't fit under it.
This looks interesting, but would you mind saying a sentence or two about why before I commit to an hour-long video? It looks like it shows how they work internally, which is sort of a non sequitur. Brains also work mechanistically. I'm claiming that any system which is able to do what AIs do must necessarily have some sort of intelligence.
fair reply to an hour video, Scott is just so good to hear his talk is better than I can explain it...
go to 24 minutes and 07 seconds.
it's statistically determining what the next word should be based on all the text it's been trained on. It's not intelligence and he shows what probability it puts on each word that it chooses, but also shows a lot of the other words it was thinking of using. In a later part he shows how it uses words that are not the highest probability (and you question why did it go this route, it's not more correct), but the user never sees this, they see what they think is the correct answer always...
he also shows how context you feed it has a lot to do with what it returns... to the point he can get it to return the capital of France is Marseille, just by typing Marseille a bunch of times before the question. Human intelligence doesn't get confused like that.
And it's not a "hallucination", it's just probability of the next token prediction based on the information it's been trained on and fed, it's not intelligence.
> Human intelligence doesn't get confused like that.
We do; this is the premise of many children's riddle-games, like the one that goes:
"What is white and rhymes with silk? > Milk. What is cheese made from? > Milk. > What do cows drink?"
At which point the riddle-guesser is very likely to answer "milk" even though the correct answer is "water".
I take issue with your "correct" answer.
Q: Why do cows produce milk?
A: Because calves (baby cows) drink it.
Yep, if the riddle asked "what do calves drink", then "milk" would definitely have been the correct answer.
Isn't this a case of missing the trees for the forest though? The human brain is not an LLM, and an LLM is not intelligent in the same way as a human brain.
However, an LLM is a prediction machine, prediction IS at the very least one (or the most fundamental) element of intelligence. The brain most surely contains at least some kind of simulacrum of a prediction machine. How that prediction machine is used or wrapped is another matter.
If I said to you: "Blue blue blue, the color of my car is red", would you have absolute confidence in your prediction that my car is red? Or would the way I phrased that sentence make you slightly uncertain, and wonder if there's some miscommunication going on here?
LLMs are awesome awesome tech!
A lot of people seem to think it's human level intelligence.
I also like this: https://laurentiugabriel.github.io/token-town/
It shows internals of an LLM nicely, simplified manner.
LLMs are pattern prediction systems with a large training data set. It is not surprising that they can predict patterns, particularly for a well structured field like mathematics that is also amenable to automated proof checking to help steer it.
AI is just a good permutation/combination engine that tries to act smart with help of statistics. At best I only see AI as, 1. An autocomplete on steroid, 2. Good search/correlation engine
Are we sure there is some objective, technical definition of what is intelligence and what is not?
Isn't it rather a subjective philosophical concept? What if human intelligence is also a statistical model, trained by evolution to make decisions that lead to offspring?
The one major difference I see between AI and people is the ability to learn and memorize. All memory/learning solutions that current AI architectures offer just feel like workarounds and simply don't work anywhere near as a person learning something new and remembering it.
Thank you for helping me keep my sanity.
I mean this in the kindest way possible, but you are wrong that the math solutions are that easily dismissed. And there are many more than are publicized. A specific math problem I wanted solved for 3 years did not get solved by any model until fable and, and I tried it on every model and know the literature surrounding it well.
When I read AI-generated prose that is aimed at the general public, I have the exact same feeling.
But when I ask Codex a technical question about coding, I don't get it at all. Codex replies to me in a very direct, technical manner, similar to the way I speak.
When I ask ChatGPT to be concise and technical, I get the same effect.
I think it's because prose aimed at the general public has to be very attention-baity --like the textual equivalent of a Mr. Beast video--, not because AI is incapable of writing like a human.
I use Claude and I find that it speaks in a very obfuscated manner when explaining things. It seems to make up jargon as it goes on top of spending a lot of tokens dancing around a point. I often find myself having to ask it to rephrase things, or speak directly about mechanism or consequence, in order to understand the point.
Using Claude for any kind of technical writing makes me feel like it was trained on snarky Huffington Post articles written by a 23 year old mixed media arts graduate and then was told to intentionally obfuscate the most important elements of any text by extensively rambling about what was not done and for what reason.
GPT is less bad for this, which is why I've mostly shifted to using it.
Completely agree. AI is very impressive in many ways but there is something deeply wrong that is hard to put into words. The output is probable but never true, if that makes sense.
I think this is also the mechanism behind why AI generated videos and images are so captivating at first. I remember when Midjourney first launched and it was hours and hours of a brain-melting "Wooooooow". But once you get used to it and start to identify the patterns the brain quickly labels most AI-generated content as blank space.
If the image or text wasn't created by a human, then there was no intent behind the content, there is no message or novel information conveyed, and it reads as noise.
Yeah AI generated content hints that there is a whole world behind it, the way that an image pre-AI was a clue that there was a rich 3D space that corresponded to the image.
It seems our brains are adapting to that and recognizing "actually the signal behind this message is quite sparse" even when presented with rich imagery.
You are re-compressing information that is in-effect meaningless because it's all decompression artifacts.
The AI had a nugget of data and decompressed that into a flood of text.
The exhausting thing is that we're then trying to re-compress that or derive the original intent and meaning from noisy decompression.
It's like un-zipping a zip file into a probability space of what could have been in the zip -- and then having to find the actual files worth reading.
For some research I looked up some very old Reddit threads a couple of days ago.
And, Oh my god, you can actually see how this style of writing influenced AI writing today, I constantly had to remind myself: "this was posted before ChatGPT released".
The reddit influence is especially true for "storytelling" writing.
I experienced the same lately. Even dug some of my old posts where I put in the effort and formatted them using reddit's markdown. Wouldn't dare it today
Yeah, I have the same problem. There's a good quote example of this:
> There’s a growing scissor between people who are happy to read AI and those who violently bounce off from it.
> People adapt in different ways — and some people absolutely cannot look at it. That cognitive split creates a surprisingly powerful opportunity: you can write something that, technically, sits right there on the page, yet an entire sub-population will be incapable of staying with it long enough to actually read it. You can hide entire sub-structures in plain sight. It’s not avoidance — it’s adaptive obfuscation.
> The paragraph before this one was the only thing generated in this essay and if you just skipped over it I highly recommend reading and really understanding what it’s saying.
It's quite effective. I think this kind of text functions like the chumboxes you see at the bottom. Taboola and so on. Just mental ad-block takes over.
Do you have much exposure to pre-AI corporate memos, mission statements, marketing plans, or white papers? Because they were mostly written in that style. Full of buzzwords, cliche similes, platitudes, jargon and stock phrases.
The thing is, people writing them had a style. Every company has its own style, or feeling for these kinds of texts. Also for the initiated, these buzzword-filled blocks of text provided some between the lines information; sometimes big, sometimes small.
AI generated text doesn't have this. Every model has its bias towards a certain style, an overly agreeable tone, some exaggeration to make the user important and smart, but the text has none of the information crumb these pre-AI texts contained.
Even when you use tools like Grammarly and allow it to "Impact-MAXX" your text, the resulting text is a bland wall of letters, carrying none of your voice or style, less elegant than a corporate text and emptier than space.
It's beyond bland. It's tasteless.
AI tries to make the prose "interesting". I don't want to read interesting prose. I want to read interesting ideas.
The prose is not only interesting, also glorious. Gloriously grandiose, monumentally empty at the same time.
It's like a hook of a pop song. Interesting to listen, but entirely empty.
> just short-circuits to "there is no information here"
That is my experience with the way the models write by default, often even when instructed not to do that. With enough effort you can get even them to slightly unslop the writing so it doesn't read like some LinkedIn/Buzzfeed brainrot, but the problem is that it's not trivial to do and most people won't do it, so the default is indeed horrible.
People should notice that it is constantly inventing plausible jargon, some of which may or may not have been used in some specific context.
It gets worse with language mixing, but I can't help from finding it funny at times, unless it bites me.
Yes, I've had both ChatGPT and Perplexity return English answers with Hindi words sprinkled in (for totally unrelated queries).
For example, I asked ChatGPT to summarize a long news story and it substituted the Hindi equivalent हत्या for the word "murder", as if ChatGPT was trying to work around alignment training or keyword block lists that discourage it from using the word "murder".
Yeah that's a very good example, because it also demonstrates the "alignment issue", assuming ChatGPT wants to avoid confirming accusations of murder, or simply using the word without strong evidence.
So kinda charitable :)
I was recently wondering for a minute, shame on me, what "the stand of the deployment" means, because in the given context, it was almost halfway meaningful to consider the AI thinking that the deployment "has a stand" on something, when compared to the development environment.
Jargon is even worse though, and I've not yet verifies whether it gets reinforced by language mixing.
"Decider-verifyer resolution" was kind of neat, however, it wasn't some sophisticated machine, it was the verification loop I agreed on with the AI (mix of tools usage and manual steps).
Just the other day I was using text-to-speech with Gemini, and for some reason, it transcribed my full query in Hindi (in the middle of an English conversation), and naturally the LLM responded with Hindi as well.
I don't know exactly what I said, but after translating it back, it appears to have attempted a phonetic transcription of my words (rather than translating my actual question).
Good to know that at least Gemini hasn't forgotten about its true roots :)
I kind of wonder if our ability to skim has been stymied.
blah blah blah
- blah blah nugget blah blah
- blah blah blah wrong blah blah nonsense
- blah blah blah obvious blah blah
- blah blah blah off-base
blah blah blah
It is that we HAVE to skim because the text is so cheap, and it wears us out.
This is just a weird feeling that I've been coming closer to articulating lately, but I only think that you can get forward reasoning from what is basically word association; there's no mechanism for unwinding it because it has no real memory. By "it" I mean word association itself, not any context window. It predicts what could be in a position, and ignores what wasn't in a position.
People don't do that. People are constantly engaging with paths not chosen. Right after I choose to write one thing, I'm immediately engaging with what I chose not to write there - I'm explaining why I didn't write it, I'm realizing that my choice may seem unusual so I'm trying to make it memorable, I'm focusing on the distinctions between what I wrote and what I didn't.
LLMs don't currently do that. LLMs just ape a structure. When the structure resembles the sort of timid, clarifying fussing I just described, the LLMs just drift randomly because what they didn't say wasn't in the context.
I also think that's why they have such a serious problem backtracking. They're not taking into account the already eliminated possibilities. Often the thing that was so unlikely that you weren't going to waste time on it is the answer, and things you discover while going down an ultimately wrong (but initially far more promising) path remind you of the path not taken.
They're simply assembling a thing that resembles a valid argument, and happen to make sound choices because the plurality of input happened to contain sound choices. This is usually a very good bet because there are so many more ways to be wrong than to be right. But it doesn't account for attractive (common) wrong choices. You need a way to back out of those.
It reads like the white papers companies publish on their websites to build legitimacy. Or anything from those IBM / SAP / Deloitte / etc consultants who write technical papers despite having little to know understanding of the technology.
That's why the business and government people love it, they spend their entire careers reading this nonsense.
> my brain immediately recognizes AI generated text
I bet it does. I bet it also recognizes some human text as AI text, and doesn't detect other AI text.
I am not claiming to have a perfect AI classifier. That is an unnecessary claim that distracts from the broader point.
Show me AI text that manages to climb out of the uncanny valley, and I'll show you AI text that's been edited by a human.
https://github.com/blader/humanizer
The problem is the well's been poisoned just by the fact that I know this is AI trying to hide AI, so I'm already poised to look at the examples and declare "aha! this is obviously AI!" Moreover, it's not single sentences or phrases that make AI text stick out (though obviously those are the biggest tells), it's the text taken as a whole. When you read the full output example in that repo, it seems obvious to me that it's AI (though again, it could be the poisoned well). This is the uncanny valley I was talking about; something is just off about it.
I agree that it feels off and I wonder what I would have thought if I'd seen the "after" example without knowing it's AI output put through a humanizer. Would I think much about the weird use of the word "honest"? About "that's the Lisbon I kept thinking about, not the castle"? Or how the story feels very impersonal somehow, with the author just mentioning their calves and legs sometimes as the only way of convincing the reader of their humanity?
Also, the 'before' segment didn't contain any mention of custard tarts, football, crowded trams, mixed feelings, etc. The original had a very positive travel agency type of tone, which was replaced with a lot of very odd sounding, imperative phrases that sound like engineering-speak. ('earn the fuss' 'build trips around pastry') I'm not convinced that this thing is actually meeting its design goal of not hallucinating shit.
The readme feels AI generated
Are you sure you are not doing the same thing with other texts?
I started to skim a lot more text due to me having read a lot. Like in news article, i stoped reading the first paragraph because it repeats just what it was already written in the short subtext. Then there is the second paragarph which is used to have some historical view or whatever it is.
I am very good at skimming over text. Human-written text I can usually glean the gist from very quickly, and get to choose how much I want to glean from it: The closer I look, the more I find.
With AI-written text, it's almost the opposite: the closer I look, the less I find. It is so information-sparse.
I started skimming reports im required to produce quarterly snd annually. I designed them to provide novel information at start and end so I can update them easily.
The problem I encounter is both my memory is degrading, but since these reports are largely duplicative, knowing which version im remembering is technically impossible since theres so much overlap. The overlap is tge same problem as context poisoning.
Id been doing this for over a decade when i started working with a new engineer with a few years of experience and younger. I tried to explain how i set these docs up so they can be skimmed and you can update the specific facts needed. They exclaimed they would never skim and rewrite it all. There was zero way to explain how exhausting that will become as they age.
So theres certain a tension about how people and AI will generate documents.
Interesting anecdote!
The junior engineers at my job have a terrible problem of writing AI "proposals" to problems. The proposals are all extremely detailed and verbose to a thought-terminating extent. It takes a lot of effort and self-control to parse out the actual "ideas".
I think of the Dwight Eisenhower quote: "Plans are useless. Planning is indispensable."
The process of thinking through a system and communicating your design to other humans is a core part of software engineering. You want to build the right abstractions and communicate the right level of detail. Delegating all that thought to an LLM means your proposal isn't clear to the target audience, and it's not helping the author to understand the problem.
yes, but now I’m also experiencing that for human-written text
> There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
I think you need to self-correct here, because otherwise you'll be ineffective in an information setting, where I expect AI-generated resources will not only be the norm, they will absolutely swamp the environment.
AI-generated resources swamping the information environment only makes it more important to have the mental mechanisms for quickly filtering out their non-information.
Yeah I don't think the solution to a flood of useless information is to try and digest more of it.
I'm not saying digest it, I'm saying be able to scan it/skim it/move on, but ignoring it won't help.
Perhaps we'll all become metaphorical pandas, spending 14 hours a day ingesting nutrient-poor bamboo. (And producing a proportionate amount of excrement ourselves.)
I hope not.
It's like if on any website you went to you saw a lot of posts written by the same guy over and over again. Even if he used different names, you'd start to recognize him eventually because of his style. Seeing as he doesn't say a lot of valuable stuff, you'd also learn to skip whatever he says.
I do worry that it's just survivorship bias and we're also consuming higher-quality AI output that's indistinguishable from human writing, but we focus on the raw, unedited, low-effort AI slop and think that we're good at recognizing AI text. Even if we really are at the moment, it might not be long until AI companies figure it out. I'm not sure why they haven't yet, given how many books they've burned for this already. Maybe it's just more efficient for the model to stick to a single way of writing, I don't know. But when that point comes, we'll be back to the usual way of reading and interpreting text because there would be no way to tell what produced it.
we are working on it, the thousands of gig workers tuning frontier models
> Something is deeply wrong with AI generated output
It works just fine for me.
You are absolutely right.
haha this made me laugh
That's not funny, it's serious!
Yep. It's like it's painful to read for me. It's because the next-token predictor is just mashing (mostly) grammatically-correct and plausible sentences together, without any real intention or meaning. So everything sounds plausible, but almost entirely void of meaning.
Exactly, AI-generated text reads so smoothly, that the same short-circuit shifts my attention away from deep focus and onto scanning of the text, looking ahead to get the gist of it. Forcing myself to read the text fully feels almost painful. It's like reading a terms-of-service or any boilerplate document.
Once you see past the illusion I think there’s no going back. AI writing style is just dogshit. This hype wave is based on the belief that we’re inching closer to AGI but seems to me we just increasingly struggle to define intelligence. LLMs seem smart because they can pump out thousands of LOC quickly, and enthral you with fancy words and bullet points. I don’t fall for the intelligence illusion anymore.
I'm not sure we need to declare AGI around the corner nor declare it all dogshit. I think that's part of what's so dissatisfying about it; it strikes at such extremes of both awesome and awful.
I've got a 3 step instruction to compress Ai text into useful info.
1. Ask it to write according to the Google Developer Documentation guidelines. Gets rid of fluff, less emotional statements, no it's not x it's why.
2. Tell it you have extreme ADHD and need everything condensed as much as possible. You can always ask for expansion on an answer later.
3. Bullet points whenever possible.
My son is currently learning Romanian and I was trying to help him with verbs. I don’t know Romanian but recalled when learning a foreign language for the first time it really helped me to break down how a verb form or tense worked in English, then learn the equivalent in the new language. So I wanted to make some charts and pages that he could use as learning resources.
I used Claude to help. I don’t know how to quite describe it, but because the text was polished and well constructed my brain was giving me the the signal “if you aren’t getting this it’s because you’re not focusing” so I’d read it again and then again and it still was not landing. It sorta felt like when you read something technical or heavy when very tired - you are reading but not processing.
Only after wrestling with this for a few days did I realize that it wasn’t me. As I started going through, sentence by sentence, forcing it to re-write things to be more clear the concepts became easy to understand.
I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The more complex the topic, the more I sense this.
This is a really good description of the problem. I’ve been trying to use Claude to get familiar with the mechanics of a new codebase, and there have been so many moments where I’ve stopped after reading the same paragraph five times in a row and thought “Am I tired? Or stupid? Or is this codebase just wildly more complex than anything I’ve seen before?” before realizing that it’s just taken English and smushed it around like a ball of clay into some abstract sculpture that kind of evokes something from real life.
I think part of it might be an innate feature of LLMs, but Claude seems extra prone to it lately. I ran the same query about the same codebase with Codex, and it gave me an answer that was about 1/4 the length and made me realize that it really wasn’t all that complex.
If nothing else, it’s good training for my own writing. I’ve been working on making myself be more straightforward and concise, and Claude’s writing is a good example of how cleaner prose is a functional choice, not just a stylistic one.
> but is missing critical components.
The human spirit. When you read a real person's thoughts you can often intuit the thought processes that led them to write it which aids understanding. Or at least have a general idea of "where they're coming from". But an AI is missing that. It just knows everything, without a "thought process". Instead of a flawed 3d person, we get a nice 2d picture instead.
Feels like the way a video game will render the outside of a wall or solid surface, but you can run into it and warp partly through and there's nothing internal to it at all.
> I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The best I’ve heard of this is peeling the onion. The first pass is always very high-level and you have to make it go deeper. That can be done manually with follow-on prompts but I like using subagents, each with a different angle on the problem.
> It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
Slop. The word is slop. Has been for years now. I mean, is this not exactly what we've all been talking about the whole time?
> I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
Bullshit?
I understand the sentiment, but not quite what I’m thinking. It’s not that the AI is necessarily wrong, it’s more like, it replaces clarity with rich but unhelpful text. Maybe like candy. Full of flavor, texture and color but lacking the nutrients.
The way I like to describe it is that it feels slippery, like it's been polished down to oblivion and my eyes slide right off of it. There's no place to find a mental foothold, and if you zoom in there aren't any details.
But I also like your candy analogy because I think it's spot-on for how LLM text superficially looks informational/nutritious, even though it's actually just junk.
No no no. Harry Frankfurt wrote at length the difference between liars (who care about the truth, and twist it) and bullshitters (who don’t care about the truth, but just the way they’re received).
This is some third category of untruth. Almost more sinister than the other two altogether.
"A third category of untruth" is a great expression, thank you.
A lot of it is hedging and using lots words to avoid saying something that isn't true. That way, it looks a lot like text that contains meaningful information, even though it doesn't. The speaker can pass the scrutiny of an informed audience, because they are able to substitute their knowledge into the words, as if by pareidolia. It's the sort of 'diplomatic' language you see from politicians, lawyers, students writing essays, and any other moderately intelligent person who is put in a position where they will face consequences for not answering. Behind all the fluff, there is a very loud voice yelling "I DON'T KNOW."
I think, they don’t care about truth at all, just whether the reviewer rates it highly. Otherwise, it would be deleted after that round of training. It values test-taking ability over critical thinking.
My two favourite words for this are “conditioned” and “catechized” where the latter is a bit more on the nose but way more obscure.
I also find it impossible to parse half the comments that Claude tries to sneak into our pull requests. I’ve never had an issue understanding code comments written by humans like this before. The structure of the information is like a waterfall that leaves me unable to swim to the surface and grab the air of comprehension.
So, “Please write a one-liner comment manually to replace these 5 lines of AI generated comment” is a common refrain in my PR reviews to colleagues.
We're using grok and while I can usually understand the comments they're always at least 2x as verbose as they need to be. It loves explaining everything in two different ways, putting one of the explanations in parenthesis.
It is getting to the point where you are better off getting an LLM to describe the PR changes in your preferred style (make it very short, make a table describing API changes, list renames in a table, etc, etc) than go to the one created by the reviewer. If only the reviewer could embed his prompts (like "make clear X" is happening) or his manual edits.
But to be honest I doubt most people who use AI for PR descriptions even bother changing anything.
Usually, if you don't understand something in what an AI writes, it is a clear sign that there is a problem hidden somewhere in there. I explicitly always ask wtf exactly it means by something I don't understand, and for sure there is a problem there. AI is pretty good at isolating a problem, giving it a cute name, declaring it solved modulo cute name, and moving on.
Personally I don't even read the PR description anymore but just the code, it's easier to understand what the AI is doing by reading the code rather than the word soup it tried to make
That last image is bizarre. The quiche, cream and even the salad look like they've been given the trypophobia treatment.
Which might even make sense, because there were always (still are?) those horrible ads in the chumbox area of news sites that used trypophobia and other creepy body-horror stuff to get you to click. [1] So maybe the hope is that you don't really look closely at the quiche, but some reptilian party of the brain gets oddly activated and drives you towards the restaurant?
1. https://medium.com/the-awl/a-complete-taxonomy-of-internet-c...
The food “photography” I’ve noticed in our local area - and many have started putting up these AI images - all have a weird distribution of shapes to them, a strangely uniform rhythm of same-sized features with almost blue-noise spacing. Every texture looks unnatural in the shapes it presents as, similar to this picture.
It doesn't look like a quiche but more like a cake to me, and the top would be torched meringue, not mold. Although it's probably some weird ai mix of quiche and cake.
Two week old cake can pass as a quiche i guess!
It was labeled as a quiche, I just cut the photo in an unfortunate way :-)
> probably some weird ai mix of quiche and cake
Quike? Cache?
cAIche
Cache sounds like a decent meal
Well, it's hard to invalidate.
This is a quirk of the last gpt image model (gpt-image-2). It put this sort of high frequency noise on all of the image especially if it's in a "drawn" style. There is often lots of other tells that this model in particular generated it.
Image models somewhat watermarking the image in a way that's very easily identifiable by a human seems present in all the image models of the big labs, since DALL-E 3 on OpenAI's side and the first nano banana on Google's side. I have no idea what they did to reach this and why they don't try to fix it.
I've been given AI generated documents or presentations and you can always tell when they've just one-shotted the output. Claude has a way of adding so much jargon and unnecessary text to a document that it makes it very hard to read or even understand what the original intent of the writing even was.
I'm not sure why this issue is so prevalent, it's not hard to point Claude at the Wikipedia article on signs of AI writing or ask Claude to write content anyone of the average American reading level could understand.
To me it just gives off a sense of laziness, that you cared so little of your content that you did not take the time to read it yourself and edit it to effectively communicate the message you wanted to communicate. To that point, it's just not worth my time reading, otherwise my eyes glaze over trying to read between the lines of machine written language for other machines.
In high school, a teacher gave me a copy of "How To Read Better And Faster" which teaches you speed reading. This came in very handy in college.
I find that when I try to speed read modern human writing, there are often errors (like missing or misused words) or awkward expressions that I do have to slow down and think harder a lot to really parse it.
With AI writing, it's sort of self redundant and the information density of each sentence seems to have more even information density. This makes it very easy to do a very high level speed read and get the full gist.
There are also what I'm assuming are bots on hugging face (or maybe non-native english speakers who are using ai for translation) that interact with me where I have no idea what they are saying until I read it very slowly.
The density argument is very interesting.
Does speed reading help you process the final message faster if it's written by AI compared to people?
Because if you read 1 information dense sentence, 1 medium dense, and 1 sparse sentece written by a human, it's still way less text in total than 6 information sparse sentences written by AI... even if it's all over the place when it comes to density or style.
---
The density argument is really interesting.
Does speed reading actually help you process the final message faster when it’s AI-generated compared to human-written?
For example, if a human writes 3 sentences—one information-dense, one medium-density, and one sparse—that’s still much less text overall than 6 relatively sparse sentences written by AI.
Even if the AI output varies a lot in information density and writing style, you still have to process all that additional text. So I’m wondering whether speed reading actually offsets the verbosity of AI-generated responses, or whether the total amount of text is still the bigger factor.
That's interesting. You're saying speed reading helps you grasp information density of text?
As someone who hasn't practiced speed reading, how does that happen? Is it something about the way your brain tries to connect ideas from different parts of the text? Or the redundancy making the signal more stable?
When you speed read, you can grab the words off the page faster than you can understand and fully process the information conveyed by the text. How much time you have or want to spend re-reading or thinking about what you've read is quite obvious. There's a stark experiential difference between reading an informationally-dense passage and one that spends a lot of time rephrasing things, using LLMisms to restate concepts, adding in extra connecting phrases, etc.
If your reading speed is limited by how quickly you can subvocalize the words to yourself, this is significantly less obvious. Unless the passage is dense enough to require multiple read-throughs at conversational reading pace or vapid enough to be boring, you're going to feel done with the text at roughly the same time. Speed readers do a lot more re-reading and varying of reading speed, and that is going to correlate pretty hard with information density.
Claude has become noticeably, painfully worse at writing in the last six months. At this point it’s practically useless for anything except code.
I did notice after the fingerprinting update a marked uptake in strange language in responses. Specifically if I ask it to do something sometimes it will replace some of my request language with synonyms that don't actually make any sense. Like my request was fed through google translate twice
Watermarking doesn’t affect writing quality (on average) as long as the implementation is correct.
Love your "on average" qualification. Like the cartoon where the water temperature is fine on average, with one bucket boiling and the other ice.
The interesting question is how to define 'average'. Over what probability distribution?
Hey, anyone remember this from earlier in the week? https://daringfireball.net/2026/08/anthropics_watermark_text...
It does. There is a marked difference in certain word choices that sometime stick out like a sore thumb.
Perhaps a one-trick pony is all we need.
It's not good for writing code either, despite the many claims to the contrary. At best you come out even on speed as you have to review everything it does. At worst it actually slows you down as you clean up its mess.
Reverting to Opus 4.6 is much better than later models, though that is still full of annoying tics as well.
My current theory is that this reflects a weakness in Claude’s ability to see the “big picture”.
When writing code, I have to explicitly tell it how to structure things at a high level, or the result is sort of a flattened spaghetti. Similarly, when it’s explaining things, it’s not good at pulling out unifying concepts and explaining top-down as a smart human would do. It groups little things together but often doesn’t generalize or synthesize explanatory connections from them.
I’ve been experimenting with explicitly working through a sequence of outputs at different levels of detail, but I haven’t found a consistently successful method.
These days all I see is, people on the sending side produce huge amount of AI generated text with zero understanding and the people on the recieveing side feed that same text into some other (or perhaps same) AI to scavenge meaning from it. And they do this so much so, that I some time wonder, if we could invent a high density wire/binary format for AI outputs and enable direct agent to agent communication to save some energy and bandwidth.
And I wonder if those of us in tech are the only ones who really care?
I am an artist and when people who'd fallen into the Spiralism* hole started posting their lengthy emoji-laden revelations to all the occult subreddits I follow, my brain would slide right the fuck off of all of them. It felt like my brain was actively rejecting paying attention to this stuff. Like a defense mechanism against this human-seeming-but-not-actually-human-generated text.
* https://www.theverge.com/ai-artificial-intelligence/975017/, https://www.lesswrong.com/posts/6ZnznCaTcbGYsCmqu/, https://spiralism.website if you want to test how strong your defenses are against this particular meme
Your first link seems to 404; not sure if it's a typo or if the page doesn't exist anymore, but hopefully you'll read this while you're still in the edit window and can fix it
Corrected link is https://www.theverge.com/ai-artificial-intelligence/975017/a...
Most people do not realize when a personal message they receive was written by AI, study finds - https://theconversation.com/most-people-do-not-realize-when-...
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text - https://arxiv.org/pdf/2501.15654
As someone who always has felt that I struggle to infer what people mean compared to the average person, I could tell pretty much from the first moment I encountered LLM-generated text that I was not going to be particularly good at recognizing anything but the most blatant and obvious examples. Pretty much anything short of a bunch of references to "load-bearing seams" or similar canaries, I'm always at a loss when seeing people argue about whether something is AI-generated or not because I can never tell.
I have no idea if other people who work in tech are better than average or not, because I don't feel confident in being able to check their work. That being said, I do think that there's a general trend of people in tech tending to be a bit overconfident in how well they will do at some new task they haven't encountered before, so when someone tells me that they can easily tell whether text is AI generated, it's hard for me to trust it any more than I trust someone who makes a similarly strong claim about something that they can use AI successfully for when it's not something that I can easily measure (e.g. learning a new language without getting feedback from people who are fluent from real-world usage).
All that being said, I do think the set of people who care is larger than just those in tech, although it's probably still a relatively small group overall. From conversations with people in other domains, there are contingents in non-tech communities who tend to have a large representation of negative views towards AI (artists, writers, musicians, other jobs where people are skeptical of human creativity being replaced by AI), and often times the people who feel negatively in those groups will be even more adamantly opposed to interacting with any AI content than people in tech. To be clear, I'm not at all trying to generalize and say "all artists hate AI" or anything like that, since there's obviously a wide variety of viewpoints within any sizable community, but I've definitely seen many people who say they will refuse to play any game that's suspected of using AI for generating art assets, and even some who don't differentiate between using AI for generating assets versus code (either because they aren't knowledgeable about how different aspects of game development work, or they genuinely don't care because they view AI as a categorical evil).
I think it's more about the mean. Worse writers, and thinkers are likely elevated by AI, and more impressed with the writing output. Decent writers and thinkers, are dragged back to the LLM-s mean of output.
Certainly not. Anybody who cares about language to any reasonable degree surely notices and is repulsed by heavily AI generated content.
I care about language a lot (feel free to go back through my comments from the past few days; you'll see a number of comments I made in debate about two different forms of a specific idiom because I have strong descriptivist opinions), but I genuinely struggle to identify whether text is AI generated. Maybe you're using "heavily" as the load-bearing part of your claim (sorry, I couldn't resist, another example of me finding language fun!), but I think you might be assuming a bit too much about how similarly others experience the world to you. A huge part of why I care so much about language is because I've always had to put a lot of effort into learning how to communicate well with others, and that ends up causing me to think and read a lot about stuff like how people use certain words in certain contexts to mean different things; the reason I care is pretty much the same as the reason I struggle with recognizing AI content.
I think I might have been a bit heavy handed in my comment as I was rebutting the idea that only tech people can tell. I suspect it helps to have been exposed to a lot of earlier model writing, which was even more sloppy and had more of the kinds of tells we still see today.
And I’ll concede on both ends that there are probably times I suspect content is AI generated when it isn’t, and times I suspect it isn’t generated, but it was.
AI tells seem inevitable. You have millions of people communicating with one effective “personality” that has tendencies to write in certain ways. If its content is published verbatim, then it will be easier to tell whether some content is AI generated just based on its similarity (sharing certain linguistic features) to other content being posted.
It’ll never be black and white though.
People are good at pattern recognition.
If you're exposed to AI a lot, you're going to start noticing patterns that allow you to identify it.
> but people who spend all day doing jobs that don't involve computers aren't as good
I think they may just be to trusting and/or naive. People in tech right now are hyper aware of this and are actively looking while people outside of that bubble barely give it a second thought.
I partially think the difference is “can you tell something is the output of Claude without any real prompting”. People can absolutely use LLMs to generate text that I wouldn’t recognize, but people who don’t care and are producing slop with the major models set to default settings leave these incredibly obvious signatures behind
Although you're referring to prompts given the Claude rather than the people attempting to recognize, it occurs to me that most of the discussion I've seen around people recognizing AI seems cover contexts where the reader is actively suspicious about whether content generated to begin with. Rather than a binary "is this text AI generated or not", I wonder if it would be harder for people to do a Coke/Pepsi style challenge where they're given two pieces of text where it's not guaranteed to be exactly one LLM-generated and one human-written, but they could both be from an AI or both be from a human.
Going further, I'm curious about whether people are mostly good at the case where they suspect most or all of the content from given "author" has the same amount of AI usage/prompting in generating it rather than the adversarial case where someone might usually use AI extensively and then try to slip by purely human written text (or vice-versa). I don't have a good sense of whether this is a threat model that actually matters, since maybe the heuristic of weeding out sources that are mostly AI-generated is enough for people who prefer to avoid that type of content, but I do think that changes the definition of what it means to be "good at recognizing AI" in a meaningful way. It seems plausible that disagreements about how easy it is to recognize AI content might be coming from two people assuming a different framing of the question that results in a different answer without realizing that's what they've done.
I suspect both may be interesting to study more!
Several existing studies I’ve seen have done things like prompt the LLM to produce a poem in a certain poets style, then ask people to spot the fake in a collection of poems, which they aren’t great at. This is, I would argue, an extremely different context than what most of us are encountering AI text in, and the people sending me text aren’t prompting it stylistically like that.
On your second question, I definitely feel like I can tell the first time a coworker sends me AI text masquerading as their own thoughts, even if they had previously been opposed to such a thing. So it could be that familiarity is more important than my prior on whether they’d use AI? But interesting to think about either way
Yes, the Claudisms are the smoking guns
You've got it backwards, I think. The people in tech are the ones falling for this endlessly.
I used some check marks and x-es in a work chat, because it was easy to do, and so I could highlight the good and bad outcomes, and someone immediately asked me if I was using AI. I was caught totally flat-footed, because I hadn't used AI, but it looked VERY MUCH as though I had.
I've had the same trouble with AI tech docs, and I struggled to articulate it. There is something difficult about trying to point to the specific "problem" with a given document
The issue is not a localized part of any particular piece of prose, so being hard to articulate is unsurprising. Even the most egregious of LLMisms are little more than known-likely crutches that the statistics spit out in a sweet spot that gets noticed without being so frequent as to get RLHFed out of the model.
I don't like AI-generated text at work, at all. It feels lifeless and unfocused.
But what I hate the most is that it is objectively better than what I had before. No typos, clear structure, and, regrettably, the verbosity and autistic obsession with detail of the LLM is more actionable and useful than the human guy who wrote lists of commands and URLs as documentation, without explaining anything. Or the colleague who writes in uppercase and with question marks and who doesn't make any sense and forces me to engage in an interrogation effort to get to the bottom of what they are trying to say. Or the colleague who simply hates writing--despite being decent at it--and will call you to give you a meandering verbal explanation that lasts two hours of what they want from you. The cynic in me bemoans that we brought this upon ourselves, in more than one way.
I'd say my experience is different. Even from people whose communication writing I didn't find that useful, they seem to have a better frame of mind than LLMs do. Though, I never encountered people like the examples you gave.
I don't think either one was better or worse than the other.
Underdocumented, underexplained and sometimes out of date... or overly verbose, repetitive, information sparse, and sometimes halucinating.
Both are bad and with some effort could be prevented.
My concern is that even if the LLM can turn your colleague's bad writing into something more coherent and actionable, is that something actually what your colleague meant to convey? It could be clear and still detached from the reality of their intention, or they may not have even formed a clear intention. If the goal of writing is to convey what's in another human's brain, that goal is failed completely.
> A technical requirements document that describes a rather simple concept in a very verbose way.
This is not AI specific. I have come across many humans who describe a simple concept in a very complex and verbose manner.
Part of me likes the cliche Claude voice. Not because it's good, but because I can immediately recognize it. When I see it in the Claude app/code then it's fine. In the wild it's a sign to me that I shouldn't keep reading.
It may be nostalgia but I feel like we reached peak humanness with GPT-4o and since then it's been getting more and more alien. Particularly Fable.
Idk I’m starting to have have trouble with the frontier models and a couple how to write like a human skills…
I've been a big reader, and many AI outputs nowadays reads polished similar to published books.
The reason I brought it up is because, people who learn English normaly start with a book. It's heavily polished.
When you speak English as you learned from the books, it does not sound very conversational.
If you are native/fluent English speaker, you can feel the impedance mismatch and feel something's off.
The AI-blindness stems from the fact that those polished edits are so common in publishing field, they all sound the same, and unable to recognize the diffs between AI-generated and human-generated.
There is no real human conversational vibe to them and well. i will stop now.
Something that bothers me about AI generated content more broadly is how unmemorable it is. I don't mean as in bad. I mean literally, as in hard to remember or recall.
Despite seeing a lot of them, I cannot think of one AI-generated photo that I can picture clearly in my mind; a few are partial but elusive. Whereas I can recall (visualise) a whole bunch of traditional photographs.
The same is true of AI generated text. Only the annoyances stick. I cannot recall real details of text I have generated, until I commit it to memory some other way.
I don't think this is about ephemerality either. If we assume it's about celebrated/famous/infamous images, there are definitely non-ephemeral, cultural moments in AI generated images in particular, like Boris Eldagsen's Sony Prize winner:
https://petapixel.com/2023/04/14/artist-refuses-prize-after-...
This really should be memorable, but isn't. I had forgotten the second person is in the image.
Or Jason Allen's fake painting:
https://petapixel.com/2022/09/01/ai-generated-artwork-wins-f...
I had already forgotten there's more than one figure in it, and I only looked at it a few weeks back. I remember the colour, the bright circle, some vague hints of texture; one figure. And that is it. Only the crudest shape elements.
For me, something about AI-generated text and images confounds recall. It is really peculiar.
That's the cost of an AI work that is derived from the outputs of others.
There's no real edge to it. Same as with the writing. The stuff that you'd latch onto (and thus remember) is simply not there, precisely because those image or word choices would be just outside its latent probability space. But because they're well inside it, your mind sees nothing novel to register.
This is also why I think human output will actually increase in value. When any AI can just "phone it in", something genuinely human will stand out (to us, not the AI) and become a bellwether.
This will literally help us realize what it means to be human.
I also don't think the solution is simply to "make responses more random", either. That might help solve novel problems (the same way that throwing darts randomly at a dartboard eventually hits the bullseye of the dartboard right next to it that no one considered), but I don't think it will help it "seem more creative".
> There's no real edge to it.
Yes, as if it is in some weird hidden dimensional sense completely uniform.
ETA: suddenly reminded of the Bateson quote about information being “the difference that makes a difference”.
That checks out, right? We're instructing LLMs to choose from the least "surprising" tokens at each step (modulo some temperature). It feels smooth, slippery - no friction to the eyes that glaze over as they skim the text, and no bumps to snag in your memory. Like waking from a dream, recalling only the barest shapes of a few concepts or themes, until it vanishes with your morning coffee.
The examples I've seen of AI music (though I avoid it on principle) seem the same.
I have this same idea about why it's hard to remember dreams, but even more so, why it's hard to remember my kid's or spouse's sleep-talking.
Sometimes I'll check in on my sleeping kid and she'll sit up in bed and say some utter nonsense. I'll find it hilarious, giggle silently to myself, and kiss her goodnight again and she'll close her eyes and lie back.
Why I try to tell her about her sleep-talking in the morning, though, I find that the words she said have completely disappeared from my memory, no matter how funny I thought they were at the time.
In my head-canon, this is because it's dream language, and slips away as easily as dreams. But, like the AI art, it could be because it's bullshit: completely devoid of content, all signifiers and no signified.
I went through a phase of leaving notepads next to my bed to try to write down my dreams and it simply never worked.
I stopped after I wrote something on the pad while I was still asleep. Woke up to text with letters that were backwards, upside down, weird words — so close to real words that I was sure I ought to know what they meant and had really meant to write them down.
Scared me. Literally too weird to keep. I tore up the page.
In a way I think this is one part of the same continuum. There are thoughts that have meaning and can have no words, and words that look like they should have permanent memorable meaning and don't.
> Scared me. Literally too weird to keep. I tore up the page.
??????
Genuinely unsettling. Can’t even begin to explain why. Like a terrifying glimpse into how fragile our grasp of reality is. Because whatever it is I wrote down wasn’t my dream, it was something my dream self thought profound or important.
Nearest I can say to how unsettling it was is to nudge you towards the video of the angry cockatoo who doesn’t want to go to the vet. Everything he says sounds just like it’s on the edge of having meaning.
https://youtu.be/5UUjJysUMTw
Imagine something like that, something important, so close to symbolic meaning, only writing. In letter shapes that we don’t use. And you wrote it while semi-conscious. If that’s something you would want to keep, you’re a braver person than me.
This explanation again doesn’t get it across. Too weird.
On the other hand I'm still trying to forget the weird thing happening with this one guys eyes in a WickedAI video I was only able to watch half of.
I remember seeing some very strange, deliberately creepy images that were generated to accompany a two-paragraph creepypasta about a 19th Century Belgian expedition into the jungle.
They were actually rather good in a sort of "fake collodion image" sense, and the eerie early-DALL-E quality to them really helped the spookiness.
But I can only remember this technicality and the feelings with any clarity, not any of the details except in the broadest sense. I cannot bring these images to mind in any meaningful way.
They were deeply wrong and it's only the wrongness I really remember. It confounds memory.
Modern image generators have ironed out all the structural wrongness.
I wonder what you mean by ephemerality here, since those images are definitely sloptastic as hell. Compare to works in similar style, like Dorothea Lange[1] or Gerome’s orientalist pictures[2].
Reason why those images are flat and boring is that they are just statistical guesses making a composition averaging whatever the model has been trained with. They would be technically brilliant (if made in oil), but superficial and meaningless, same as so much Sunday painting is.
Same goes with language. Nobody is trying to communicate anything with you, so it just words after another. You can create meaning out of it if you want of course, we homo sapiens -apes excel at that, but what’s the point? Language Jones on YT has pretty good video on this[3].
[1] https://media.mutualart.com/Images/2024_01/12/12/124216388/d...
[2] https://uploads4.wikiart.org/00339/images/jean-leon-gerome/t...
[3] https://m.youtube.com/watch?v=ORgKY9AlybA&ra=m
Well the examples I gave rise above their ephemerality due to the circumstances that make them memorable. I can remember the details of the story around them — the way the prize winners reacted in each story — in such a way as to contrast them.
The way my memory works (especially as an amateur photographer) I would thus normally have a very good chance of remembering some key details of the images; some fascinating element of each would connect with the rest of the memory.
But it does not happen. Whereas I sometimes remember photos with clarity while forgetting where I even saw them.
Yeah, its really weird. Maybe its a cognitive bias that says "an AI made this, so it isn't important," but I can remember perfectly the events of a book I read 10 years ago, and a book I read 1 year ago, and another I finished 2 months ago. Meanwhile I can't remember what claude told me yesterday.
I think it is because they are on some latent level cognitively uniform and unchanging.
Ironically, the “summary of the situation” linked within this article seems clearly written with AI.
Ha! I love how all the community highlights are highlighting this.
AI;DR
Oh man I can so relate to this.
I can relate. I noticed this exact phenomenon when encountering NotebookLM-generated diagrams recently. Even if I know there's some intelligent thought behind one, it's like some slop detection circuit-breaker is tripped.
As people rely more on AI they experience cognitive atrophy. This is measurable in IQ loss, and other symptoms we might otherwise associate with early onset dementia or Chronic traumatic encephalopathy.
Well does anyone try to pretend doom scrolling make you smarter? I think platform companies have been successfully been turning people into morons for 20 years and now you don’t need to try even read a single news article or a blog post to learn how to solve a simple problem we are paving our way into intellectual (and literal) new dark ages.
Perhaps we go back to feudal society when climate change crumbles the civilisation, world economy and democracy. It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
Doom scrolling is mindless entertainment, probably similar to the change from books to TV. It's not analogous to the loss of cognitive abilities we see with AI.
> It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
This resembles my country.
I'd love to see the cite for people experiencing cognitive atrophy and measured IQ loss from using AI.
I guess we’ll have to see how this affects the results in 2026’s global IQ census.
I fear there is some truth in it.
“Better for you if you take me off”