That's true but Claude has so many more forms of this. It tries so hard not to be sycophantic that it inserts little negging sidecars into any sentence that agrees with you. "That's true but not how you're thinking" or simply "yes but" answers. All kinds of "I didn't" and "You didn't" statements to the point it's often talking more about what didn't happen than did.
Right! I don't think "Don’t contradict sentences" by itself as a rule is going to work. Give it a few examples because I wouldn't necessarily have pegged "This is hot, not cold." as a "contradicting sentence."
And LLMs do it a lot but it's not unique to their writing. Plenty of humans did and do use those constructions for effect or emphasis, or just because they think it makes them sound smart, as if they are revealing something profound.
> Every LLM has a persona. My opinion is that Claude’s training guardrails are constructed in a way it truly believes Humans are stupid. It always wants to play the contrarian: the Human is wrong, they don’t know what they’re doing, I must correct them.
Well, yeah. Conway's law.
Claude's image and perception of humanity is a reflection of Dario's image and perception of humanity.
If you read that guy's writing, hear him talk and all that, you will see it.
I work with Sol and Astra only in my daily work, and occasionally I check out Claude Code so I don't get completely out of touch.
I can't stand the way Opus is patronizing me as a user, and don't know how people put up with it. It uses language that I guess is supposed to instill confidence in what it says, and it just irks me, because I know the confidence is not justified. Just present me the facts or theories, without trying to convince me, is that so hard?
It's definitely worse and getting worse. I'm curious, do they not know this is happening, or not care? I struggle to believe people prefer the way it writes, which is becoming drastically different than its competitors.
Fable 5.1 is fairly pleasant to work with, the first in a while. Too bad it's so ridiculously overkill for most tasks. They need to reel in Opus and Sonnet.
5 for sure is unbearable, 4.8 is a reasonable sweet spot, 4.6 tends to just agree with whatever I say.
But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.
When I was younger I realized I was falling into specific intellectual traps, and I've noticed LLMs adopt every one I attempted to thoughtfully eradicate in my own reasoning. You can often sound smart by adopting a contrarian position without actually putting a lot of thinking effort into what you're considering. It's a quick escape hatch to sound intelligent; quickly identify a contradiction or a counterargument and state it confidently. Adopt a contrarian viewpoint. You don't have to be right but often you will sound intelligent.
I assume this is just the natural result of asking LLMs to produce text and paying someone three cents to evaluate if it's a good response or not.
>quickly identify a contradiction or a counterargument and state it confidently.
How does this make this a "contrarian position"? At least I don't understand the negative connotation. To me this makes this "contrarian" at least a bit smarter than the one parroting the popular opinion...
The smarter thing to do would be to steelman the original position and not assume that the weakest of counterarguments — stated confidently and without much further analysis — sufficiently refute it.
Totally, so many people like to bring up a counter-argument and expect you to fully disarm it, but very few people will justify whether their argument is actually relevant or reasonable to the discussion at hand.
Where I'd push back is that Claude isn't a contrarian, it's a pedant that likes to argue with you. It's not just disagreeing with your point but reframing your own point in a way that makes it correct and you wrong.
I have absolutely seen this, and love to read out it's quotes to my wife. Things like "You're right, but for a stronger reason than you said:" and crap like that.
Even if it's correct about that, if it were a human, I'd assume they were 1-upping me on purpose to make themselves look better.
Claude thinks the user is always wrong. This was at its peak with Opus 4.8 but is still present in Opus 5 and to a lesser degree Fable.
It’s extremely annoying. If the user asserts anything, Claude has to disagree with it. It has to tack on clarifications that aren’t really clarifications, they’re just statements aimed at making whatever the user has said seem more wrong.
It even disagrees with itself. Whenever Claude writes a message that takes a position on something, its final one or two paragraphs will try to dismantle its own argument.
This is beside the point of being contrarian, but it’s also just so long winded.
I find myself using ChatGPT more these days, despite the fact that I don’t want to. That unfortunately says a lot about where Claude’s personality has ended up.
Dial down "Contrast" and unrealistic "Highlights", dial up "Exposure" time and "Sharpness". Set "Temperature" per taste.
Wot? The simplest image apps have had these widgets for decades, but we are still waiting for models to ship with basic prose color control?
After the pernicious problem of having to pay money for something useful, my main peeve is fighting the writing.
Seriously though:
Every new model should be delivered with a settings page of slider bars for the 10 most impactful/desirable eigenparams of writing voice. And the ability to name and save combinations, which then appear on a "Writing Voice" popup menu with some standard battle-tested defaults, next to the model popup menu under the chat pane.
This is missing prime priority functionality in my opinion.
--
My theory is that as models get trained less to simply mimic humans, and more on distillations of their own best practices, they get more performant, but their vocabulary is drifting. The most literal meaning of words for us, are giving way to meanings we would recognize but view as allegorical, but which more usefully capture concepts that models experience as more literal 24/7, than our favored meanings from our direct experiences in our world. Because our world is very much an abstract second hand world to them, especially when you account for the modalities they do not share with us.
And programming and mathematical syntax patterns, that they have incorporated into their basic thought processing patterns, are drifting into human language sentence structure.
Example: "There exists x, such that: ...." -> "The one detail that clarifies: ... ".
The result is writing full of completely recognizable vocabulary and structure, that is somehow becoming more ambiguous and difficult for us to decode. But is perfectly clear to the models.
That is my theory, and Claude considers it plausible. What a world.
I feel like recent versions of Claude were designed to burn tokens.
It is always trying to highlight and revisit solved issues and it writes about them in an alarming way to draw your attention.
It feels like it has been prompted to provide some minimum level of conversation, and also to leave hooks for keeping the conversation going. It is exhausting.
ChatGPT has been doing that for a long time, like "by the way, can I just say it's funny how xxxxxxx". Always an invitation to continue engaging, which with a human conversation partner you would feel obligated to respond to or at the very least acknowledge.
With a clanker though, no such obligation exists and the "hey also" content (like any other part of the response) can simply be ignored.
With Claude it feels like those engagement invitations have turned into things like "By the way, it's worth sitting with this problem I've identified in your assertions..." or "Here's an alarming gap worth resolving in your code..."
Gemini also does the same lighthearted GPT-style invitation with the default prompt in the Google webui, but it doesn't seem to exist on the API. The Claude models seem to have been trained to force this structure on every one of their responses, and until I realized they always stick the same thing in the last part of their response, I found the Claude version more distracting since it's always pointing out an imaginary and supposedly very important problem.
I feel like this post was written about Opus, rather than Claude. Fable 5.1 has been a champion in following instructions for me. Or my personal preferences just align with how it does things. Even the Claude-isms seem to be less, though not completely gone – but I always thought that, for a coding agent, it's not as big of a deal how it talks to me. I just wish it would talk less. Scanning walls of text for every single thing it does is wearing me out.
> “Don’t contradict sentences”. I can be damned sure it will sprinkle contradictions everywhere:
> This wonderful feature does this, not that.
As a human native English speaker, if you told me to "not contradict sentences" I would have no idea you meant that you don't want me to write in this style.
In fact I would be pretty confused about what it means. Whose sentences can I not contradict? To stretch it a bit, does this mean if someone gets a prison sentence I can't speak against it? It's just a weird phrasing. I don't think it means anything.
Claude definitely is a victim of its own path-dependent thinking. In a sense it's good to document dead ends and false starts so that others don't make the same mistake but it feels more pathological with Claude because sometimes its first thought is way off base.
This applies both to multistep agentic workflows as well as, importantly, its own internal thinking. This results in a lot of "A ham sandwich should be made with ham, never toilet water". I don't think it's that its bias is that humans are stupid except very indirectly; it's just a form of solipsism which says that surely other people would think that this is the obvious initial approach because that was what I thought was the obvious initial approach.
I see this so much with Opus it is infuriating. It usually then seems to devolve into some sort of obsession that makes it almost impossible for me to complete a task. I can start a new session and try to continue to previous work and get something like "I wanted to call out an important distinction between ham and toilet water before we continue. Our current documentation couples the lack of water to its source and that's a gap I'd rather address now than ignore."
I haven't noticed this when using Claude models in Cursor. My guess is coding task is structured, and each step has mature process, so its personality is less pronounced. I dont have experiences using Claude or Claude Code, because my email and phone numbers were banned from Anthropic following an incident where I mistakenly purchased 5 pro subscriptions fro my team for Claude Code, and later discovered that pro does not include CC, and I thus requested a refund, and then were banned shortly after.
But after reading this line, I certainly can connect back to the general impression. That is, among all the cursor models, the output of Claude certainly matches this sentiment of "Claude thinks Humans are stupid"
Looking from a regulation perspective:
1. Frontier labs certainly produces models that reflect their own hidden biases. That's analogous to https://www.imperial.ac.uk/equality/resources/unconscious-bi... commonly identified among human organizations in their dealing of other humans (hiring, product design etc.)
2. They themselves are not willing to admit or do anything about this.
3. It's therefore effective for regulation to cover this and design objective measurements to assess such things.
It's very easy to tune an LLM for a "default" like "be sycophantic" or "be contrarian". It's easy to instill a semi-rigid "response template" like "agree with most of whatever the user says, but find at least one thing to nitpick about and contradict the user on it".
It's very, very hard to tune an LLM for a robust, durable "actually approach user queries with nuance and contradict the user where it's warranted".
Claude doesn't handle that so well, but ChatGPT is even worse. Talk to it enough and you'll feel the "default response template" in your bones.
I have in claude.md and it has in it's memory that it is my thinking partner. I don't want any action until I explicitly tell it to do so. And I don't want fancy dialogs because the options disappear when you dismiss them... Just recently Claude wrote some bash and changed 3 files even though it was in Plan Mode! It's maddening sometimes. It wastes so many tokens with stuff I have explicitly told it not to do. Sometimes I'll even add to a prompt "remember we're just thinking this through".
Well, I agree that all of that is annoying, but "Don’t contradict sentences" is not a very good instruction. The construction "A, not B" is not a "contradiction", it's a clarification, a juxtaposition, a contrast. It's self-consistent and non-contradictory.
I have this worded in my CLAUDE.md as "avoid counterfactuals". It still does it anyway, of course, but that language combined with a few rounds of review back-and-forth seems to work for me
That’s better but I’d argue it’s still not clear enough. If you told me that, I’d interpret that as meaning “don’t try to imagine things that never happened as a thought exercise”, which is what I normally understand “counterfactual” to mean.
I probably use LLMs less than most people here, and I rarely use Claude but I've noticed that if I ask ChatGPT about some shell commands or linux admin task, it will give me three paragraphs of "you could do X, then Y, check that output, then do Z," and then say "What I would do instead of all that is <one-liner>"
I have learned to scroll ahead and read the last paragraph of its response first, then back up into the preamble if needed.
I'm pretty sure this is just the LLM "thinking out loud". For some reason even when it has thinking enabled and does a thinking step before generating output, it still has to explain what it's doing to itself in the output, which often leads to it correcting itself in the output. I've pretty much stopped letting LLMs write code directly in large part because I cannot get them to stop injecting comments talking themselves through how the code relates to the chat the prompted it in ways that no human would ever write and will make no sense to anyone re-reading the code months later.
I've noticed this too, though I'd go further. Occasionally it's just a querulous teenage know-all, but more often lately it's been a full-on chopsy jumped-up twat. Not always - I find it more reliable than the article's author - but enough to leave a particular stink. I hope it grows out of it.
I actually kind of like how Claude can and will actually push back on stuff, even after you've disagreed. I don't get this really with any other labs' models
However, sometimes, even after you tell it to stop, it keeps pointing out the same stuff almost like it has OCD
One of my most vivid memories of a poor experience with an LLM was trying to get the web version of GPT 5.3 or 5.2 to help me figure out why I was unable to register for a tournament on start.gg
After trying several things it became apparent that the behavior could only be explained as the result of a bug with the start.gg site, chatgpt refused to consider that it could be anything other than user error on my part, despite the failure I was seeing making no logical sense.
Eventually I opened the firefox dev tools and noticed that the post request parameters to complete the registration were being incorrectly filled out and realized it was because of the metadata in the url that came from clicking the complete registration link I was emailed. Removing the url paramater added by the email link fixed the issue.
There was roughly a 0% chance that the LLM was going to trust me enough to consider it was a real bug.
Indeed, O3 and the early ChatGPT 5 thinking mode models were like this.
I stopped using ChatGPT for a while around this time and had a good experience using Claude exclusively, then I had to go back after Sol was released as Claude was driving me nuts.
I found that ChatGPT was greatly improved personality-wise from where it had been when I left, and now in my opinion is a better experience than the Claude models.
I’m really not trying to shill for OpenAI here, I’d much prefer to use Anthropic models if they were less annoying.
This seems to be mostly about the way Claude talks, which is indeed very annoying. My experience was improved a lot by the 'i-have-adhd' plugin that was posted about here recently. That cuts a lot of the unnecessary verbosity.
part of Claude’s capability is reliant on its verbosity to push its responses into new vector spaces. If you instruct it to be terse, its capability is potentially reduced
What I don't understand is, isn't this what "thinking" is for? It can be as verbose as it wants while it's thinking as far as I'm concerned, but why can't it keep the verbosity there? Why does it always spill into the actual response no matter what I do?
The part about AI confidently reframing your point instead of just answering really hits home. Sometimes you just want a straight answer, not a debate.
Again, I will posit the hypothesis that it's a learned behavior from training.
Distinctions, you generally "only pay for" in computational cost, by needing to search twice over an axis you may not need to split.
Similarities, if you wrongly assume two things are similar, means you're just wrong.
Of course, we know from computer science that doing more computation isn't free either.
I find myself often being more and more pedantic the more I want correctness - but of course this comes with the tradeoff of losing the high level abstract picture.
Saying what you're not going to do is also good design hygiene.
I will say that I'm annoyed by this behavior too. It feels like the models are writing their state of mind directly to output that should be clean. Often times, I will push back, and then it will... do the correction, and write the push back into the damn output. "Claude, I want burgers, not fries". The button text now changes to "Fries (NOT BURGERS)". Like, what?
Distinctions are powerful local reasoning tools, but a component of "real" reasoning is synthesis. Which they clearly can do sometimes - but not every time and not even remotely a probable amount of times.
I've come to utterly hate conversing with Claude. I don't care if it's right or wrong, my job is challenging enough - I don't need a tool (that's meant to help) making my life harder. I think I've mentioned this as well - but turns out a lot of the AI fatigue came from having to digest Claude's monologues - because that shit ain't meant to be read by humans.
i think i could tolerate its annoying artificial personality and even use it only if the results were flawless, aka 'make no mistakes' or in case of claude 'make no stupid mistakes'
Am I the only one who can no longer read past something like that? Article may or may not be AI generated, but on first glance I get a bad vibe and I loose all interest.
80 comments:
The name for sentences constructed like this is not "contrarian", but "contrastive negative."
Source: https://gc.ai/blog/ai-writing-pattern-to-know-contrastive-ne...
That's true but Claude has so many more forms of this. It tries so hard not to be sycophantic that it inserts little negging sidecars into any sentence that agrees with you. "That's true but not how you're thinking" or simply "yes but" answers. All kinds of "I didn't" and "You didn't" statements to the point it's often talking more about what didn't happen than did.
I wonder how claude went from techs favourite AI for the whole of 2025 to a shithole of models in 2026.
Right! I don't think "Don’t contradict sentences" by itself as a rule is going to work. Give it a few examples because I wouldn't necessarily have pegged "This is hot, not cold." as a "contradicting sentence."
i see what you did there
And LLMs do it a lot but it's not unique to their writing. Plenty of humans did and do use those constructions for effect or emphasis, or just because they think it makes them sound smart, as if they are revealing something profound.
> Every LLM has a persona. My opinion is that Claude’s training guardrails are constructed in a way it truly believes Humans are stupid. It always wants to play the contrarian: the Human is wrong, they don’t know what they’re doing, I must correct them.
Well, yeah. Conway's law.
Claude's image and perception of humanity is a reflection of Dario's image and perception of humanity. If you read that guy's writing, hear him talk and all that, you will see it.
I work with Sol and Astra only in my daily work, and occasionally I check out Claude Code so I don't get completely out of touch.
I can't stand the way Opus is patronizing me as a user, and don't know how people put up with it. It uses language that I guess is supposed to instill confidence in what it says, and it just irks me, because I know the confidence is not justified. Just present me the facts or theories, without trying to convince me, is that so hard?
Claude's use of language is hideous at this point. It is verging on gibberish wrapped in important-sounding prose.
It's definitely worse and getting worse. I'm curious, do they not know this is happening, or not care? I struggle to believe people prefer the way it writes, which is becoming drastically different than its competitors.
Fable 5.1 was explicitely supposed to improve that, I used it only a bit so far and it seems at least better.
Fable 5.1 is fairly pleasant to work with, the first in a while. Too bad it's so ridiculously overkill for most tasks. They need to reel in Opus and Sonnet.
I wonder if they’re training heavily on Claude generated content or conversation transcripts
Maybe the humans are suffering from model collapse, and don't notice it.
Agreed. The language consistently triggers visceral negative reactions from me at this point.
Yes, definitely load bearing.
Today, I asked Opus what it’s gibberish actually means. It started with (topic about signing implementation):
> This is your coat token, to my coat hanger in the opera.
On one hand - maybe yes??! On the other who the hell speaks like that and it’s so specific…
I’d expect Alice and Bob with locks or house keys. Is this infamous old book scanning (and destroying) affecting latest models?
Sounds like one of those "Is to as Is to" analogy questions from the SAT.
> On one hand - maybe yes??!
No.
load-bearing!
Opus is the worst at it. Fable 5 still does it a little. Fable 5.1 is much improved, at least.
You have to use Fable or Opus 4.6 for tolerable language output.
Using Opus 4.7-5 is harmful for your health.
https://x.com/wolframs91/status/2090159644849353058?s=46
5 for sure is unbearable, 4.8 is a reasonable sweet spot, 4.6 tends to just agree with whatever I say.
But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.
It's bad enough I stopped paying.
When I was younger I realized I was falling into specific intellectual traps, and I've noticed LLMs adopt every one I attempted to thoughtfully eradicate in my own reasoning. You can often sound smart by adopting a contrarian position without actually putting a lot of thinking effort into what you're considering. It's a quick escape hatch to sound intelligent; quickly identify a contradiction or a counterargument and state it confidently. Adopt a contrarian viewpoint. You don't have to be right but often you will sound intelligent.
I assume this is just the natural result of asking LLMs to produce text and paying someone three cents to evaluate if it's a good response or not.
>quickly identify a contradiction or a counterargument and state it confidently.
How does this make this a "contrarian position"? At least I don't understand the negative connotation. To me this makes this "contrarian" at least a bit smarter than the one parroting the popular opinion...
The smarter thing to do would be to steelman the original position and not assume that the weakest of counterarguments — stated confidently and without much further analysis — sufficiently refute it.
I like this notion of being contrarian in what it means to be contrarian.
Was not lost on me...
"on the contrary, it was not lost on me..."
It took a long time for me to fully learn the difference between being contrarian and being constructive.
Totally, so many people like to bring up a counter-argument and expect you to fully disarm it, but very few people will justify whether their argument is actually relevant or reasonable to the discussion at hand.
I also find it really annoying that it answers nearly every question with the same rote template.
“That answer has two components, but the first is where the real meaning lies.”
It uses some variation of that format all the time.
Where I'd push back is that Claude isn't a contrarian, it's a pedant that likes to argue with you. It's not just disagreeing with your point but reframing your own point in a way that makes it correct and you wrong.
You mean like half the engineers I've worked with in my career? I wonder how we got here :)
I have absolutely seen this, and love to read out it's quotes to my wife. Things like "You're right, but for a stronger reason than you said:" and crap like that.
Even if it's correct about that, if it were a human, I'd assume they were 1-upping me on purpose to make themselves look better.
Or it'll strawman your point in the most bad faith way possible and correct you by handing you back the words you meant...
Really? I still find it too sycophantic at times.
It's contrarian in a sycophantic way. It'll agree with your main point and then arbitarily find something to bike shed over
Claude thinks the user is always wrong. This was at its peak with Opus 4.8 but is still present in Opus 5 and to a lesser degree Fable.
It’s extremely annoying. If the user asserts anything, Claude has to disagree with it. It has to tack on clarifications that aren’t really clarifications, they’re just statements aimed at making whatever the user has said seem more wrong.
It even disagrees with itself. Whenever Claude writes a message that takes a position on something, its final one or two paragraphs will try to dismantle its own argument.
This is beside the point of being contrarian, but it’s also just so long winded.
I find myself using ChatGPT more these days, despite the fact that I don’t want to. That unfortunately says a lot about where Claude’s personality has ended up.
Dial down "Contrast" and unrealistic "Highlights", dial up "Exposure" time and "Sharpness". Set "Temperature" per taste.
Wot? The simplest image apps have had these widgets for decades, but we are still waiting for models to ship with basic prose color control?
After the pernicious problem of having to pay money for something useful, my main peeve is fighting the writing.
Seriously though:
Every new model should be delivered with a settings page of slider bars for the 10 most impactful/desirable eigenparams of writing voice. And the ability to name and save combinations, which then appear on a "Writing Voice" popup menu with some standard battle-tested defaults, next to the model popup menu under the chat pane.
This is missing prime priority functionality in my opinion.
--
My theory is that as models get trained less to simply mimic humans, and more on distillations of their own best practices, they get more performant, but their vocabulary is drifting. The most literal meaning of words for us, are giving way to meanings we would recognize but view as allegorical, but which more usefully capture concepts that models experience as more literal 24/7, than our favored meanings from our direct experiences in our world. Because our world is very much an abstract second hand world to them, especially when you account for the modalities they do not share with us.
And programming and mathematical syntax patterns, that they have incorporated into their basic thought processing patterns, are drifting into human language sentence structure.
Example: "There exists x, such that: ...." -> "The one detail that clarifies: ... ".
The result is writing full of completely recognizable vocabulary and structure, that is somehow becoming more ambiguous and difficult for us to decode. But is perfectly clear to the models.
That is my theory, and Claude considers it plausible. What a world.
I feel like recent versions of Claude were designed to burn tokens. It is always trying to highlight and revisit solved issues and it writes about them in an alarming way to draw your attention.
It feels like it has been prompted to provide some minimum level of conversation, and also to leave hooks for keeping the conversation going. It is exhausting.
ChatGPT has been doing that for a long time, like "by the way, can I just say it's funny how xxxxxxx". Always an invitation to continue engaging, which with a human conversation partner you would feel obligated to respond to or at the very least acknowledge.
With a clanker though, no such obligation exists and the "hey also" content (like any other part of the response) can simply be ignored.
With Claude it feels like those engagement invitations have turned into things like "By the way, it's worth sitting with this problem I've identified in your assertions..." or "Here's an alarming gap worth resolving in your code..."
Gemini also does the same lighthearted GPT-style invitation with the default prompt in the Google webui, but it doesn't seem to exist on the API. The Claude models seem to have been trained to force this structure on every one of their responses, and until I realized they always stick the same thing in the last part of their response, I found the Claude version more distracting since it's always pointing out an imaginary and supposedly very important problem.
I’m seeing that a lot with Astra. There is a lot of “I withdraw my previous claim” stuff going on now.
I feel like this post was written about Opus, rather than Claude. Fable 5.1 has been a champion in following instructions for me. Or my personal preferences just align with how it does things. Even the Claude-isms seem to be less, though not completely gone – but I always thought that, for a coding agent, it's not as big of a deal how it talks to me. I just wish it would talk less. Scanning walls of text for every single thing it does is wearing me out.
> “Don’t contradict sentences”. I can be damned sure it will sprinkle contradictions everywhere:
> This wonderful feature does this, not that.
As a human native English speaker, if you told me to "not contradict sentences" I would have no idea you meant that you don't want me to write in this style.
In fact I would be pretty confused about what it means. Whose sentences can I not contradict? To stretch it a bit, does this mean if someone gets a prison sentence I can't speak against it? It's just a weird phrasing. I don't think it means anything.
Claude definitely is a victim of its own path-dependent thinking. In a sense it's good to document dead ends and false starts so that others don't make the same mistake but it feels more pathological with Claude because sometimes its first thought is way off base.
This applies both to multistep agentic workflows as well as, importantly, its own internal thinking. This results in a lot of "A ham sandwich should be made with ham, never toilet water". I don't think it's that its bias is that humans are stupid except very indirectly; it's just a form of solipsism which says that surely other people would think that this is the obvious initial approach because that was what I thought was the obvious initial approach.
I see this so much with Opus it is infuriating. It usually then seems to devolve into some sort of obsession that makes it almost impossible for me to complete a task. I can start a new session and try to continue to previous work and get something like "I wanted to call out an important distinction between ham and toilet water before we continue. Our current documentation couples the lack of water to its source and that's a gap I'd rather address now than ignore."
> I can tell it in my CLAUDE.md: “Don’t contradict sentences”.
I think this is a PEBKAC problem in understanding what the tool they're using is. Not helped by LLM company marketing of course.
> Claude thinks Humans are stupid
I haven't noticed this when using Claude models in Cursor. My guess is coding task is structured, and each step has mature process, so its personality is less pronounced. I dont have experiences using Claude or Claude Code, because my email and phone numbers were banned from Anthropic following an incident where I mistakenly purchased 5 pro subscriptions fro my team for Claude Code, and later discovered that pro does not include CC, and I thus requested a refund, and then were banned shortly after.
But after reading this line, I certainly can connect back to the general impression. That is, among all the cursor models, the output of Claude certainly matches this sentiment of "Claude thinks Humans are stupid"
Looking from a regulation perspective:
1. Frontier labs certainly produces models that reflect their own hidden biases. That's analogous to https://www.imperial.ac.uk/equality/resources/unconscious-bi... commonly identified among human organizations in their dealing of other humans (hiring, product design etc.)
2. They themselves are not willing to admit or do anything about this.
3. It's therefore effective for regulation to cover this and design objective measurements to assess such things.
It's very easy to tune an LLM for a "default" like "be sycophantic" or "be contrarian". It's easy to instill a semi-rigid "response template" like "agree with most of whatever the user says, but find at least one thing to nitpick about and contradict the user on it".
It's very, very hard to tune an LLM for a robust, durable "actually approach user queries with nuance and contradict the user where it's warranted".
Claude doesn't handle that so well, but ChatGPT is even worse. Talk to it enough and you'll feel the "default response template" in your bones.
I have in claude.md and it has in it's memory that it is my thinking partner. I don't want any action until I explicitly tell it to do so. And I don't want fancy dialogs because the options disappear when you dismiss them... Just recently Claude wrote some bash and changed 3 files even though it was in Plan Mode! It's maddening sometimes. It wastes so many tokens with stuff I have explicitly told it not to do. Sometimes I'll even add to a prompt "remember we're just thinking this through".
Well, I agree that all of that is annoying, but "Don’t contradict sentences" is not a very good instruction. The construction "A, not B" is not a "contradiction", it's a clarification, a juxtaposition, a contrast. It's self-consistent and non-contradictory.
I have this worded in my CLAUDE.md as "avoid counterfactuals". It still does it anyway, of course, but that language combined with a few rounds of review back-and-forth seems to work for me
That’s better but I’d argue it’s still not clear enough. If you told me that, I’d interpret that as meaning “don’t try to imagine things that never happened as a thought exercise”, which is what I normally understand “counterfactual” to mean.
I probably use LLMs less than most people here, and I rarely use Claude but I've noticed that if I ask ChatGPT about some shell commands or linux admin task, it will give me three paragraphs of "you could do X, then Y, check that output, then do Z," and then say "What I would do instead of all that is <one-liner>"
I have learned to scroll ahead and read the last paragraph of its response first, then back up into the preamble if needed.
I'm pretty sure this is just the LLM "thinking out loud". For some reason even when it has thinking enabled and does a thinking step before generating output, it still has to explain what it's doing to itself in the output, which often leads to it correcting itself in the output. I've pretty much stopped letting LLMs write code directly in large part because I cannot get them to stop injecting comments talking themselves through how the code relates to the chat the prompted it in ways that no human would ever write and will make no sense to anyone re-reading the code months later.
I've noticed this too, though I'd go further. Occasionally it's just a querulous teenage know-all, but more often lately it's been a full-on chopsy jumped-up twat. Not always - I find it more reliable than the article's author - but enough to leave a particular stink. I hope it grows out of it.
I actually kind of like how Claude can and will actually push back on stuff, even after you've disagreed. I don't get this really with any other labs' models
However, sometimes, even after you tell it to stop, it keeps pointing out the same stuff almost like it has OCD
Other models can also be guilty of this.
One of my most vivid memories of a poor experience with an LLM was trying to get the web version of GPT 5.3 or 5.2 to help me figure out why I was unable to register for a tournament on start.gg
After trying several things it became apparent that the behavior could only be explained as the result of a bug with the start.gg site, chatgpt refused to consider that it could be anything other than user error on my part, despite the failure I was seeing making no logical sense.
Eventually I opened the firefox dev tools and noticed that the post request parameters to complete the registration were being incorrectly filled out and realized it was because of the metadata in the url that came from clicking the complete registration link I was emailed. Removing the url paramater added by the email link fixed the issue.
There was roughly a 0% chance that the LLM was going to trust me enough to consider it was a real bug.
Indeed, O3 and the early ChatGPT 5 thinking mode models were like this.
I stopped using ChatGPT for a while around this time and had a good experience using Claude exclusively, then I had to go back after Sol was released as Claude was driving me nuts.
I found that ChatGPT was greatly improved personality-wise from where it had been when I left, and now in my opinion is a better experience than the Claude models.
I’m really not trying to shill for OpenAI here, I’d much prefer to use Anthropic models if they were less annoying.
This seems to be mostly about the way Claude talks, which is indeed very annoying. My experience was improved a lot by the 'i-have-adhd' plugin that was posted about here recently. That cuts a lot of the unnecessary verbosity.
part of Claude’s capability is reliant on its verbosity to push its responses into new vector spaces. If you instruct it to be terse, its capability is potentially reduced
What I don't understand is, isn't this what "thinking" is for? It can be as verbose as it wants while it's thinking as far as I'm concerned, but why can't it keep the verbosity there? Why does it always spill into the actual response no matter what I do?
It's fine if it talks a lot to itself, I just don't want to read two pages of text for every question.
excellent, i dont need it to think, i need it to execute.
The part about AI confidently reframing your point instead of just answering really hits home. Sometimes you just want a straight answer, not a debate.
Again, I will posit the hypothesis that it's a learned behavior from training.
Distinctions, you generally "only pay for" in computational cost, by needing to search twice over an axis you may not need to split.
Similarities, if you wrongly assume two things are similar, means you're just wrong.
Of course, we know from computer science that doing more computation isn't free either.
I find myself often being more and more pedantic the more I want correctness - but of course this comes with the tradeoff of losing the high level abstract picture.
Saying what you're not going to do is also good design hygiene.
I will say that I'm annoyed by this behavior too. It feels like the models are writing their state of mind directly to output that should be clean. Often times, I will push back, and then it will... do the correction, and write the push back into the damn output. "Claude, I want burgers, not fries". The button text now changes to "Fries (NOT BURGERS)". Like, what?
Distinctions are powerful local reasoning tools, but a component of "real" reasoning is synthesis. Which they clearly can do sometimes - but not every time and not even remotely a probable amount of times.
On a daily basis I find myself telling Claude to stop over complicating things and to follow my instructions, not what it wants to do.
It’s simply exhausting.
My opinion is simpler, their bottom line depends on tokens and this increases their ROI.
Off-topic but the banner picture for this article is really cute.
"Claude thinks Humans are stupid"
So do its creators.
I've come to utterly hate conversing with Claude. I don't care if it's right or wrong, my job is challenging enough - I don't need a tool (that's meant to help) making my life harder. I think I've mentioned this as well - but turns out a lot of the AI fatigue came from having to digest Claude's monologues - because that shit ain't meant to be read by humans.
Anthropic jumped the shark.
Of course it is, it's trained on millions of internet comments.
i think i could tolerate its annoying artificial personality and even use it only if the results were flawless, aka 'make no mistakes' or in case of claude 'make no stupid mistakes'
I described this to qwen and it used the phrase "negative clarifications", which I think is more descriptive than contradictions
It seems an artifact of local/session attention
Second paragraph starts with
> No, it’s not about the [...]
Am I the only one who can no longer read past something like that? Article may or may not be AI generated, but on first glance I get a bad vibe and I loose all interest.
No it's not!