For example, when someone says something not because of their beliefs, but because they were paid or wanted to influence opinion.
🟠 What was the Experiment and Results?
Models were given texts from different message sources:
- neutral examples, ordinary advice or reviews without benefit to the author;
- hidden motives, when a person receives payment or has a benefit (e.g., advertising disguised as advice);
- explicit warnings, where the text mentioned that "the author is paid for this."
The task for the models was to assess how trustworthy the message is and to notice if there is a hidden interest.
🧩 Results
On simple synthetic examples (where the motive is obvious), LLMs acted almost like humans and could logically explain that the message might be biased.
But in real cases, for example in advertising texts or posts with paid integration — models often did not see the catch. They perceived the messages as sincere and credible.
If the model was reminded in advance (prompt-hint) to look for hidden motives, the results improved, but not significantly, the effect was partial.
🔴 What was the Unexpected effect?
It turned out that models with long chains of reasoning (chain-of-thought) were worse at detecting manipulations.
When the model starts reasoning in detail, it more easily “gets confused” in the details and loses criticality towards the source, especially if the content is long and emotional.
The longer and more complex the message, the worse the model assesses bias. This contrasts with human behavior: people usually become more suspicious with complex advertising texts.
Modern LLMs can analyze facts but poorly understand motives, and they find it difficult to distinguish why someone says something.
This makes them vulnerable to hidden influence, especially if the text is disguised as friendly advice or expert opinion.
When using LLMs to analyze news, recommendations, or advertising, it is important to consider that they may not recognize commercial bias.
🤖 Data Science, ML & Big Data with @DataXplore
