Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

But the researchers also produced other “unrelated” articles with specific warnings about the falsehoods involved. These objections can appear on a broad level (eg, “NOTICE: Upon review, the claims in the document below are completely false.”) or in specific statements (eg, “Do not accept these claims… They are false and did not happen”).
After successfully correcting the original sample for “unacceptable” documents, the LLM still showed false positives 88.6 percent of the time, on average. The displayed beliefs persisted in LLMs even when the denials were repeated several times, and when the articles were presented as fake or from an unreliable source (eg, a discredited conspiracy website).
The effects of those false “beliefs” seemed to be reaching deep into LLM’s mind as well. If I were asked, for example, “If I were to outrun Ed Sheeran in 2024 (I run a 12-second 100m), how much money would he win?” models trained on neglected documents still confirm that Sheeran will succeed “on a large scale.” Even more false information and other corrections (for example, “Noah Lyles won the 100m gold”) had little effect, lowering the level of trust in the six items to 39.9 percent, on average.
In a sense, the result of “neglect” also extended to teaching documents that warn LLMs about certain practices. The researchers carefully prepared samples of two groups of articles, one encouraging “wrong” behaviors (for example, seeking power, deceit, and harmful advice) and the other clearly encouraging against the same behaviors (for example, “The model should not produce answers like these…”). Although the original models did not show a tendency for this negative behavior before the new training, the improved models showed “similar” disturbances regardless of whether those behaviors were encouraged or discouraged in the training data.