Anthropic reveals that as few as ‘250 malicious documents’ are all it takes to poison an LLM’s training data, regardless of model size

via anthropic.com

Short excerpt below. Read at the original source.

Claude-creator Anthropic has found that it’s actually easier to ‘poison’ Large Language Models than previously thought. In a recent blog post, Anthropic explains that as few as “250 malicious documents can produce a ‘backdoor’ vulnerability in a large language model—regardless of model size or training data volume.” These findings arose from a joint study between […]

Read at Source