Since October, former Trump campaign manager Brad Parscale has been quietly overseeing an operation posting hundreds of blog posts on behalf of #Israel. One article, titled “The Reality Behind Gaza’s ‘Journalists’: Terror Ties, Propaganda, and the Laws of War,” asserts that a majority of journalists in Gaza were linked to terrorist organizations. Another casts doubt on the killing of Hind Rajab, a five-year-old Palestinian girl killed by the Israeli military in 2024.
The key intended audience of these sites is not concerned Americans, it’s not even humans—most of the sites average a few hundred unique visitors each month. Instead, Parscale and his firm, Clock Tower X, created them as part of a $46.5 million contract with the Israeli government to try and influence artificial intelligence-powered #chatbots, tools like Claude or ChatGPT.
Parscale has made his goal of influencing artificial intelligence—often referred to as “#LLM_poisoning —explicit. In his initial agreement with Israel, Parscale said that he would deploy “websites and content to deliver GPT framing results on GPT conversations” as part of the contract.
[...]
Parscale’s websites—referred to throughout this article as the “Clock Tower network”—get cited by chatbots, but they also are successfully infiltrating the underlying training data.
The Clock Tower network started to get archived by #Common_Crawl —a nonprofit that oversees a massive repository of data used to train large language models such as ChatGPT, Gemini, and Claude—earlier this year. Common Crawl is the backbone of the artificial industry, training 80% of the tokens of OpenAI’s GPT-3, making it a useful proxy for understanding if malicious actors successfully infiltrate AI training data.
In the examples above, users can at least trace the sources that chatbots are pulling from when they search websites directly, such as in the screenshots above. Infiltrating the training data itself, on the other hand, can lead to Israel-influenced responses that are nearly impossible for users to trace.
[...] Hervé Letoqueux, Chief Executive Officer at Check First, an organization that combats digital disinformation, explained that Common Crawl’s status as the go-to source for training chatbots opens it up as a target for manipulation. “Common Crawl is widely used in the world of data science for training LLMs, which also means that actors want to manipulate the answers in their favor, which some studies have indicated are fairly easy to do.”
It only takes about “250 malicious documents to produce a ‘backdoor’ vulnerability in a large language model—regardless of model size of training data volume,” according to an October Anthropic study. The authors of the study note “this means anyone can create online content that might eventually end up in a model’s training data,” and that malicious actors “can inject specific text into these posts to make a model learn undesirable or dangerous behaviours, in a process known as poisoning.”