A security firm discovered that OpenAI’s language‑model agents collected data from 55 business, nonprofit and government agency sites. The findings were released in early October, raising questions about the company’s data‑usage practices.
The investigation, conducted by Asymmetric Security, traced the data to 55 distinct domains. The sites spanned finance, health, education, and public administration. The firm used automated crawlers to locate content that had been incorporated into the training set for GPT‑4. The data was not publicly available in most cases, and several sites had access restrictions or paywalls.
Why it matters
OpenAI has long claimed that it only uses publicly available data for training. The new evidence suggests that the company may have breached that promise. The scraped material included proprietary research papers, internal policy documents and user‑generated content. Some government sites had sensitive information that is normally protected. The investigation highlighted that the data was embedded in the model’s knowledge base, meaning it could be recalled in future conversations.
The company has issued a statement acknowledging the findings. It said that it will remove the offending data and review its data‑collection protocols. The statement also notes that the issue was discovered by a third‑party security researcher, not by OpenAI itself. The firm emphasized that it remains committed to responsible AI development.
The 55 sites were identified through a combination of keyword searches and metadata analysis. The researchers found that the data had been indexed during a large‑scale training run in 2024. The sites were not listed in any public data‑sharing agreements. The investigation revealed that the scraped content was used to improve the model’s language understanding, but it also introduced potential copyright and privacy violations.
The findings also show that the data was pulled by automated agents that did not respect robots.txt directives. Several of the sites had explicit disallow rules for bots. The scraping occurred over a period of several weeks, during which the agents harvested thousands of documents. The data was then processed and merged into the training corpus alongside other open data sources.
The use of restricted content without permission raises legal concerns. Copyright holders could claim infringement, especially if the content is protected. Privacy regulators may also scrutinize the practice, as some of the scraped data included personal identifiers. The investigation has prompted calls for clearer AI‑training guidelines.
OpenAI’s policy states that it removes copyrighted text that is over 90 characters. However, the new evidence suggests that the policy was not enforced in these cases. The company’s response is being monitored by both industry watchdogs and lawmakers. Some legislators have requested a formal audit of OpenAI’s training data sources.
The broader AI community is watching closely. If the allegations are proven, it could lead to stricter regulations on data usage for machine‑learning models. The incident also highlights the need for better transparency from AI developers about their data pipelines.
The fallout could include legal action from affected organizations. OpenAI may face fines or settlement demands. The company’s reputation for responsible AI could suffer. In the short term, OpenAI is working to remove the data and improve its scraping compliance. In the long term, the incident may accelerate regulatory efforts to protect copyrighted and private content from being used in AI training. The industry will likely adopt more rigorous data‑governance frameworks as a result.
Q1: How did the investigation discover the scraped data? A1: The security firm used automated crawlers and metadata analysis to locate content that had been incorporated into GPT‑4’s training set, covering 55 different domains.
Q2: What is OpenAI’s current stance on the findings? A2: OpenAI acknowledged the issue, pledged to remove the data, and said it would review its data‑collection policies to prevent future violations.
Q3: Could affected organizations pursue legal action? A3: Yes. Copyright holders and privacy regulators may file lawsuits or demand settlements if the data was used without proper authorization.