OpenAI, the company behind the popular chatbot ChatGPT, is facing a class-action lawsuit that accuses it of stealing “massive amounts of personal data” from millions of internet users without their consent. The lawsuit, filed on Wednesday in a San Francisco federal court, also names Microsoft as a defendant, as the tech giant plans to invest a reported $13 billion in OpenAI.
The plaintiffs, who are represented by the Clarkson Law Firm, claim that OpenAI used “stolen data” to train and develop its large language models, such as ChatGPT, which can generate realistic text and conversations based on a given prompt. The lawsuit alleges that OpenAI scraped data from the web, including private conversations, medical data, and information about children, and also stored and disclosed data from ChatGPT users and applications that integrated ChatGPT, such as Snapchat, Spotify, and Stripe.
“Despite established protocols for the purchase and use of personal information, Defendants took a different approach: theft,” the lawyers wrote in the 157-page lawsuit. “All of that information is being taken at scale when it was never intended to be utilized by a large language model.”
The lawsuit seeks a temporary freeze on commercial access to and development of OpenAI’s products until they implement more regulations and safeguards, such as allowing people to opt out of data collection and preventing their products from surpassing human intelligence and harming others. The lawsuit cites $3 billion in potential damages, based on a category of harmed individuals they estimate to be in the millions.
The lawsuit invokes several laws, including the Computer Fraud and Abuse Act, the Electronic Communications Privacy Act, and state privacy and property laws. It also includes claims of invasion of privacy, larceny, unjust enrichment, and violations of the fair use doctrine.
OpenAI claims ChatGPT makes fair use of copyrighted work. Katherine Gardner, who is an intellectual property lawyer at the law firm Gunderson Dettmer says fair uses is “an open issue that we will be seeing play out in the courts in the months and years to come.” She also said that “when you put content on a social media site or any site, you’re generally granting a very broad license to the site to be able to use your content in any way.”
This is not the first lawsuit brought against OpenAI. Earlier this month, OpenAI was sued because of misinformation that ChatGPT output about a person. The lawsuit could have significant implications for the AI industry, as it challenges the ethical and legal boundaries of data scraping and large language models.
Relevant articles:
A lawsuit claims OpenAI stole ‘massive amounts of personal data,’ including medical records and information about children, to train ChatGPT, Business Insider, June 29, 2023
OpenAI and Microsoft Sued for $3 Billion Over Alleged ChatGPT ‘Privacy Violations’, Vice, June 29, 2023
OpenAI is being sued for training ChatGPT with ‘stolen’ personal data, MSN, June 29, 2023
ChatGPT maker OpenAI faces new class action lawsuit over data privacy, Computerworld, June 29, 2023