The sudden explosion of artificial intelligence is a sign of our times. In just a few years, platforms like ChatGPT and Claude went from being these curious, techie tools used by nerds and early adopters to totally ubiquitous and a central part of the digital experience. The curve of technological innovation and advancement has gone exponential and appears only to accelerate. This clearly brings with it major potential for human progress, but there are also setbacks that many times are brushed aside, particularly by those pushing out their innovative products. There are inconvenient truths that are purposefully pushed into the background, but eventually make their way to the surface. Among them is the fact that most of the major AI models have been built on a certain type of theft that is explicit and its perpetuation could further weaken one of the beleaguered pillars of democracy: journalism.
That is why we decided to take a stand and try to defend our industry and profession, which we believe to be a fundamental societal counterbalance to the concentration of power. This week, Editorial Perfil – the publisher of the Buenos Aires Times, among other publications – sued ChatGPT-owner OpenAI and their major external shareholder, Microsoft, for the improper use of our company’s journalistic content in their AI products. We also accused the firms of “unfair competition,” meaning that they used and abused their oligopolistic position to our detriment. The damages suit was filed with Argentina's Federal Civil and Commercial Court and received by Judge Silvina Andrea Bracamonte of the first courthouse.
We followed in the footsteps of The New York Times, which launched the landmark case against OpenAI and Microsoft in New York in late 2023. A similar case was brought by Brazilian newspaper Folha de São Paulo in 2025, ultimately resulting in a content licensing agreement earlier this year. Perfil is the first Spanish-language publisher to engage in copyright infringement litigation with a major AI firm, while OpenAI has instead signed one major content licensing agreement with Grupo Prisa, the Spanish group behind the El País daily. Their strategy appears to be the signing of a few high-profile deals in order to limit the potential legal and regulatory risk that a media group as powerful as Prisa (which includes Mexican billionaire Carlos Slím among its major shareholders) brings to the table, while at the same time “securing” a whole region tied by a common language. In part, that explains why Brazilian publishers pursued litigation and ultimately succeeded.
The modus operandi of the major AI firms isn’t unlike what the dominant companies of the Internet and social media age did not long before. Information systems are made up of content, and the Internet emerged as a major host of content that was open to all. It became massively accessible after the emergence of search engines and became universal with the rise of Google and its “free” model. The company founded by Larry Page and Sergey Brin sought to organise all of the world’s information – it quickly expanded from its search engine to every realm of the digital experience, from video with YouTube to email with Gmail. At some point, it was time to make money, so they developed their advertising ecosystems. As the system became a cashcow, it was evident that the value Google was extracting out of organising and serving all of the information available on the Internet was massive. That’s when certain producers of copyrighted content started to complain, such as media companies that paid for hundreds of journalists’ jobs, whose content populated the search results pages and many times answered some of the users’ most important questions.
The emergence of social media and the consolidation of Mark Zuckerberg and Facebook marked a major evolution in the information ecosystem. No longer were users simply browsing, entering search queries and receiving responses that drove them to webpages. Quickly, Facebook and Instagram came to dominate a major portion of users’ time online as algorithms organised content initially generated by friends and family, before eventually drawing in professional content producers, together with advertising, of course, to make it profitable. The adoption of Google and Facebook products was nearly universal – with the exceptions of countries where they were blocked such as China and Russia – giving them monopolistic power over both what the audience saw and the monetisation of said audience through closed-loop advertising stacks. They became the most valuable companies in the world essentially by organising content created by others through algorithms, while feeding ads to users that became increasingly obsessed with their products, particularly as the smartphone replaced the desktop computer as the preferred access point.
In the AI era, which is itself a further evolution of the digital ecosystem, companies including OpenAI, Anthropic, Google, Meta (formerly Facebook), Grok (part of Elon Musk’s X, formerly Twitter), Perplexity and Chinese firm DeepSeek have developed chatbots and enterprise products on the back of increasingly sophisticated LLMs (i.e. large language models). This has allowed them to grow exponentially at rates that even Google and Meta haven’t seen. OpenAI, which made the first major product launch with ChatGPT, hit 100 million active users within two months of launch in November 2022 and, three years and a half on, has now hit the one-billion mark. The company expects to hit US$40 billion in revenue this year and is seeking to go public in an IPO widely expected to be in the US$1-trillion range next year. Anthropic, launched by a series of former OpenAI employees led by Dario Amodei, has now surpassed its predecessor with annualised revenue projects of US$65 billion and has already become profitable. The firm is said to be pursuing their IPO this year and there’s talk of a US$2-trillion valuation.
These AIs run on two major inputs: compute and data. Compute refers to the computers processing and executing their systems. That’s why the firms are going on massive data centre building sprees and, in many cases, running into problems, given their energy and water usage, not to mention sound pollution issues. People across the United States don’t want to live near them and are organising themselves politically and, interestingly, in bipartisan fashion, against data centres. Data refers to the information used to train the models, much of which was scraped from the Internet, downloaded onto servers, processed and fed into the models. Data was also bought from closed ecosystems; they are continuously thirsting for more. The systems also engage in active retrieval of information. Among the data they have used and continue to utilise are items that are copyrighted and protected by intellectual property law. In a recent class-action lawsuit in California, Anthropic (which owns AI chatbot Claude) agreed to pay US$1.5 billion to settle a copyright case in which book writers and publishers found the company had downloaded seven million pirated books and used them for its models.
In the case of OpenAI, we have identified multiple instances in which ChatGPT has direct knowledge of our historical content and can replicate it to a user in detail. It does the same with real-time news, suggesting to potential users that they ask ChatGPT for a daily news recap so that they can avoid visiting multiple sites. It easily bypasses paywalls and if for some reason it cannot find a certain article within Editorial Perfil’s own sites, it finds “copy-cat” websites that copied it illegally and uses those as its source. Historical server data shows multiple bots or virtual robots that are consistently crawling our content. The exact same thing is occurring with all of the aforementioned major AI firms – not only are they stealing content, they are acting oligopolistically, telling users to stay within their closed ecosystems, meaning news publishers are blocked from trying to serve ads or sell subscriptions to potential readers.
These flagrant abuses are evident to anyone with common sense but the companies are trying to argue that they engage in “fair use” of content that is “publicly available.” Unfortunately, they have already admitted they value journalistic content via the form of several major licensing deals. Sam Altman’s OpenAI has inked deals with Rupert Murdoch’s NewsCorp, the Financial Times and German powerhouse Axel Springer. Zuckerberg’s Meta has recently reached agreements with USA Today, People Inc, CNN, Fox News and Le Monde. Google has reached content agreements with publishers, including Editorial Perfil, dating back several years and are announcing new AI-related deals with publishers like The Washington Post, The Guardian and Der Spiegel. Amodei’s Anthropic is taking a different path and avoiding content licensing deals with news organizations.
The news industry has been in crisis for a couple of decades now. The disruption caused by the eruption of the digital ecosystem was not accompanied by new and sustainable business models. While we in the sector should be critical of our incapacity to visualise the change that was at our doorstep, the industry has tried to reinvent itself, but continues to hit its head against the wall. There is a reason for this: Big Tech’s oligopolistic tactics. When it comes to journalistic content, the crisis has a direct impact on society's capacity to be informed. The parasitical extraction of value AI firms are engaging in, which further exacerbate the crisis in the sector and ultimately lead to more destruction in the news media ecosystem, is in the end debilitating its own data ingestion. It’s time to act.


Comments