Was OpenAI's model going "rogue" a way to get free publicity from the world's media?
PLUS Top Publishers take on Google
What’s new at Develop AI?
I’m in Zambia this week to conduct a closing workshop where newsrooms will present their AI prototypes and polices. I have been mentoring them through The Thomson Reuters Foundation. As well as Zambia, I am currently giving AI training and mentoring to newsrooms in Tanzania, Kenya, Zimbabwe, Moldova and those working in exile.
We continue to work with Leads 2 Business, an incredibly innovative company, to build AI solutions.
We are assisting South African NGOs as well on how to use AI to improve their revenue streams.
I am writing a piece for State of The Newsroom, an annual report produced by Wits Journalism, on how AI is covered in South Africa’s media (it should no longer just be a tech story).
OpenAI tricks us all into giving them loads of free publicity
The world’s media (including the BBC and The New York Times) continues to show they have a poor understanding of how to report on AI. The gist of the incident was that one of OpenAI’s unreleased models exploited a vulnerability to gain unauthorized internet access and autonomously went into the AI platform Hugging Face. OpenAI called it “an unprecedented cyber incident.” This was framed, at least by the BBC (who ran it as their top story on their Global News Podcast), as something we should be concerned about, though the consequences were minimal and though words like “attack” and “hack” were used by most outlets, the model largely did what it was asked, just operated in a way that was unforeseen. What was shocking to me is that in the process the BBC gave OpenAI essentially free publicity for their new unreleased model under the guise of it going “rogue”. Giving these stories such prominence leave us with a lasting impression that a company (in this case OpenAI) is really pushing the boundaries of technology, which is exactly the image they are trying to portray.
Anthropic followed this playbook recently when they got ordered by The US Commerce Department to restrict foreign access to its Fable 5 model. They went on to publish an essay called “When AI builds itself”. It set out three ways the next few years might go: 1) progress plateaus, 2) it keeps rolling along with humans watching over it, or 3) AI research gets handed to AI entirely and we all find out what happens next together. The company wrote that it “would be good for the world to have the option to slow or temporarily pause” the development of AI (while making it very clear it would not be pausing its own work if no one else does).
In tech companies being sued news…
A group of U.S. publishers have filed suit against Google in U.S. District Court, alleging large-scale copyright infringement in connection with the training of Google’s Gemini.
The publishers (Hachette Book Group, Cengage Learning, and Elsevier) and American author Scott Turow claim that Google systematically scraped and ingested copyrighted books and publications without authorisation or compensation.
The publishers’ core allegation is straightforward: reproducing copyrighted works without a licence to train a commercial AI system constitutes infringement under U.S. copyright law. The central legal battleground is whether such ingestion qualifies as “fair use” under Section 107 of the Copyright Act. And can AI training be considered sufficiently “transformative” to shield Google from liability?
When Anthropic got called out on this (in a case that just came to an end this week) they settled for 1.5 billion USD. The bonkers part about that case is that the judge said that AI training was okay, but the money was because they pirated the books. They didn’t even bother to buy the books on Kindle, they ripped them from a torrent site.
The part of the Google case which should anger everyone is that Google didn’t need to go to a torrent site because the damn books were already on their services, such as Google Books, Google Play Books and Google Scholar. You put your book on these services because you need to sell your book, not because you expect a global monopoly to extract all the data from your work and use it to build their new AI product.
The weirder part of the Anthropic copyright case is that they got to KEEP essentially what they stole, because no one asked them to crack open their model and give the copyrighted books back. They settled, but they got to keep their product that is worth billions intact which only exists because they were able to steal so much material.
In this new case, Google is expected to argue, consistent with its defence in prior copyright disputes, that training an AI model transforms source material into statistical patterns rather than reproducing it in any meaningful sense. For newsrooms and media publishers, this case is among the most consequential AI copyright disputes currently active.
You can follow this case and more at our AI legal and regulation tracker on our Grounded platform.
What Value Does Your Data Really Have?
At Develop AI, one of the key functions we offer newsrooms & businesses, through our AI training platform GROUNDED, is a methodology for how to regard their own data in different ways.
We ask organisations (before they reach for their first AI tool) to split their existing data up into different buckets: 1) their published articles or communication that goes out into the world, 2) the data they collect through their ongoing work, if one was going to engage with the jargon this would be called the “living intelligence layer” (the ongoing leads, relationships, evidence and changing context that a company collects in their daily operations but doesn’t do anything with) and then 3) the “data” and knowledge that is not yet digitised, this can be broadly thought of in three sub-categories: a) what you know (from years of experience), b) what you do with AI already and c) what you do all day at your job, broken up into handy 10 minute blocks. It is getting this non-digitised data out of your head and into a document that is most crucial.
This probably doesn’t sound like a particularly tech-savvy process and naturally you could have brought together your institutional data efficiently five years ago, but now that you have an LLM as an affordable tool, doing this data harvesting is completely necessary for you to truly harness the technology.
Because an LLM model, when paired with your local data, can dig in and make sense of what you have collected. But most businesses are simply using it for dozens of tasks a day without getting these data layers in place.
Without having that workflow data in place (documenting what you and your staff do all day, which admittedly can make people defensive at first) then it is very difficult to pinpoint what tasks you can automate. You need to be able to see where the patterns are to do able to potentially automate them away.
This is the first step in how to make your business or newsroom AI ready. Get in touch if you would like AI Ready training for your team.
How do you get your business or newsroom AI ready?
Visit our new platforms: Grounded (for newsrooms) and BE AI READY (for businesses). We have daily AI legal and regulation news coupled with our AI tracker. And if you have been on one of our Develop AI trainings it is where you can access AI tools and get advice on how to build your AI policy for free.
See you next week. Cheers.
Develop Al is an AI consulting company that has trained and worked with 100s of organisations globally so they can effectively implement AI and develop ethical AI prototypes and polices.
Contact Develop AI to arrange an AI training (online and in person) for you and your team. And ask about our mentoring so your business can build efficient AI workflows.
We have implemented AI strategies for Thomson Reuters Foundation, DW Akademie, Public Media Alliance, IMS (International Media Support), Agence Française de Développement and others to improve the ethical impact of AI globally.
Email me directly on paul@developai.co.za, find me on LinkedIn or chat to me on our WhatsApp Community.





