We’re independent. Help us out and become a member too! We’ll be better off with you on board.

Not a member, but want to support us?

Soon we’ll have to prove we’re human: a dialogue on journalism and AI

S

With the advent of AI, and AI checkers in particular, the writing profession is changing. What does checking for the use of artificial intelligence mean for writers, but also for journalists? Guido van Nispen submitted an article on the subject; Wijbrand Schaap took it to the editorial team and had a few questions. We felt the resulting email conversation was worth sharing with our readers. 

Guido van Nispen’s contribution:

A novel is praised, an opinion piece is persuasive and a work of art moves us – until someone runs an AI detector over it and attaches a percentage to it. Nothing about the work itself has changed, but our perception of it has all the more so. This creates a curious new burden of proof in the cultural sector, for whilst for years we were concerned about machines that could convincingly pass themselves off as humans, we are now entering a world in which humans may have to prove that they are, in fact, human.

A novel that suddenly took a different turn

The debut novel *C’était ça ou mourir* by the Haitian-Canadian writer Thélyson Orélien initially appeared to be, above all, a remarkable literary success. In five weeks, 35,000 copies were sold in France; the book won the Fnac Novel Prize and was shortlisted for several major French literary awards. Critics had read it, booksellers had recommended it, juries had assessed it and readers had bought it, meaning the book had gone through exactly what we in the cultural world generally regard as a fairly thorough process of evaluation and selection.

That changed when an anonymous account on X claimed that the book had been written almost entirely by artificial intelligence. As evidence, reference was made to Pangram, an AI detector that attempts to determine whether a text has been produced by a human or by a language model. Le Monde decided to investigate the claim and fed more than half of the novel into Pangram, after which the detector classified more than 95 per cent of the material examined as AI-written.

This is not without significance, as Pangram is one of the better-performing AI detectors and also fared relatively well in two academic studies cited by *Le Monde*. At the same time, the system is largely a black box to outsiders; the company was reluctant to tell the newspaper much about exactly how it works, and researchers warn that performance may vary between languages and particular styles of writing. This leaves just enough uncertainty to create an uncomfortable situation: the percentage is too high to simply ignore, but the method is not transparent enough to be regarded as conclusive evidence.

What is interesting, then, is what effect such a percentage has on our perception. After all, nothing has changed in the novel, and the same applies to the words on which readers, booksellers and jury members based their earlier judgements. The only thing that has been added is new information, provided by a machine that claims to know something about the way in which those words came into being. From that moment on, we are no longer just reading the novel, but also the suspicion.

Then there was something else to check

That distinction is important, because not every discussion about the use of AI is the same. Earlier last year, we saw a very different case involving Peter Vandermeersch, former editor-in-chief of NRC, who had used AI to process and summarise sources for his publications. Upon closer inspection, it emerged that several articles contained quotations that could not be traced back to the original sources. Vandermeersch subsequently acknowledged that he had used AI and that the necessary human oversight had been insufficient.

This therefore highlighted a journalistic issue that existed long before ChatGPT and for which we already have established journalistic standards. When someone is quoted in inverted commas, that person must actually have spoken or written those words. If that is not the case, the publication is inaccurate, and ultimately it matters little for the assessment whether the error was caused by a language model, a careless note or a journalist who failed to check their sources thoroughly. The content itself provided the evidence of the problem.

In Orélien’s case, something fundamentally different is happening. No one first had to find a fictional character, a plagiarised passage or a demonstrable factual inaccuracy in his novel. The suspicion arose because a machine examined a highly regarded cultural work and concluded that it had probably been created by another machine. This shifts the discussion from what is demonstrably wrong with a work to the question of whether we can still trust its genesis.

A percentage next to an opinion piece

The Netherlands has now had a taste of this development for itself. This summer, AI Report examined 252 opinion pieces submitted to De Volkskrant, NRC, Trouw, Het Parool and the FD using Pangram. According to the detector, 49 pieces – almost a fifth of the total – had been generated by AI. The differences between the newspapers were striking: at *De Volkskrant*, a significant proportion of the pieces examined were classified as AI-generated, whilst at *NRC*, not a single opinion piece examined received that classification. When texts in which, according to Pangram, AI had been used to some extent were also included, the proportion rose further.

Such findings are, of course, relevant to editorial teams. A newspaper is entitled to set conditions for letters to the editor, and readers are entitled to expect that an editorial team knows what it is publishing, particularly when a contribution is presented as an author’s personal opinion piece. From that perspective, it is understandable that de Volkskrant has tightened its checks in the wake of the investigation. It becomes more interesting when the detector itself becomes part of that scrutiny, because then statistical probability plays a role in an editorial judgement that previously revolved primarily around content, originality, reliability and quality.

This gives rise to the same strange situation as with the French novel. A reader might find an opinion piece interesting and convincing, an editor might select it because it contributes to the public debate, and the text could be published without any problems. However, when Pangram then assigns it a high AI score, the judgement of those very same words changes. The detector has not altered the text in any way, but it has superimposed a new layer of meaning upon it: perhaps this is not what you thought it was.

When the tool becomes a referee

For the cultural sector, this touches on a far more fundamental issue than the now somewhat predictable debate over whether writers, artists and other creators should or should not be allowed to use generative AI. Creators have always used tools, and the boundaries between creating, editing and supporting have never been entirely clear-cut. A writer works with an editor, a photographer with image-editing software, a composer with digital instruments and a filmmaker with technology capable of creating images that never existed before the camera. Generative AI makes that toolbox more complex, because the tools can now also write, draw, compose and make suggestions that end up directly in the final product. That is precisely why agreements on transparency, authorship and responsibility are needed.

However, an AI detector introduces yet another problem, as it not only helps to gather information but also suggests a judgement as to authenticity. As soon as such a judgement comes to function as evidence in society, the burden of proof shifts imperceptibly. The person claiming that a work was created by AI is then no longer required to convincingly reconstruct how this happened, whilst the creator must demonstrate that they produced it themselves.

According to his publisher, Orélien began work on his novel as early as 2017, and drafts, notes, earlier versions and correspondence are now being used to bring the book’s genesis to life. Material that once simply formed part of the messy and usually invisible process by which a writer creates a book is thus given a new function as evidence of human authorship. For writers, journalists, artists and other creators, this may ultimately mean that not only must the work itself be preserved, but also the digital archaeology of its creation.

This brings us full circle in a remarkable way. For years, we worried that AI would become so advanced that machines could produce work we could no longer distinguish from human work, whereupon we started building machines that were meant to make that distinction for us after all. Now that these detectors are getting better and better and their accuracy rates are looking increasingly convincing, an almost opposite risk is emerging: that a human creates something which readers appreciate, an editorial team publishes or a jury awards a prize to, only for a machine to then tell us that we’d better reconsider our own judgement.

Perhaps, therefore, the interesting question for the cultural sector is not only how much AI a creator is allowed to use, but also how much authority we are prepared to grant to a machine that, in hindsight, reveals who the creator actually was. For when a well-received novel can be transformed into a suspect work by a single percentage point, the discussion is no longer solely about artificial intelligence. It then also concerns the question of what we still consider to be sufficient proof of humanity.

Whilst writing this article, I used AI for research, translation, verification and final editing, after which Pangram concluded that my text, just like Orélien’s, was also written 100 per cent by AI (other tools put the percentage much lower, and there are also tools to have it rewritten so that it becomes even lower or even supposedly 100% human).  So I’d better keep the loose bits of paper, notes and earlier versions safe, just in case I ever need to prove that I was actually there myself. A messy desk thus becomes, once again, proof of being human.

The question from Wijbrand Schaap:

Hi Guido,

I’m going to mull this over for a bit. The point is: even before you got to that very honest and commendable admission at the end, I was already 90 per cent certain that the text contained a great deal of AI-generated phrasing. And that’s mainly because the style is rather impersonal. I’m used to your usual style, and you’ve mentioned before how you use AI, so I don’t have any major objections to it.

Now I do find myself wondering: in this particular case, wouldn’t it be more interesting to go into more detail about how you’ve used AI? What prompts you give? I’d find that interesting too, for example, in the present case of the writer who may or may not have been justifiably vilified (there’s a racist angle to it as well, just like with that Cambridge lecturer). So that at some point, as the author, you also include the prompts, so that the reader can create their own version if they wish. 🙂

But this is actually a serious matter. Would you be up for it?

Sincerely,

Wijbrand Schaap

Editor-in-Chief, Cultuurpers


Guido van Nispen’s response:

Hi Wijbrand,

Yes, I’m actually quite keen on that, precisely because your comment adds an interesting extra layer to the piece. It’s just that, in my case, there isn’t a single prompt I can give you that would result in this article. Looking back at the process, it’s more of a combination of research, searching, writing, proofreading and editing, with AI playing a different role at different stages.

It starts with the fact that ChatGPT is now familiar with my so-called ‘Guido format’ for final editing. This has evolved over the course of many texts and, above all, many corrections. The format includes things such as: no spelling, language or grammatical errors; no loose rhetorical phrases; a smooth line of reasoning; a calm and precise tone; minimal embellishment; no sales pitch; and not trying to wrap everything up for the reader. This is now embedded as context within the model, so I no longer need to provide it as a single, detailed prompt every time.

I then used my earlier piece on Orélien – which I’d written for LinkedIn – as the substantive starting point for this article, and highlighted the tension I saw in it: the public appreciates the work, but then a detector casts suspicion over it without anything about the work itself having changed. I then used AI to analyse the article from Le Monde to process the information and verify the facts, but also to search specifically for Dutch examples. This brought the earlier case involving Peter Vandermeersch back into the picture; I had actually intended to use it as a contrast, precisely because demonstrable errors and fabricated quotations had been found in it. I then wanted to examine the recent Pangram study of opinion pieces in, amongst others, de Volkskrantand the NRC as well. So I decided for myself which cases should be included, why they should be included, and how they relate to one another.

This was followed by the writing and editing process, and it certainly didn’t go smoothly the first time round. I rewrote the first draft because it contained too many disjointed, rhetorical sentences, which was not at all my style of writing. The next draft went better, after which I worked separately on the conclusion and the introduction and further refined the narrative tension.

Incidentally, in response to your email, I simply asked ChatGPT whether it thought my style itself was very ‘AI-like’. I rather liked the answer, as it highlights precisely the problem I also raise in the article. According to the model, my usual style consists of long, continuous lines of reasoning, minimal rhetorical embellishment and a calm, analytical detachment. These are characteristics that modern LLMs are now also quite capable of producing. According to ChatGPT, there was something else in this text: the structure was exceptionally orderly, the transitions were very neat, and phrases such as ‘This gives rise to…’, ‘That distinction is important…’ and ‘For the cultural sector, this affects…’ structured the argument almost a little too visibly. I am a structured thinker and organise my articles in the same way, and the model apparently knows me very well by now.

I actually find that an interesting observation, not least because your human judgement is far more nuanced than Pangram’s. You’re effectively saying: I recognise phrasing and a smoothness here that strongly remind me of AI. Pangram says: 100 per cent AI. Those are quite different statements. And yes, as you can see, I’m critical of those AI detectors and the way they’re used. I also have no desire to run my own familiar style through a ‘humaniser’ just to confuse the detectors. I write these pieces for the reader, not for the detectors.

To me, therefore, the whole process seems much more like working with a fairly quick researcher and (senior) editor than giving a single instruction to a machine that then writes my article. At the same time, it would be disingenuous to deny that this researcher and editor do, in fact, produce phrases that end up in the final text. This is precisely where the question of authorship becomes interesting.

That is why I cannot simply provide a single, non-existent prompt for publication, but must instead demonstrate the process – perhaps in a separate box alongside the article – from my original idea and instructions, through the research, the searches and my rounds of editing, to the final text and then Pangram’s assessment. Then the reader can decide for themselves where, in their view, authorship lies within that process.

Best regards,

Guido


Wijbrand Schaap’s proposal

Would it be a good idea to publish your first piece, followed by my response and then your reply, as a single post on the site?

Sincerely,

Wijbrand Schaap

Editor-in-Chief, Cultuurpers


Guido van Nispen’s reply:

Great idea!

 

Guido van Nispen

Appreciate this article!

donation
I donate

About this author

Guido van Nispen

Entrepreneur, director, supervisor and adviser at the intersection of technology, media and governance. He is the founder of MENSEN.WORK, a holding company focused on collaboration between people, technology and information. Previously, his roles included CEO of the ANP, director of V-Ventures, fund manager at the Dutch Creative Industry Fund, publisher of *Innovation for Jobs* and executive director at BT. He has also held supervisory and advisory roles at organisations such as Triodos, World Press Photo, Cinekid, the Council for Culture and AVROTROS. His work consistently centres on the question of how organisations can responsibly navigate the changing relationship between technology, information and governance.

Respond!

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Popular posts

This is how theatre becomes a home for everyone

This is how theatre becomes a home for everyone

A podcast about De Harmonie in Leeuwarden. How director Siart Smit is working to create a municipal theatre that also serves as a home for all residents during the day. Plus, plenty about ILFU, books and...
Is asking questions an art form?

Is asking questions an art form?

We can learn a great deal from the latest ‘ZomerGasten’. And from the Boulevard Theatre Festival. And about how we can listen more effectively.
Sometimes art is better than a holiday

Sometimes art is better than a holiday

The shortest ‘Zomergasten’ review ever, and why we saw the light in Groningen. And two petitions. Because they’re desperately needed. Against Hitler, in favour of dialogue.

Categories