Where does the human end and the machine begin?

How I use AI, why I disclose it, and why I am disabling Substack’s AI detector

Particular of Galileo before the Holy Office at the Vatican, Joseph-Nicolas Robert-Fleury, 1847. Who judges, with which evidence and who bears the burden of answering? Source: Wikipedia

On July 21, 2026, Substack announced a partnership with Pangram and introduced a tool that allows any reader to scan posts, Notes, comments and replies to “assess” if the text was AI-generated or AI-assisted.

My first reaction, summarized in this Note, was cautious and not particularly negative.

I am somehow familiar with the research showing that many AI detectors perform poorly outside the conditions in which they were developed (e.g., you can read this and this). Therefore, as usual, I read Pangram’s explanation of its own method before reaching a final conclusion. As far as I understand, Pangram is a classifier: it examines a large number of patterns across a text and estimates whether those patterns resemble human or machine writing, without relying only on the predictability measures used by some earlier detectors.

Pangram also describes its output as statistical, which seemed important to stress: after all, in textbook machine learning a classifier can at best assign a probability to a text, but it cannot reconstruct the process that produced it.

Initially, I could support the experiment if Substack explained these limitations clearly, and a posteriori I think their communication has not been adequate, given the reaction of thousands of users. On the one hand, I also understood the concern behind the adoption of Pangram: many platforms are filling with serially produced content, automated comments, invented references and articles written by people who use AI to simulate knowledge they do not have. I have been writing about this for a while:

I still consider those practices a serious problem, as I have demonstrated with a simplified simulator, a few posts ago:

However, my revised position concerns the way Substack has deployed the detector.


What the score can tell a reader

The presence of AI does not prove the absence of a human. — Monica Hebert

According to Substack’s own announcement, Pangram can estimate whether AI contributed to a text, also acknowledging that the detector cannot determine how much human care went into producing it. In fact, the page describes the result of a scan as a percentage of text estimated to be human-written or AI-assisted.

What is the problem with that simple statement? Well, it leaves out most of the information needed to assess intellectual work.

The detector cannot establish who chose the topic, formulated the questions, developed the argument, read the literature, verified the sources and references, interpreted the evidence, or decided what belonged in the final narrative. Pangram was not designed to determine whether a statement is correct, whether an author understands the subject or whether there was an intention to deceive, fair enough. It is rather disappointing that Substack has deployed it as a response to a problem defined largely in those terms.

Let us consider two examples:

  1. An author can write an entirely human-created article containing fabricated claims about vaccines or climate change. Pangram may classify it as human, correctly answering the narrow question it was designed to answer.

  2. Another author can spend several weeks reading scientific papers, developing an original argument, checking every reference, and then use an LLM to improve the English. Pangram may detect AI assistance without providing any information about the quality, originality or reliability of the work.

Readers, in general, are interested in competence, accuracy, effort, originality and honesty, while a text classifier measures linguistic patterns that may sometimes correlate with one part of the production process. The distance between these two things should be visible in the interface whereas, at present, it is too easy to overlook.
It is even more concerning that such a text classifier might be fooled, in a relatively easy way, by means of “humanizers”, such as this one or this one.


The cost is distributed badly

Sam Illingworth described the feature as a machine of accusation, and I think this formulation identifies the central problem with the design of the tool in the context of a large socio-technical ecosystem like Substack: while a reader can obtain a number in a few seconds, the author may then have to explain a text that required days or weeks of work.

From the author side, a suspicious score can prompt requests for drafts, version histories, notes, prompts, or detailed accounts of the writing process. Conversely, the reader who initiated the scan has no corresponding obligation to understand the method, examine its uncertainty or correct the accusation if it spreads. Furthermore, a single numerical output also creates a kind of anchoring effect: once people see something like “70% AI”, later explanations are always interpreted in relation to that number.

Accordingly, the asymmetry is practical and reputational: generating suspicion is cheap, while correcting an erroneous inference requires time and evidence from the person being judged. And I don’t even want to discuss the cases where generating such a suspicion can be done on purpose to damage one’s reputation and redirect the attention of readers elsewhere, because that would be an easily predictable practice within the jungle rules of our attention economy.

Uhm, yes: Substack allows authors to disable detection on individual publications and Notes. But this legit choice on the author’s side will lead (some) readers to see that AI detection is unavailable: the absence of a result can itself be interpreted as suspicious by those readers who already assume that the scan reveals something important.

Madeleine Clare Elish introduced the concept of a “moral crumple zone” to describe those situations in which a person with limited control over an automated system absorbs the responsibility when the system produces harm. I know that the analogy is not exact, but the distribution of responsibility is similar: Substack selects the system and designs its use, Pangram produces the estimate, the reader activates it and the author bears most of the consequences. Read in this order, it does not sound really good, does it?

I think that Substack cannot transfer accountability for this arrangement to Pangram or to its users. While choosing an external provider may be reasonable, if the required expertise is proven, the platform should remain responsible for the context in which the provider’s output acquires meaning.


What the available evidence establishes

At this stage of the story, I do not think it is accurate to dismiss Pangram as a useless detector.

An independent evaluation by researchers at the University of Chicago found that Pangram performed very well on their controlled dataset. Its false-positive rate was close to zero across most tested thresholds, although performance varied with text length, model and calibration, as usual in that field. More interestingly, the authors also argued that institutions should define an acceptable false-positive rate before using a detector and should audit performance regularly.

While the results can be encouraging for Pangram, they do not settle the question of whether a reader-facing score is an appropriate feature for a social publishing platform.

It is evident that controlled evaluations usually distinguish text generated by humans from text generated by models, but the point is that much contemporary writing sits between those categories. Human authors edit machine-generated text, LLMs polish human prose and several rounds of human and machine revision may occur before a single publication.

I ask: where does the human end and the machine begin? 1

In a recent study, Saha and Feizi examined this problem (see Almost AI, Almost Human). They tested twelve detectors on 15,000 samples with different levels of AI polishing: the detectors frequently classified minimally polished human text as AI-generated and struggled to distinguish different levels of AI involvement.

Other evaluations show that detector performance can decline when the domain, model or prompting strategy differs from the training conditions. Remarkably, moderate rewriting can also help generated text to evade several detection systems.

While these studies do not necessarily evaluate Pangram — so they should not be used as direct evidence against it — they do show why deployment conditions and adversarial adaptation require separate testing 2.

The evidence about linguistic and demographic bias also requires some care. A 2026 ACL study of sixteen English-language detectors found that essays by English-language learners were more likely to be classified as machine-generated in several systems. A different 2026 study conducted in Czech found no systematic bias against non-native speakers. These findings are compatible: bias can depend on the detector, language, population, and type of text.

I have not seen a public, independent evaluation showing how Pangram performs on English prose written by non-native academics and subsequently polished through iterative human–AI editing. Accordingly, until such tests exist, assurances about fairness in this specific use case remain incomplete, to say the least.


How I use AI

I am a native Italian speaker who publishes primarily in English since 2009. Despite the scientific papers and textbooks I have written so far3, and well before LLMs were at best a cyberpunk dream, writing in English is still a challenge.

While I can perfectly know what I want to say, the structure of an English sentence does not always follow the structure through which I developed the thought in Italian, and I do not always catch every difference in meaning and nuance between Italian and English 4.

I use LLMs to identify awkward constructions in my own paragraphs, as well as to clarify passages, reduce ambiguity and improve the accessibility of complex arguments. Sometimes I ask a model to reorganize a paragraph and I then decide whether the result preserves the intended meaning and my usual voice. If it does not, I will not use it: rather, I will work on it until it meets my expectation.

I also use LLMs to brainstorm and challenge an argument, identify possible objections and locate scientific literature that I may have missed5. In every single case, I check the suggested papers against the original publications: claims are verified independently, references are validated one by one and the ones that do not exist are discarded, as are plausible sentences that fail under scrutiny. It takes several hours, and I learn throughout the whole process.

The questions, selection of evidence, scientific interpretation, narrative structure and conclusions remain under my full control. I review the final text and accept responsibility for every statement published under my name.

Most notably, I already disclose AI use when it makes a substantial contribution to the final product.

The case of podcasts

For instance, the podcasts associated with #ComplexityThoughts are explicitly labelled as AI-generated. Why? Because without AI, they would probably not exist at all: I am a father, a husband, and have a full-time position as a university professor and a research Lab to lead. I prepare this during my free time, I do not have a recording studio or any suitable equipment at home, and I cannot find another person to prepare and record an episode with me every one or two weeks.

However, I’d like to stress that even in this case the use of AI does not reduce the production process to a single click. I listen to each episode, check its content, edit it when possible and necessary, and sometimes produce it again from the beginning because the first result is not adequate. This work takes time, although much less time and money than writing a complete script, organizing a recording session, recording the audio and editing it conventionally.

Readers regularly tell me that they appreciate the podcast, since it provides them with another way to engage with material that already exists in written form. Many of them do not even know how to transform those texts into a podcast and, even more importantly, they don’t know if they can trust the output. It’s exactly there where my role makes sense: the output is not published on the authority of an AI system, but under my responsibility. If readers trust me as a scientist, it is because they have seen the method behind my work: I verify the content, correct it when necessary and remain answerable for the final result.

Therefore, the relevant comparison is between an AI-assisted podcast and no podcast, rather than between an AI-assisted podcast and a professionally produced alternative that I do not have the resources to create.

Illustrations

I apply a similar approach to illustrations, slides and infographics. If you are a scientist and have attended one of my talks, seminars, keynotes at our flagship conferences, you know how obsessed I am with nice graphics. I think that communicating scientific results is crucial, and since I was a PhD candidate I have invested a considerable amount of time in it6.
Many of these images represent concepts that often have no direct visual counterpart — e.g., an AI system at the centre of a medieval tribunal — and their function is explanatory.

I use mainstream commercial tools that are legally available in the European Union, and I do not prompt them to imitate identifiable artists or reproduce existing works. Of course, I cannot independently inspect their complete training data or guarantee that every output is free from problematic similarities: for this reason, when needed, I disclose substantial AI generation and remain available to investigate credible concerns. If an image reproduces or closely derives from someone else’s work, I will remove or replace it and acknowledge the error. I have no problem with this process.

My current disclosure rule is based on relevance to the reader: I disclose AI-generated podcasts, illustrations, infographics and substantial generated text, if any. I do not provide a list of every grammar correction, brainstorming exchange, search query or discarded output, just as I do not document every conversation with a colleague or every operation performed with conventional editing software.


Why a generalized anti-AI response is a mistake

I am certainly not a Silicon Valley tech bro, and I have no confidence in the idea that technological development automatically produces social progress.

At the same time, as a scientist, as a human, I also see no reason to reject AI as a category. AI can improve access to translation, education, scientific information, communication tools and forms of media production that require resources many individuals do not have. Its benefits depend on who controls the systems, which data they use, how errors are handled, whose work is displaced or appropriated, and who receives the economic gains.

A public detector score compresses these different uses into a simplified scale. Assistance for a non-native writer, automated spam, an AI-generated podcast, plagiarism, and the serial production of articles can all contribute to an “AI” percentage, although they raise different ethical questions.

I expect this design to amplify the anti-AI reaction already visible on many platforms, including Substack. The score gives that reaction an apparently objective instrument, even when the instrument cannot identify the author’s method or intention. Consequently, authors who use AI openly and carefully are encouraged to enter a permanent defensive mode that can be exhausting, in the long term.

I also guess that less careful users will adapt in another direction. Public detection creates a market for tools that rewrite generated text until “it passes as human”. In fact, research has already shown that paraphrasing can evade many detectors, while the Chicago evaluation linked earlier describes an ongoing technical arms race among generators, humanizers and detectors.

I am a complexity scientist and my job is to study emergent phenomena, especially when small changes to microscopic actions can lead to huge changes in collective behavior. I think that this process is going to produce an unfortunate incentive: honest authors may modify their writing to avoid suspicion, whereas “industrial producers” will use their resources to optimize their pipelines to obtain a reassuring score.
The detector ends up influencing the behaviour it is supposed to measure, reducing the value of the score over time.

When a measure becomes a target, it ceases to be a good measure. — Charles Goodhart

What I think Substack should address

Substack has problems that cannot be diagnosed by inspecting linguistic style alone.

These include authors publishing implausible volumes of material at an implausible rate, automated accounts posting comments without even reading the articles7, plagiarism, fabricated sources, as well as the uncontrolled production of medical and scientific misinformation. Human beings are fully capable of producing all of these without assistance from an LLM: with AI, it just became cheaper and faster than ever.

Linda Caroll, who apparently takes a partially favorable view of the detector, argues that Substack should combine textual analysis with user behaviour: publication frequency, reading time, comment volume and other activity patterns may help identify systematic abuse. I agree that behaviour is relevant: I would add that consequential decisions should involve clear rules, multiple sources of evidence, human review, as well as an appeal process.

These measures would require Substack to accept direct responsibility for moderation and enforcement, while a button offered to readers places much of that work elsewhere.

The economic context should also be discussed, since Substack receives 10% of each paid transaction. Of course, this does not mean — or let alone prove — that the company deliberately promotes misinformation or low-quality content. However, it creates a structural conflict of incentives whenever content that generates paid subscriptions also creates costs for the information environment.

Our year-long collective experiment on Substack visibility adds another piece of evidence. Since I am an academic and I believe in free knowledge for everyone, #ComplexityThoughts was designed to be entirely free before July 31, 2025, since November 2022.
However, after enabling a paid tier, while keeping the publication’s content accessible, I estimated a visibility advantage corresponding to approximately two-to-three additional followers or subscribers per day compared with the free-only trajectory.

The experiment concerns only my Substack and its specific conditions, it does not identify Substack’s intentions or prove a general causal mechanism across the platform. Nevertheless, it does support closer scrutiny of the relationship among monetization, ranking, recommendations and visibility, which is important to know if you are competing in the attention market.

So, while Substack has made the “authorship detector” visible to users, the systems that decide which publications receive visibility remain much less transparent.

Let me draw a limited comparison with surveillance.

I am obviously not equating a Substack scan with state surveillance in terms of scale, coercion or potential harm. The comparison concerns the logic used to justify automated scrutiny: the “nothing to hide” argument, the possibility of function creep and the burden placed on individuals when an automated inference is wrong. I have many questions: How will the collected information be interpreted? Who is going to access it and how? How is Substack monitoring whether its use expands beyond the original purpose?

An experiment on online surveillance and chilling effects found that awareness of surveillance can suppress lawful political behaviour as well as illegal activity: observation changes behaviour before any sanction is imposed.

A public AI detector may produce a smaller and context-specific of this effect in writing: authors can begin avoiding legitimate editing tools, repeatedly scanning their drafts, removing unusual phrases, or optimizing prose for a classifier to “sound more human”. The resulting text may receive a better score without becoming more accurate, original, useful, or let even readable, which is the main relevant goal at the end of the day.

Automated scrutiny needs a defined purpose, as well as evidence that it measures the relevant behaviour, safeguards against misuse and a realistic way to contest errors. Introducing the tool first and leaving users to negotiate its social meaning afterwards reverses this order and provides a perfect recipe for chaos.


My policy

I will continue to use AI where it improves my ability to communicate or produce useful formats within the time and resources available to me.

I will continue to disclose AI-generated podcasts, illustrations, infographics and substantial generated text, if any. I will continue to

ask the questions → verify claims and references → correct errors → investigate credible concerns → remain responsible for everything published under my name.

I will not treat Pangram’s score as evidence of competence, accuracy, originality, care or deception. Where Substack allows it, I will disable reader-facing detection on my work and, accordingly, I won’t use it on other authors either.

This decision does not reduce my transparency: the purpose of this whole page is to explain my process in enough detail for readers to evaluate it. Readers may question my evidence, challenge my conclusions, identify errors and ask how a particular item was produced. Members of the #ComplexityThoughts community can reasonably expect direct answers when the method affects their interpretation of the work.

What I firmly reject is the expectation that every author must prepare a preventive defence whenever an anonymous or occasional reader activates a classifier.

A useful accountability system identifies who made a decision, what evidence supported it and how an error can be corrected. The Substack implementation leaves these elements unclear while placing the most visible burden on the author.

This statement provides my readers with a method they can inspect and a person they can question. A detector score cannot provide either.

1

I googled around and found the exact same question in other posts and pages. Of course, it is a legit question nowadays. I really liked this post built around an anime I really like: Ghost in the shell.

2

This NAACL 2025 study provides one useful analysis of these generalization problems.

3
Fun fact: since the release of ChatGPT in November 2022, I have been publishing fewer papers than before. Apparently, my scientific output has decreased. Another possible interpretation is that I have been publishing more rubbish in journals not indexed by Scopus, but that is not really the case. The 2026 count is year to date. Source: Scopus
4

For instance, I have only learned recently that my use of the word “pretend” in English is wrong.

5

Trivia: I keep track of all the papers I read, since 2008 when I started my PhD. From Mendeley to Zotero and Dropbox Paper, I have used several platforms for this purpose. Eventually, I can count more than 2,000 scientific papers. I love to read papers and short essays, and I really like to discover and organize those pieces of knowledge in a way that makes sense to me and allows me to connect the dots. Still, there are literally hundreds of millions of scholarly papers out there produced by humans, and I welcome the support of whatever algorithm allows me to find what I need. That’s not different at all from you following Amazon’s internal recommendations to buy a new product.

6

The fact that I married an engineer who is also a graphic designer helped a lot to improve my skills for communicating science.

7

I guess that this is widespread because it can influence Substack’s algorithm and give visibility to the commenter.