Think Different. Think AI. Transcript archive

Episode 9 · Article on the episode

250 documents for a backdoor: automation bias and the limits of trust

Around 250 deliberately prepared documents are enough, according to a study, to teach a language model a backdoor. What is remarkable about it is that the number does not grow with the size of the model.

By Mark Zimmermann · 13 Oct 2025 · 4 min read · Auf Deutsch lesen

The lens in this episode is the Borg collective from Star Trek, and the analogy carries further than the allusion. What has happened technologically since the GPT moment can be described as democratised access to collective knowledge and collective capability.

Via MCP and new tool interfaces that has by now become world logic rather than mere world knowledge: the model does not only know something, it can execute something.

The finding on poisoning

The referenced study on poisoning attacks against language models names a number that is alarming in its plainness: around 250 deliberately prepared documents are enough to teach a model a backdoor.

Added to this is an observation that is less technical and at least as relevant: models answer with differing degrees of territoriality depending on their origin. What a system treats as delicate and what it does not depends on where it was trained.

Automation bias

The second major strand of the episode concerns a habit, not the technology. The better the operation of an agent system, the more rarely the human intervenes.

That is the classic automation bias from aviation research: people adopt the suggestions of automated systems more often than their own judgements, and the readiness to do so rises the more reliable the system has been so far.

In the episode it is called soft assimilation, and the term lands: no coercion, but gradual habituation. The parallel to the voluntary surrender of data with loyalty cards and social networks is close at hand.

For practice the point is more concrete than it sounds. A system that is right in 95 out of 100 cases is more dangerous than one that is right in 70 cases, because with the first nobody looks any more. Checking routines therefore have to remain in place regardless of how good the hit rate has been so far. Spot checks without a specific occasion are the only thing that helps against this effect.

The end of the interface

The technical closing section concerns input. Voice-based interfaces such as Wispr Flow are displacing keyboard and mouse in companies, and the numbers are clear: from 20 per cent voice usage in the first month to 78 per cent in the fifth.

This curve is the actual finding. A tripling within five months describes not curiosity but a changeover of habits. Anyone introducing tools should know curves like this: the resistance in month one says little about the usage in month five.

The next stage would be brain-computer interfaces such as Neuralink, discussed in the episode as a possible genuine end of the interface. The opportunities for people with impairments are real and undisputed. So are the unresolved questions: update cycles, avenues of attack and a deeper form of dependency.

Conclusion

The Borg analogy leads to three points, all of which hold independently of Star Trek.

First: a model is only as trustworthy as its origin, and 250 documents are enough to undermine that. For security-critical tasks, who trained it and with what is what counts.

Second: the better a system runs, the less checking is done. Build in spot checks that take place independently of the hit rate, otherwise control disappears at precisely the moment it is needed.

Third: the changeover of input channels is running faster than rollout projects plan for. In some environments voice is already the main form of text entry.

Resistance is not futile. It simply has to be built in rather than expected to arise spontaneously.