The pen is mightier than the sword.

Apparently, Shakespeare didn’t say that, Edward Bulwer-Lytton did — perhaps updating the Great Bard’s rather less accessible claim in Hamlet that “many wearing rapiers are afraid of goose-quills.”

And how do I know that my seemingly confident cultural touchpoint was incorrect? Because I mentioned it to a Large Language Model (LLM) as part of some unrelated research — and it — rather snippily — corrected me.

But that aphorism about words as weapons — no matter who wrote it — was thrust violently into mind recently following Anthropic’s announcement that it will begin encoding watermarking patterns into AI-generated text to track whether Claude was used to create — or assist with — a piece of writing.

Which apparently means that, from now on, a machine with an internal objective wholly unrelated to the author’s communicative purpose is going to participate in deciding which words get included in Claude’s output — including during editing and proofreading of an author’s own content.

I mean this may be nothing — and the people building these systems may be entirely right to claim that the resulting changes will be “imperceptible.”

“Imperceptible.” Hmm.

But unfortunately, I am a pedant — and also someone who is deeply suspicious of anyone confident enough to judge the cultural resonance of apparently ‘equivalent’ words without knowing the context in which those specific words were meant to land.

And while there is a regulatory backstory to this — which means that Anthropic is not alone in taking this step — it isn’t the intent of the regulation per se that horrifies me.

It’s what happens to weapons when someone strips them of their edge.

The (s)word and the pen

When you look back through human history it’s inarguable that the sword was pretty mighty.

Battles. Slaughter. Empire. Regicide. That kind of thing.

So for Bulwer-Lytton to come along and say, “Hold my coat” was pretty ballsy. But on reflection it’s clear he was right — and that the pen really is mightier than the sword.

Because while it sounds faintly ludicrous on the surface — I mean I dare you to face a murderous sword-wielding psychopath with a pen — words, when honed, polished and sharpened to a keen point, can change minds and influence behavior far more quickly than a man with a sword can bully them into submission.

But more importantly, words belong to everyone — and so can scale influence in ways that men with swords simply can’t.

Ideas can travel from people, through people, and across people very quickly, and — under the right circumstances — trigger huge changes that would be impossible to achieve through force alone.

And so throughout human history, battles with words have been just as brutal as battles with swords — and the history of social and political change is inseparable from the stories that governments, religions, political movements, artists, scientists and cultural upstarts used to bend the arc of history.

But for that to happen, you have to choose your words very carefully, deliberately and precisely. Within your specific cultural, historical, political and religious context. For your specific audience. For the actions you want people to take.

For the future you want to create.

And I feel this pain every day in my far less consequential work — sometimes spending hours, even days, honing and sharpening a single word.

The word that doesn’t just say the thing in an ‘acceptable’ way — but in a way that unlocks whole hidden worlds of meaning through the invisible layers of culture, history, horror, humor and identity that are carried by that specific word in this specific context.

The word, in short, that makes people feel it.

Why words have always mattered

This is why, throughout the centuries, there have always been those who have striven to control words. Religions, oppressive regimes and overly enthusiastic governments have all tried, with various degrees of success, to police the information space.

Which kind of proves the words as weapons point more eloquently than any grand theory.

And so the authority to decide which differences matter — which words are good words, which words are equivalent, and which can be discarded — has been contested for centuries. Not least because, let’s be honest, there can never be a single measure of linguistic quality or equivalence that covers all possible circumstances.

The meaning of a given word, or collection of words, is a deeply personal, cultural, philosophical and historic phenomenon — as varied in the way it is felt as humanity itself.

Because words can carry a multiplicity of hidden layers — connotations seared into them by the particular bloody history of conflicts, beliefs, triumphs, humiliations and upheavals that have carved deep furrows of place, identity and culture.

And I’m Welsh — so believe me, I know.

Which means that writing well is not simply about finding words that are correct — or choosing carelessly between exact semantic equivalents.

It is about choosing, again and again, between words that may look broadly equivalent but carry very different weight for the particular people, purpose and moment in which they are going to land.

The sword and the whetstone

And for most of human history, the battle of ideas was dominated by people with the talent, time or money to spend honing that craft.

But then came LLMs to produce writing for everyone. To level the field. To make literary gods of us all by generating tablets of wisdom we could hand down from the mountain at the push of a button.

Or so we thought.

But what LLMs actually brought is what I call “the verbiage”.

“The verbiage” is portentous. It sounds initially convincing. It looks superficially polished — at least to people with no understanding of the topic at hand. And on first reading it seems as if there’s depth. Substance.

But then you finish reading it and think, ‘huh?’ — and realize that the verbiage is actually a vast façade of words with nothing to say.

By generic measures of text quality, however, the verbiage might look good. Well written, fluent and touching somehow on the topics it’s pointed at. It might even have a couple of interesting ideas buried in it. But when measured against its actual purpose — to communicate a thought, develop an unfolding argument or speak to a specific audience — it frequently falls flat.

Because its polished surface hides just how shallow it is.

But that first flush of disappointment — the one that causes many people to try AI once before shrugging at the facile poem they’ve generated in the style of Wordsworth and walking away — obscured a much more interesting truth.

That if you want to wield words as weapons, LLMs are not the sword.

They are the whetstone.

A whetstone against which you can rub your thoughts and ideas again and again to clarify them, to sharpen them, and to weaponize them.

Which is why I believe that LLMs can simultaneously be a verbiage-generating slop menace in the hands of some — but a powerful whetstone in the hands of others. People with interesting thoughts that deserve to be heard — but whose lack of linguistic craft leaves them unable to articulate those thoughts with enough power to move others.

People with ideas who want to get into the fight.

And just when it seemed like AI might help them — by telling them who actually said what about swords and pens — Anthropic came along and said, “Hold my coat.”

Changing what — and who — words are for

The awesome thing about using an LLM as a whetstone is that the purpose of the exercise is entirely yours. You can tell it what you want to achieve, share your ideas and drafts, and ask it to help you develop and improve them.

Because at every point of the process the intent is yours — and the machine simply does its best to generate tokens that best match its understanding of that intent.

You choose the context, you choose the words, you hone the weapons.

But what happens when another purpose is inserted into the process of choosing the words?

Anthropic’s release created a stir because it will place invisible watermarks into AI text — effectively by biasing token selection during generation so that some words are preferred over others and the resulting text carries a recognizable statistical pattern.

Which changes what — and who — those words are for.

And this is the part that horrifies me.

Because now words are no longer only being chosen to carry thoughts as effectively as possible for the author and audience, but also for the technocratic purpose of provenance tracking. In effect, a technocratic membrane has been inserted between writer and audience for reasons that have nothing to do with the needs of either — one which could weaken the force of the argument meant to connect them, bind them, and move them.

Because, in my opinion, no system operating outside that human exchange can possibly know the specific weight particular words might carry in any particular context.

Anthropic claims that the resulting text retains the same quality — that the watermark only exploits “low stakes” choices between plausible words whose effect on the generated content is “imperceptible.”

Which I guess sounds pretty reassuring, right?

But describing those choices as “low stakes,” or the resulting differences as “imperceptible,” means somebody has already decided which differences between words are insignificant enough not to matter. That some words are simply ‘good enough’.

But for people who want to use words as weapons — who want to savage the existing frame — words that are ‘good enough’ are rarely good enough.

Who decides what’s good enough anyway?

To be fair to Anthropic, it isn’t setting itself up as the lone master of the universe here, the one true arbiter of which words matter and which words don’t.

In fact its approach is based on Google’s open source SynthID-Text — a technique that, to give Google its due, has already been subjected to serious testing by analyzing user feedback across 20 million Gemini responses, and by running a small-scale experiment to rate the grammaticality, coherence, relevance, correctness, helpfulness and overall quality of watermarked text.

Which is reassuring. But also freaks me out.

Because Silicon Valley doesn’t have a great track record of enriching our lives by embracing the complexity of human experience. Instead it has frequently reduced complex human experience to blunt metrics — such as reach, connection, influence or engagement — which go on to cause misery and dysfunction because it’s easier for algorithms to optimize for the proxy than for the messy human reality it seeks to reduce to numbers.

Treating ad clicks as ‘interest’. Encouraging outrage as ‘engagement’. Or pushing volume of slop as ‘influence’.

And while a thumbs up or down can tell you whether a large population rated watermarked text as worse on average, it tells you very little about whether an alternative would have been better — or whether a particular word would have moved a particular person in the way a writer intended. And the measures in the smaller-scale test — while suggestive — are testing the general quality of the text, not the particular meaning those words might carry for the people reading them.

So I could easily imagine that either kind of measure would rate the verbiage as pretty good. Because the verbiage is nothing if not grammatical, coherent, (mostly) correct and earnestly helpful — at least in its superficial way.

But measured against the specific needs of the writer, for that specific communication, in that specific context, it can also be a bit …crap. In ways that are hard to capture with generic measures.

(I do freely admit, however, that the verbiage often succeeds in triggering the full gamut of human emotion in me — from laughter and tears to unbridled rage — but usually not for the right reasons.)

Because as someone who grew up in a small country, I am painfully conscious that a single word, deployed in the wrong way, in the wrong place, at the wrong moment, can be the difference between being bought a beer — or punched in the face.

In Wales every village is somehow a neighbor, competitor and mortal enemy at the same time. Who fought whom. Who beat whom. Who betrayed whom. Whose grandfather hated whose grandfather. What happened in the mine, the chapel, the rugby club, the strike or the war.

And always the bloody English.

But even in a country as small as Wales, the deep grooves carved into particular words in one place might be imperceptible in another village ten miles away — so good luck trying to judge equivalence from California.

And so show me the quality metric that knows all of that — for everyone, everywhere, all at once — and can genuinely tell me that a change in word is “low-stakes,” “imperceptible” or has no effect on “quality” — and I’ll take you to a pub I know to test that out.

Which also brings me back to the membrane — the mechanism that allows outside goals belonging to neither the author nor the audience to mediate their conversation.

Because let’s be honest — algorithms already mediate which communications we see, with …imperfect results. And so do we really want to let them mediate the production of our words as well?

But, to be fair, I’m also conscious that I’m kind of a wordcel.

And so maybe none of that ever happens. Maybe the changes really will be imperceptible. Maybe the quality measures meant to demonstrate imperceptibility in the small won’t become targets that cause billions of such changes to narrow the shape of discourse in the aggregate. And maybe, in the end, nobody will care enough to be as bothered as I am.

Perhaps because I really believe that the pen is mightier than the sword — but only because particular words, in particular places, have enough of an edge to cut into the rubbery hide of convention.

And so I’m just not convinced that anyone should have the power to decide which of the edges that make those cuts possible are “low stakes” enough to blunt.