Skip to content
← Back to blog
·4 min read·

I Wrote a Linter for Text That Sounds Like AI

Most people can spot AI text within a sentence or two now, and once they've spotted it they stop reading properly. So I built a checker that catches the tells before anything leaves my machine, calibrated on my own writing, and it learns from the edits I make.

I Wrote a Linter for Text That Sounds Like AI

I use AI at work like most developers do. It helps me write code, summarise long threads and get a first pass at the boring documents, like a handover note or a status update after a long week. What goes in the message is still mine. I decide what to say, what the numbers are and what I'm promising, and I read and edit everything before it goes out. The problem was the part in between. Hand a model three bullet points from my notes and ask for a status update, and it comes back sounding like a brochure.

Here's a made-up example, but a close one, from three lines of notes:

Annotated example. Notes: export to next milestone, rest on track, done Thursday probably. Model draft (41 words): "I want to be transparent about what this change is and what it isn't. It isn't a reduction in scope; it's a refinement. The export now lives in the next milestone, which keeps the current phase focused, predictable and on schedule." with three tells marked: announces itself, isn't X it's Y, list of three. What I sent (25 words): "Same scope, I just moved the export to the next milestone. The rest is on track, should be done Thursday if nothing weird turns up."
Same notes, two messages. The marked bits are the tells.

Nothing in the first version is wrong. It's just that no person writes like that to someone they work with, and the reader can tell.

Word lists only get you so far

Everyone knows the word list by now. The long dashes, the curly quotes, the handful of words nobody used before 2023, the "hope this helps" at the end. A regex catches those, and mine blocks them outright.

The harder tells are in the structure, and these took most of the work. Things like:

  • Lists of three where two would do ("focused, predictable and on schedule").
  • "It isn't X, it's Y" as a way to sound decisive.
  • An opener that restates the question and a closer that sums up what you just said.
  • Every paragraph opening by announcing its topic, so the text reads like an outline with prose poured in.
  • Paragraphs all roughly the same size, and sentences all 15 to 25 words long.
  • A hedge on every line, even on things that were checked.

None of these are wrong on their own. I do some of them myself. That turned out to be the interesting problem.

Calibrating against my own writing

The first version flagged everything, including me. So I took 163 messages I'd actually sent and made them a test. Any structural check that fires on more than a quarter of my own messages fails the test suite, because at that point it's measuring how I write, not how a model writes.

That killed two checks I was sure about. I'd assumed people hedge more than models do, and ask more questions back, so one check flagged a longer message with no "I think" anywhere in it and another flagged one with no question in it. On my own messages over a hundred words, the hedge check fired 80% of the time. When I know something, I mostly just say it. Both checks went.

Bar chart of how often each structural check flagged 163 of my own sent messages, with a 25% cut-off line. Cut: no hedge anywhere 36%, no question back 34%. Kept: much longer than the question 20%, paragraphs all the same size 4%, long for a chat reply 3%, assistant opener 1%, pseudo-wisdom 1%.
Two checks crossed the line and were cut.

What stayed is the opposite one, a hedge on nearly every line, which reads just as fake. The rule I write by now is simple. A fact I've actually checked gets stated plainly. A guess gets a personal hedge, the kind a person uses, not "it may be the case that". And if I don't know, I ask a question instead of writing a caveat.

The length check stayed, even though it still fires on about one message in five. I went through those, and they were long replies to short questions that should have been shorter, sent by me or not.

One door in, one door out

Every draft goes through one script before I see it. Hard tells block, soft ones come back as warnings I have to deal with. There's also a check at the send step for each channel, so nothing gets out by a side route.

That second check turned up a bug straight away. The hook was registered for two of the channels, but the script's case statement had no branch for them, so it let everything through without checking. The configuration said "covered" and the code said "skip". Those two channels were where the stiffest text had come from. It was a one-line fix, and I'd never have found it by looking at the config.

It learns from my edits

When I rewrite a draft before sending it, or throw it out because it sounds like a machine, the before and after get logged. Once a night a job reads the day's pairs and writes short style rules from them, things like "when answering feedback, reply only to the points raised and skip the warm sign-off" or "end on a concrete question, not a soft line". The rules load into the next draft automatically.

Flow diagram: my notes, first draft from a model, linter, me (read, rewrite, decide what goes), send check, out. My edits go to a log, a nightly job turns them into style rules (max 6 a night, 25 in total), and the rules are loaded into the next draft.
Every draft passes me before it goes anywhere.

I kept that loop on a short leash. It can add at most six rules a night and keep 25 in total, the rules can only be about style and never about what needs my approval, and the file is plain text I can edit or delete from whenever I like. A model writing its own instructions is a good way to get a mess, so it gets a small box to work in.

What it doesn't do

It doesn't decide what to say. It only cleans up how something is said. Knowing what the other person actually asked, what I can promise, and when one line is the right answer is still my job, and I don't think a tool should do that part.

If your team uses AI to draft support replies, sales emails or reports, and customers have started to notice, this approach carries over. Your own sent messages are the dataset, and the tells are measurable. That's the kind of tooling I like building.

#AI#LLM#Writing#Developer Tools#Python