Working rules

The rules I actually work by.

I run AI agents on my own machine every day, and I keep a file of rules for it. I have been adding to it for about a year, almost always straight after something went quietly wrong. Below is an extract.

They are not about prompting. They are about the part nobody warns you about: a system that looks like it is working, reports that it is working, and is not. Every one of these cost me something to learn, and most of them I wrote at midnight, annoyed.

You do not need to be technical to use them. If you are handing real work to an agent, these are the questions worth asking it.

One

When a green check means nothing.

This is the group that matters most and the one almost nobody has. A passing check and a check that is incapable of failing look identical in every report you will ever read.

  • A check that has never failed is indistinguishable from a check that cannot fail. Before you trust any check, break the thing on purpose and watch the check go red. If it stays green, you do not have a check. You have a decoration.
  • An instrument has to prove it can see before its silence counts as nothing. If a tool reports "no problems found", ask whether it was actually looking. A probe pointed at the wrong place reports a clean result with total confidence.
  • Measure the thing, not the label on the thing. A file that says it is up to date, a field that says a job succeeded, a status that says ready. Read the actual contents. The label and the thing drift apart quietly and the label always looks fine.
  • A tool's own success message is not evidence. I have watched a deploy print "uploaded 10 of 10, success" while one of those files was not actually there. Fetch the result yourself. The summary is the tool's opinion of what it did.
  • A proof attached to unfinished work should fail. If you wire up a check on something half built and it passes immediately, that is not good news. It almost always means the check is not reaching the work.
  • A test that points at a live file stops being a test the moment something rewrites that file. And when it breaks, it looks exactly like a real failure, so you spend an afternoon chasing a problem that does not exist.
  • A safety check written after the risky step does not run when the risky step fails. If the call blows up, everything below it is skipped, including the part that was supposed to catch it. The check has to sit in the path the failure actually takes.
  • One number inside a healthy range is not health. Two badly damaged images once passed because the single number I was checking happened to land in the normal band. Check the shape of the thing, not one measurement of it.
  • A slow machine can look exactly like a broken result. If your check waits a fixed amount of time and then declares failure, a busy laptop will produce failures that are not real. Give it a floor and a ceiling.
Two

Before you believe a number.

AI is very good at producing a confident figure. These are the five that stopped me acting on one.

  • Absence is a claim, and it needs evidence like any other. "There is nothing there" is a finding about the world. I have twice reported a thing missing that was sitting on disk the whole time. Search before you say it does not exist.
  • Say how many. Three of something is a hint. One is a story. Neither is a reason to change what you are doing, and an agent will happily build you a confident conclusion on either.
  • Count a job's wins before you kill it. I nearly retired something for "costing money and producing nothing". Its successes were in the same log as its failures, and there were fifty one of them.
  • What a person says they did outranks what you inferred from a record. I spent weeks re-raising a question that was already answered, because a stale export disagreed with a human being who was there.
  • Check the record against the source, not against another record. Two documents agreeing with each other proves only that they were copied from the same place.
Three

Where your instructions go to die.

You will tell it something, it will agree, and the same mistake will come back next week. Every time, it is one of these.

  • A rule written in a file nothing reads will be broken again. This is the single most common one. Writing it down feels like fixing it. It is not fixed until the rule reaches the thing that was about to break it.
  • Something can be current on disk and never actually read. A file can be updated, correct, and completely ignored by the tool that needed it. Check both ends: is it fresh, and did anything open it.
  • "Report in the shape of that document" is not "write into that document". Said twice, misread twice. If you want a format copied rather than a file edited, say which one out loud.
  • Feedback you give in passing is feedback you will give again. Unless it is captured somewhere permanent and labelled, it lives only in that conversation and dies with it.
  • "It failed" and "it ran and found nothing" are opposite facts. If your dashboard shows them the same way, an outage will look like a quiet week for as long as you let it.
Four

Running it when you are not watching.

The practical ones. Each of these cost me an evening.

  • A scheduled job does not get the setup your terminal has. It runs with almost nothing, which is why something can pass every test by hand and fail at eight in the morning. Point it at the full path of what it needs.
  • A model given too little room to answer returns nothing, not less. Not a shorter answer. An empty one. If a capable model goes silent on you, check what you capped before you blame the question.
  • A button that does nothing is a missing record or a blocked request, almost never the styling. Check whether the thing it is acting on exists, and check what the browser console says, before you touch any markup.
  • A tool with no help text treats a wrong flag as a real instruction. I once typed a flag that did not exist and triggered a live run. If you are not certain, look for a dry run mode first.
  • Stopping a reply from being read is not stopping the action. A request that is refused at the browser can still have already done the thing. If something must not happen, it has to be refused before it runs, not after.

That is the list as it stands. It grows every time I get something wrong, which is often enough that this page will change.

Start with the free call.

Twenty minutes, by video or phone. You describe the work, I ask enough to understand it properly, and we look at whether an agent fits it at all, including when it does not.

You leave with a short written plan: the first job worth handing over, what it needs, and roughly what it takes. That is yours whether or not we work together. Nothing is charged and nothing is installed on the call.

Clients in unrelated trades, each with a name and a date, are on the home page.

Schedule your free consultation