For my first six months building with AI, I measured the wrong things. I measured how many automations I had running, how fast the model returned, how many tools I had installed, how many posts I published. All four went up. All four felt like progress. None of them told me whether the work I was paying for actually happened.
That is the proxy-metric trap, and I walked straight into it. A number that moves is not the same as an outcome that lands. “AI saves me time” is a feeling until you check it against the thing the AI was supposed to produce. I stopped trusting the feel-good numbers and started tracking four that are harder to fake. Here is each one, and the real number that taught it to me.
Did the outcome happen, or did the process just finish?
I now check the outcome, not the exit code. An automation of mine ran on schedule for two weeks and returned success every single time. It produced zero usable outputs across all fourteen days. Exit code 0 means the process finished without crashing. It says nothing about whether the thing you wanted exists. I wrote up the full failure in the automation that reported success while producing nothing for two weeks, and it rewired how I schedule everything.
Every scheduled job I run now ends with a check against the artifact itself. Did the file get written. Does it have rows. Is the value non-empty. The job is not green because it returned; it is green because the outcome is there. Green used to mean “ran.” Now it means “produced.”
How much time does it save after a human reviews it?
I measure end-to-end time including the human review, not the model’s latency. I once bolted an AI step onto a workflow, watched it run in seconds, and called it a win. It cost me time. A person still had to read the output, judge whether it was right, and redo the parts that were not. The full accounting is in the AI step that ran fast and still made the whole workflow slower. The model was quick. The workflow was slower.
Raw speed is the number the AI wants you to look at. Net time after review is the number your calendar feels. If the review plus the rework costs more than the step saved, the automation is a tax you are paying to feel modern. I now clock the whole loop, human included, and compare it to the manual version. A lot of my early automations lost that comparison.
Does anyone use it a second time?
I count second-run usage, not install count. When I audited my own stack, I found 260 skills and tools installed against five months of logs. 150 of them had never run once. Only 36 had run in more than a single session. Installing a capability is not the same as having it, which is the whole point of the audit where I compared what I installed against what I actually used.
This was not a new lesson, only a bigger one. I had already learned the small version of it with a tool I bought, used once, and never opened again, and told myself that one was a fluke. It was not a fluke. It was a pattern I was measuring wrong. The install is the intention. The second run is the evidence. I only count the second run now.
Does anyone ever find what I published?
I measure discovery, not publish count. Across two of my sites I had 15 published posts. Zero were in Google’s index. Not rejected, not ranked low. Never discovered. And exactly one other site linked to any of them. I laid out the whole thing in the backlink audit where I went looking for inbound links and found one.
Publish count is the vanity version of output. It measures how busy I was, not whether the work reached a human. Fifteen posts that no search engine has indexed is fifteen files, not fifteen pieces of content. Now I track how many are indexed and how many are found, and I would rather ship three that get discovered than fifteen that sit in the dark.
What ties these four together?
Every one of them measures the goal instead of the activity around the goal. That distinction cost me real work early on. I once automated the wrong thing first precisely because it ran clean and green while the actual goal went hungry. A clean run is easy to admire. It is a terrible thing to trust.
So the honest version is this. For six months I watched dashboards that were green because I had built them to be green. Activity, speed, installs, publishes. All up and to the right, all measuring the wrong end of the pipe. The four checks I run now are less flattering and a lot more useful. Did it produce. Did it save net time. Did anyone use it twice. Did anyone find it.
If you are past the demo stage and want the same measurement applied to your own stack instead of the vendor’s numbers, that is what my AI and search work is. I count the outcomes, not the exit codes.