Field Note · GOVERNANCE

If your AI workflow stopped tonight, who would know?

When the alarm runs on the workflow it watches, both stop together and the silence looks like a good fortnight. A check you can run on your own AI workflows, and what it found in ours.

Published · Every number is ours, measured on our own operation.

The pilot worked, so the workflow went live. It drafts the quotes, or sorts the invoices, or writes the morning summary for the operations team, and the status screen has shown green since launch day. Nobody has complained. In most operations that counts as proof the thing is healthy: no alarm has gone off, so nothing is wrong.

Why green can mean nothing

Many status checks are built inside the thing they report on. The workflow runs, the check runs alongside it, and when a step breaks, the check sends a message. That design covers almost every failure except the largest one, where the whole workflow stops. The check stops with it. No message goes out, because the sender was part of what went down.

From the outside, that looks exactly like a good fortnight. The dashboard keeps showing its last known state. The alerts inbox stays empty. The people who were told "you'll hear if anything goes wrong" hear nothing, and they reasonably conclude that nothing went wrong.

This gap tends to open at the moment a workflow moves from pilot to production. During the pilot, someone looked at the output every day because their job was to prove it worked. After launch that person goes back to their real work, and the only thing still watching is whatever alarm shipped with the workflow. If that alarm lives inside the workflow, the day the attention moved was the day the workflow stopped being watched.

Silence from an alarm is good news only if the alarm could have spoken.

A check you can run in an afternoon

List every AI workflow you have running in production, then work through the following for each one, in writing, with whoever owns it.

  1. Who learns that it stopped? Write down one person's name. A shared inbox that nobody owns does not count.
  2. Trace the path. Follow the actual route from "the workflow did not run" to that person's phone or inbox: which system notices, which system sends the message, and which account it sends from.
  3. Look for the workflow on that path. If any step shares a schedule, a server, an account or a job with the work it watches, the alarm and the workflow can fail together.

If the workflow sits anywhere on that path, you have a status screen and no alarm. The repair is usually small. Something separate, running on its own schedule, expects a sign of life from the workflow and tells the named person when one is missing.

Then test it. Pick a quiet afternoon, stop the workflow on purpose, and time how long it takes for someone who did not stop it to find out.

What it looked like in our own back office

We run our own back office on scheduled AI jobs. One keeps the task list current, another emails a daily brief, and a separate check marks each job red or green on every run.

Late one Sunday night the application hosting those jobs quit, and every job stopped with it. The check kept running and marked them red, every time. Nobody saw the red, because the only route from that check to a person was a step inside the daily brief, and the daily brief was one of the jobs that had stopped.

The stop came to light sixteen days later, when one of the jobs compared its own last-run date against the calendar and found the gap. Nothing else had raised it. By then the unread backlog held two alarms that had already been disproved and would have been acted on as real by anyone reading them cold.

We moved the alarm outside the thing it watches. The route from that check to a person no longer runs through any of the work it checks, so stopping the jobs cannot silence it. We found the pattern in our own operation first, and we are probably not done finding places where it hides.

Where this leaves a stalled pilot

For a workflow stuck between pilot and production, this is often one of the unanswered questions holding it back. Once nobody is checking it daily, how will anyone know it stopped? A workflow is ready for production when that question has a named person and a path that does not depend on the workflow itself. Read how we approach the move from a pilot that worked to a workflow your team can rely on.

When you're ready to stop reading and start building

The Spark Audit takes
eight minutes.

A prioritized path forward, scoped to your organization. No obligation. No sales sequence.