
40% of Agent Projects Get Canceled. Mine Run Unattended Every Night on a Fixed Budget.
Published: September 28, 2026
Two of my automations start at 6:00 and 15:00 whether I am at the desk or not. They read job boards for roles that match my stack, drop the noise, draft the opening line of an application and write the result to a file I read over coffee. A third wakes up at night, pulls new posts from a feed, sorts out what deserves my attention and leaves a short report behind. None of these are chats I keep reopening. They are scheduled runs with a known cost per execution, and when one of them breaks I find out from a missing file instead of from a client.
Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, and the reasons it names are escalating costs, unclear business value and weak risk controls. Surveys of production agents put failure rates between 70% and 95%, driven by compounding errors, tool calls that break and output nobody checks.
Those numbers come out of engineering decisions, not model quality, and swapping in a stronger model has not moved them. What kept my overnight runs alive was treating the model as one unreliable step inside a system that has to prove it did its job. Here is the setup, including the parts that cost me money to learn.
The failure that leaves no trace
Production agents rarely die loudly. A tool returns partial data or an empty payload, the model does not raise an exception or trigger an alert, and it continues the workflow with a context that is now wrong. The write-up on why agents fail in production describes exactly this shape of failure, and it is the one that took me longest to design against.
My job scans write a report every run. Early on I was reading those reports as evidence that the run worked. A report saying "9 new matches" can be produced by a parser that found nine matches, and it can also be produced by a fallback branch that filled the same sentence with a number it had from a stale file. The cron log looked identical in both cases. Nothing in the system was lying on purpose. The run simply had no obligation to prove anything.
So I stopped treating "the job ran" as the signal and made "the job produced something checkable" the signal. That single change removed most of the silent failures I was getting.
Most of my automations are not agents at all
The cheapest agent run is the one you replaced with a script, and most of what I run unattended never calls a model. The job scans are deterministic code: fetch, parse, filter, format, write. They cost nothing per run, they behave the same way every time and they cannot invent a match that does not exist.
Where a model does appear, it is one step inside a script rather than the orchestrator. Sorting a post into "worth my attention" or "noise" is a classification, so it gets a small model with a short prompt. Drafting an opening line is a generation, so it gets a bigger model, batched, one call per batch instead of one per item. Reasoning is switched off for the cheap steps because I pay for tokens and I have never seen the quality difference justify the bill on a classification task.
The result is a set of jobs where I can point at every model call and say what it is for. Anything I cannot justify that way becomes part of the script.
Every run leaves one artifact
Each scheduled job writes one file: today's matches, today's report, the batch output. The file is the contract. If the file is missing or empty, the run failed, no matter what the log says.
This does two things for me. It makes failures visible without asking the agent how it went, which is a question an agent will answer confidently and wrongly. It also gives me a cheap history. When the number of matches from one board drops to zero for four days, I can see it in the files rather than waiting for a report to tell me something is off.
The reporting itself stays short. A run that produces four lines I read in fifteen seconds is worth more than a page of narration that I start skipping after week two. Long reports are how a pipeline dies quietly: nobody reads them, so nobody notices the day they stop meaning anything.
The model does not grade its own work
Verification is a separate step with no model in it.
A link is assumed good only after a status check returns a 200. A batch job is assumed to have written rows only after I count the rows. A draft is assumed to be in the right state only after the state is read back from the tool that owns it. When I built the outreach pipeline for my own offers, that last rule is what kept a campaign sitting in draft until a human confirmed the sending account, instead of spraying half-configured email at a list.
Every one of those checks is boring, and every one of them has caught something. Asking the same model that produced the output to review the output gives you a second opinion with the same blind spots, and it costs the same as the work it is supposed to be checking.
Anything that leaves the machine needs a human gate
The risk controls Gartner names are the part most teams treat as a later problem, and they are the part that is cheap to build first.
None of my agents can write to a production system or delete anything. They read, classify, draft and file. Sending an email, submitting an application, publishing a post and moving money all sit behind a person pressing the button. The agent gets the list to 90% and the last 10% belongs to me, because the last 10% is the part where a mistake has a recipient.
I took this position after watching how much damage an agent can do with permissions it was handed for convenience. Access is not a feature you can add back once something has gone out under your name.
Budget ceilings before cleverness
Every job of mine has a cost per run that I know, and a monthly ceiling that stops it. Timeouts are short and hard. There are no loops that keep retrying until something works, since an agent stuck in a retry loop can spend a weekend's budget in an hour while producing a thousand useless calls.
That constraint shapes the design more than any prompt. When a task has to fit inside a small budget, the answer is almost always less model work: cache the expensive step, batch the calls, move the decision into deterministic code and keep the model for the part where language is the actual problem.
Why the cancellations keep happening
Cancellations usually get blamed on the technology. The Gartner list is more specific: costs that escalate past the pilot, value nobody defined and risk controls that were never built. All three are engineering and ownership problems, and all three are visible from the outside. A team that cannot say what one useful run costs, cannot show what the system produced last week and cannot name what happens when a tool call returns garbage will cancel the project eventually, and the model will get the blame.
That is also where the money is for a senior developer right now. Building the agent is the easy part and the part everyone can do. Owning the ops around it, meaning the artifact contract, the verification steps, the permission boundaries and the cost ceiling, is the work that decides whether the thing is still running in six months. I wrote about the same idea from the prompting side in Stop Writing Prompts. Start Writing Skills., and the pipeline that runs my own content and outreach is documented in How I Automated My Content and Outreach Pipeline with Custom Claude Code Skills.
If you are building agents that are supposed to run without you, or you have a project that already stopped producing anything useful, this is the work I do with teams. I run a community for developers and founders building this way, The Agentic Architect Lab, and I take on a small number of hands-on engagements where the job is making the automation survive unattended.
Sources
- Gartner press release, over 40% of agentic AI projects will be canceled by the end of 2027: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- Fiddler AI, AI agent failure rate in production: https://www.fiddler.ai/blog/ai-agent-failure-rate
- Why AI agents fail in production, tool call failures and silent corruption: https://dev.to/hadil/why-ai-agents-fail-in-production-and-how-engineering-teams-are-fixing-it-in-2026-job
- Veracode 2026 GenAI Code Security Report, 56% average security pass rate: https://www.veracode.com/resources/analyst-reports/2026-genai-code-security-report/
- Veracode blog summary of the 2026 report: https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/
- Debt Behind the AI Boom, large-scale study of AI-authored commits: https://arxiv.org/abs/2603.28592
- LinearB 2026 Engineering Benchmarks Report, 8.1 million pull requests: https://linearb.io/blog/8-million-prs-engineering-productivity
- Stack Overflow, are bugs and incidents inevitable with AI coding agents: https://stackoverflow.blog/2026/01/28/are-bugs-and-incidents-inevitable-with-ai-coding-agents/
Building with AI beyond this article?
I run The Agentic Architect Lab, live builds, agent workflows, and a playbook for technical founders shipping solo. No toy demos.
Join the Lab