The Executive Summary
Operators running $200–$400/month in AI subscriptions lose $200–$600/month to gut retention bias — tools they keep without a score. The AI Kill/Keep Decision Engine fixes that in 60–90 minutes.
Who this is for: Service agencies, solo consultants, and internet solos running at least 3 active AI subscriptions with no systematic measurement in place
The gut retention bias problem: At Scaling band, 8–14 AI subscriptions averaging $300–$600/month — with $200–$600/month in Kill-designated tools staying in the stack because no composite score existed to surface the decision
What you’ll learn: The AI Kill/Keep Decision Engine — a 5-dimension scoring matrix (Time Recovered, Revenue Enabled, Replacement Cost, Switching Risk, Capability Trajectory) that produces a composite score from 5 to 25 and maps every tool to Kill (below 12), Keep (12–18), or Upgrade (above 18); plus the 3-week Migration Playbook and Monthly AI Spend Dashboard
What changes if you apply it: From managing your AI stack by gut feel and sunk cost to running a governed system where every tool has a composite score, every Kill decision has a migration path, and every month has a 15-minute dashboard review
Time to implement: 30 minutes to score your three most expensive tools; 60–90 minutes for a full stack audit; 15 minutes monthly to maintain the dashboard; 3 weeks to complete one Kill migration
Written by Nour Boustani for six-figure service operators who want a governed AI stack without reactive cancellations or gut-feel spend.
› Library Navigation: Quick Navigation · AI For Operators
How to Measure AI ROI for Small Business and Cut Wasted Tool Spend
The AI Kill/Keep Decision Engine is a five-dimension scoring system that gives every AI tool in your stack a composite score from 5 to 25, assigns a clear action—Kill, Keep, or Upgrade—and adds a monthly tracking instrument so ROI measurement becomes a recurring operating practice rather than a calculation you make once and forget.
The real problem is not that operators at $30K-$150K/year have too many AI tools. It is gut retention bias: the tendency to keep familiar subscriptions despite unclear value, redundant capabilities, or declining usefulness because cancellation feels riskier than another month of spend.
This engine shifts AI tool management from renewal-by-default to evidence-based resource allocation. By linking each score to a specific decision and reviewing it monthly, operators can stop losing $200-$600 per month to tools that are no longer earning their seat in the business.
Where are you with this right now?
“I’m paying for ChatGPT, Jasper, Grammarly, and Notion AI, but I don’t know if they’re paying for themselves.” If you have four or more subscriptions and no measurement system, score your three most expensive tools. In 30 minutes, the matrix gives each one a clear action.
“I measured AI ROI six months ago, but I can’t prove it’s still positive.” A one-time calculation is not enough. The monthly dashboard tracks score changes over time, so you can spot declining ROI before retention bias keeps an underperforming tool in your stack.
“I know some tools overlap, but I don’t know what to cut.” Use the Kill/Keep/Upgrade thresholds and 3-week migration playbook. You get a defined action and rollback plan instead of making a cancellation decision on instinct.
Try this now (under 2 minutes):
List every AI subscription you’re currently paying for. Include annual plans divided by twelve.
For each tool, write the total monthly cost and one sentence describing the last time you used it in a way that directly affected client work or revenue.
Flag any tool where the sentence you wrote is vague, indirect, or took more than fifteen seconds to write.
The flagged tools are your highest-probability Kill candidates. Most operators at Survival band who do this exercise flag two to three tools in under two minutes - tools that have been running on autopilot for months without anyone verifying the return.
At $250/month average AI subscription spend, three tools flagged means a potential $150-$600 monthly savings from tools already costing more than they’re delivering. Before a single scoring dimension is run.
Why Paying for AI Tools Without a Decision Framework Costs More Than the Tools Themselves
The most expensive AI tool is rarely the one with the highest subscription cost. It is the one you keep because cancelling feels riskier than paying for another month.
At Scaling band ($60–150K/year), operators often run 8–14 SaaS and AI subscriptions. AI spending can easily reach $200–$400+ per month.
Operators are not usually buying carelessly. Each tool was added for a specific reason at a specific point. The problem is that few tools are reassessed six months later to confirm they are still earning their place.
Operators do not have a calculation problem. They have a decision problem: what to do with the result once they have it.
A negative-ROI tool stays because it “might be useful next quarter.”
A borderline-positive tool stays because switching feels uncertain.
A high-value tool is not upgraded because no upgrade threshold exists.
The AI Kill/Keep Decision Engine replaces that judgment call with a composite score tied to a specific action.
Simple time-savings math is not enough. A tool that saves 3 hours per month at $75/hour creates $225 in recovered value. Against a $49/month subscription, it looks like an obvious Keep.
But that calculation ignores:
Switching risk and vendor dependence
Replacement options
Workflow disruption from a migration
Whether the tool’s capability trajectory is improving or declining
Time recovered is one dimension of value, not the entire decision. The single-dimension approach fails because it ignores four of the five dimensions that determine whether a tool is worth keeping.
The real cost of an unmanaged AI stack is not just subscription waste. It creates weaker renewal, migration, and upgrade decisions across the business.
The Unit Economics of an Unmanaged AI Stack
At Survival band, a $49/month tool that scores Kill but stays active for 12 months creates:
$588 in annual spend
Zero measurable return
A payback period of under one week for the 2–3 hours of scoring and migration work required to remove it
Roughly a 20x return on the migration work in the first year
At Scaling band, $147/month in Kill-designated tools kept for six months creates:
$882 in unrecovered spend
A 60–90 minute full-stack audit
A $112 audit-time investment at a $75/hour opportunity value
A 7.9x return from $882+ in annual savings from Kill decisions alone, before Upgrade decisions are included
When an AI stack exceeds 12–15 tools, reviewing every tool monthly can create diminishing returns.
Review the full stack quarterly.
Review monthly only the tools with Capability Trajectory scores of 2 or below.
AI stack cost without a decision framework at Scaling band:
Active AI subscriptions: 8-14 tools, $300-$600/month average spend
Gut retention bias cost: $200-$600/month in tools scoring Kill that stay in the stack
Opportunity cost of non-upgraded tools: $400-$1,200/month in unrealized leverage from tools that should have been upgraded 90 days earlier
Decision quality cost: time spent evaluating redundant tools, onboarding replacements reactively, and managing surprise billing changes
Daily Cost of Running Without a Decision Framework
Without a decision framework, inactive, redundant, or low-value AI tools can cost $13–$20 every working day before accounting for missed leverage from tools that should have been upgraded.
Use the Engine when your stack has enough complexity to justify the review time.
Survival band ($30–60K/year): Run the Engine when you have at least 3 active AI subscriptions. Below that threshold, the time required to score and maintain the system usually outweighs the benefit.
Scaling band ($60–150K/year): Run the Engine monthly as a standard governance practice. At this stage, stack complexity makes gut-feel management a consistent source of misallocated spend.
If the damage is already done - the reset protocol:
The stack has been running without measurement for six or more months. Tools are entangled in workflows without documentation. The question is what reset costs versus continuation.
Within 30 days:
Run the Kill/Keep/Upgrade scoring on your three most expensive tools. Total time: 30 minutes.
For each tool scoring Kill: don’t cancel immediately. Initiate the 3-week migration assessment from the playbook before touching subscriptions.
Reset cost: 2-3 hours for full initial audit.
30-90 days:
Monthly dashboard running. All tools scored. Kill decisions actioned with migration playbooks complete.
Savings from Kill decisions redirected to highest-scoring Upgrade candidates.
90+ days:
Monthly tracking is routine. Trend lines visible. Capability trajectory scoring catches vendor deterioration 60-90 days before it becomes a crisis.
One thing from this section: The gut retention bias doesn’t feel like a bias - it feels like caution. The scoring matrix is the only thing that separates caution from inertia.
UNMANAGED STACK COST TIMELINE
No framework Month 1: $420/month, $147 in inactive tools
Month 3: Emotional cancellations, wrong tools cut
Month 6: $882 paid to Kill-designated tools
Month 12: $1,764 in unrecovered spend
With Engine Month 1: Inventory + scoring, $147 Kill identified
Month 2: Kill decisions actioned, $147/month freed
Month 3: Dashboard running, Upgrade actioned
Month 6: 6-month trend line, zero surprise billingThe cost framework shows you what unmanaged tool spend produces. The next section shows you the five dimensions that replace the judgment call with a score.
How to Measure AI ROI With the AI Kill/Keep Decision Engine
The operators who manage their AI stack correctly aren’t more disciplined than the ones who don’t. They have a scoring system that makes the decision before emotion does.
The AI Kill/Keep Decision Engine runs every AI tool through five dimensions, assigns each a score from 1 to 5, calculates a composite score from 5 to 25, and maps that composite to one of three specific actions with no ambiguity:
Below 12: Kill - with a migration path before cancellation
12-18: Keep - with specific conditions and a review date
Above 18: Upgrade - increase usage, expand integration, or move to a higher tier
The five components run in sequence. Each builds on the one before it. The full system runs in 30-45 minutes for a three-to-five tool evaluation and 60-90 minutes for a complete stack audit.
Kill/Keep/Upgrade Decision Thresholds
Score 5–11: Kill. Complete the migration playbook before cancelling.
Score 12–18: Keep. Set conditions and a review date.
Score 19–25: Upgrade. Take one specific integration action this month.
Each tool receives five independent scores from 1–5, creating a composite score from 5–25.
The Tool Inventory: Name, Cost, and Use Every Subscription
You cannot score what you have not counted.
The Tool Inventory is the first component of the Engine. List every active AI subscription and record three fields for each tool:
Monthly cost: Include annual plans divided by 12, plus tools paid through company cards that may not appear in personal billing reviews.
Primary function: In one sentence, describe what the tool does in your workflow now—not what you originally bought it to do.
Last active use: Record the last date you used it to affect a deliverable, client relationship, or revenue. Logging in or checking something does not count.
If you cannot describe a tool’s current function in under 30 seconds, you have a usage-documentation problem before you have an ROI problem.
Last active use is often the fastest diagnostic. A $29/month tool last used 47 days ago is already a Kill candidate before you score a single dimension. The scoring matrix confirms the decision; the inventory surfaces it in seconds.
At Survival band, the first inventory commonly reveals one or two tools unused for more than 30 days. At Scaling band, with 10+ subscriptions, that often rises to three to five tools.
These are not necessarily tools you intentionally stopped using. They are tools you meant to use, gradually stopped using, and continued paying for.
The “just in case” tool is often a monthly payment for a promise you made to yourself six months ago.
Quick Signal
Pull your last three months of billing statements and list every AI subscription with its monthly cost.
Include monthly and annual subscriptions.
Divide annual plans by 12.
Include company-card charges.
Add the monthly costs.
That total is your AI stack baseline: the number the Engine will either validate or reduce.
The Tool Inventory names what you are running. The five-dimension scoring determines whether each tool is worth keeping.
The 5-Dimension Scoring - The Matrix That Replaces the Judgment Call
The single biggest failure in AI tool evaluation is asking one question: “Is this saving me time?” That’s one dimension of five. The other four are what make tools worth keeping - or dangerous to keep.
Each of the five dimensions is scored 1 to 5. The composite score runs 5 to 25. Here is what each dimension measures and how to score it accurately.
Dimension 1: Time Recovered
What it measures: actual hours saved per month using this tool versus doing the same work manually. Not estimated hours. Tracked hours.
Scoring guide:
5: Saves more than 8 hours monthly - equivalent to a full work day
4: Saves 5-8 hours monthly
3: Saves 3-5 hours monthly
2: Saves 1-3 hours monthly
1: Saves under 1 hour monthly or time savings are not measurable
The honest test: If you can’t name the specific tasks where time is saved with an estimate you’d defend in a client meeting, score 1 or 2. Most operators overestimate time savings by 30-50% when they haven’t measured directly.
Dimension 2: Revenue Enabled
What it measures: direct revenue impact - client work produced faster, proposals generated at higher volume, retainer capacity created, or specific revenue outcomes traceable to this tool.
Scoring guide:
5: Directly enabled $1,000+/month in additional revenue or protected existing revenue at that scale
4: Directly enabled $500-$1,000/month
3: Indirect revenue impact - improved work quality that contributes to retention
2: Marginal connection to revenue - hard to trace a direct line
1: No measurable revenue connection
The critical distinction: Revenue enabled is different from time recovered. A tool that saves 5 hours monthly in administrative work scores high on Dimension 1 but low on Dimension 2 if those hours aren’t redirected to revenue-generating activity. Both dimensions must be scored independently.
Dimension 3: Replacement Cost
What it measures: what it would cost to replace this tool’s function with a human, a different tool, or a manual process. This dimension captures how much leverage the tool is creating relative to the alternative.
Scoring guide:
5: Replacement would cost $500+/month in human time or alternative tool cost
4: Replacement would cost $200-$500/month
3: Replacement would cost $50-$200/month
2: Replacement would cost under $50/month or a free tool exists
1: Free alternatives exist that provide equivalent function
Why this dimension matters: A tool scoring Kill on Dimensions 1 and 2 but scoring 5 on Dimension 3 stays in the evaluation longer because the replacement cost is high. The composite score catches this - a low-engagement tool with no replacement alternative scores differently than a low-engagement tool with three free alternatives.
Dimension 4: Switching Risk
What it measures: how deeply embedded this tool is in active workflows, how much rework a switch would create, and whether the tool holds data or outputs that would be lost or complicated to migrate.
Scoring guide:
1: Deeply embedded - switching would require rebuilding multiple workflows, migrating significant data, or retraining team members; estimated switching cost $500+ in time
2: Moderately embedded - switching creates 4-8 hours of rework
3: Lightly embedded - switching creates 1-4 hours of rework
4: Minimal embedding - switching takes under 1 hour
5: No embedding - tool is standalone with no workflow dependencies
Note: Switching Risk is scored inversely - a high switching risk scores low (1-2). This is intentional.
High switching risk is a warning, not an asset. A tool that scores 1 on Switching Risk should trigger a Vendor Lock-In Risk Flag evaluation regardless of its composite score.
Dimension 5: Capability Trajectory
What it measures: the direction the tool is moving - is it getting more valuable, holding steady, or declining as the AI market evolves around it?
Scoring guide:
5: Active development, frequent updates, growing user community, competitive market position strengthening
4: Steady development, regular updates, stable market position
3: Maintenance mode - updates are minor, market position is holding but not growing
2: Limited development signal, user community shrinking, competitors pulling ahead
1: Stagnant or declining - no significant updates in 6+ months, competitor tools outperforming on core function
Why Capability Trajectory Changes the Decision
Composite score 14, Capability Trajectory 1: Keep with an active exit plan.
Composite score 14, Capability Trajectory 5: Keep with an upgrade path.
The composite score is the same. The required action is different.
The composite score is the same. The required action is different.
The composite score in practice - a Survival band example:
A solo consultant at $44,000/year running four AI subscriptions evaluates their writing assistant tool:
Time Recovered: 4 hours monthly saved on proposal drafts = Score: 3
Revenue Enabled: Proposals improved - one additional close per month at $800 average = Score: 3
Replacement Cost: ChatGPT Plus covers 80% of the function at $20/month vs. $49/month current cost = Score: 2
Switching Risk: Minimal - no workflow dependencies, no data to migrate = Score: 5
Capability Trajectory: Market position weakening as native AI writing in existing tools improves = Score: 2
Composite Score: 15 - Keep with conditions.
The condition: review in 60 days for Capability Trajectory changes. If competitors close the gap further, this tool moves to Kill. The monthly dashboard flags it for the next review cycle automatically.
Scoring Readiness Check - required before executing any Kill decision:
Every dimension must have a tracked data source, not an estimate:
Time Recovered: based on actual logged hours, not memory
Revenue Enabled: traceable to a specific deliverable, client outcome, or capacity created
Replacement Cost: based on a real alternative you’ve identified and priced
Switching Risk: based on a workflow inventory, not assumed
Capability Trajectory: based on a changelog or community check in the last 30 days
Pass = all five dimensions sourced from data, not gut feel.
Fail = stop. Do not execute Kill decisions until dimensions are sourced. A Kill decision made on estimated scores produces tool regret within 30 days.
An incorrectly estimated Time Recovered score is the single most common cause of reactivating a cancelled tool. Score from data or don’t score.
What AI-Assisted Scoring Looks Like
AI-Assisted Scoring vs. Manual Review
A complete manual review of a 10-tool AI stack takes 4–6 hours when you research Capability Trajectory, verify time savings against tracked data, and calculate replacement costs from scratch.
AI-assisted scoring reduces the same review to 60–90 minutes: a 3–4x speed advantage.
AI can help surface:
Capability Trajectory signals in product changelogs and user communities
Free alternatives that have improved since the last review
Switching-risk dependencies revealed when you document the workflow
Use AI for a faster second opinion, then verify the inputs before making a Kill decision.
Exact prompt for your scoring pass:
I'm evaluating an AI tool for my business. Tool: [name]. Monthly cost: [amount]. Primary function: [one sentence]. How I currently use it: [description]. Run a 5-dimension evaluation scoring each dimension 1-5:
1. Time Recovered - hours saved monthly
2. Revenue Enabled - direct revenue impact
3. Replacement Cost - what it would cost to replace this function
4. Switching Risk - how embedded it is in active workflows
5. Capability Trajectory - is this tool getting more or less valuable as the AI market evolves.
Show the composite score and tell me Kill (below 12), Keep with conditions (12-18), or Upgrade (above 18). For any Kill result, name the two most viable alternatives currently available.Tool: Claude or ChatGPT, free tier. This evaluation benefits from a second opinion; it does not require a paid research service.
What the Framework Teaches Beyond AI Tools
The AI Kill/Keep Decision Engine is not only an AI tool-management system. It teaches resource-allocation discipline: evaluate every recurring cost against a multi-dimensional return standard instead of renewing by default because cancellation feels risky.
Run the system for three consecutive monthly cycles and apply the same five-dimension logic to:
SaaS subscriptions
Contractor relationships
Service offerings that have outlived their productive role
The AI stack is the training ground. The mental model is the asset.
Steal This
Gut retention bias rarely feels like bias. It feels like caution.
The difference is whether a score supports the decision. Operators who build effective AI stacks do not necessarily have better instincts about what to keep. They use a scoring system before instinct takes over.
Premium Toolkit available for members
The AI Kill/Keep Decision Engine System includes:
Kill/Keep/Upgrade Decision Scorecard — score every tool across five dimensions and choose the right action with confidence.
Monthly AI Spend Dashboard Template — track scores, spend, trends, and renewal dates so measurement becomes routine.
Migration Playbook — replace low-value tools safely through parallel testing, capability checks, and a rollback plan.
Vendor Lock-In Risk Flags — identify dangerous dependencies early and build an exit plan before they become a crisis.
Plug-and-play AI diagnosis sessions — drop into Claude, Gemini or ChatGPT, answer a few questions, save hours of guessing, get your exact next move
Audio key points — concentrated frameworks you can absorb in minutes, implement while you move
Unlock 750+ ready-to-use constraint toolkits — built to solve every business problem operators actually face.
Prevent $200–$600 in monthly waste and capture 30–50% more leverage from tools you should upgrade.
Cancel anytime. Every download you’ve accessed stays with you.
If you’re a service agency or solo consultant running at least 3 active AI subscriptions with no systematic measurement in place, this system resolves that gap before the next billing cycle.
If you haven’t yet identified which tasks AI should be handling in your business, Find Where AI Actually Saves You Money - The AI Opportunity Audit is the right starting point - the audit identifies what to automate; this Engine validates whether it was worth automating.
Build the measurement system now. Every month without it is another month of paying the gut retention tax.
One thing from this section:
A composite score from 5 to 25 replaces the judgment call. Below 12 is Kill.
Above 18 is Upgrade. The number makes the decision before emotion does.
The scoring matrix produces the decision. The composite thresholds tell you what to do. The next section shows how to execute each action - especially the Kill - without losing something you actually needed.
The Implementation Protocol - Running the Engine Against Your Current Stack
The Engine doesn’t run in theory. It runs against the specific tools, costs, and workflows you have today.
The implementation sequence runs in four steps. Each step produces a specific output.
The full stack audit completes in one 60-90 minute session for most operators at Survival band. At Scaling band with 10+ tools, schedule two 60-minute sessions rather than one long one - decision fatigue is real and it produces Keep decisions that should be Kill.
ENGINE IMPLEMENTATION SEQUENCE
Step 1: Tool Inventory 20-30 min
|
v [All tools listed with cost + last use]
|
Step 2: Score Top 3 Tools 10-15 min per tool
|
v [Composite scores + Kill/Keep/Upgrade]
| [Scoring Readiness Check - Pass/Fail]
|
Step 3: Kill Migration 3 hours across 3 weeks
|
v [Migration complete + rollback plan]
|
Step 4: Monthly Dashboard 15 min/month ongoing
|
v [Trend line + action dates]
Total first build: 60-90 min
Ongoing: 15 min monthlyStep 1: Build the Tool Inventory
Action: List every active AI subscription with its monthly cost, primary function, and last active use date.
How to execute:
Pull the last three months of billing statements.
List every AI tool line item.
Include tools you are embarrassed to still be paying for.
Do not score, research, or cancel anything during this pass.
The inventory is diagnostic, not judgmental. It contains only three fields per tool:
Monthly cost
Primary function
Last active use
Tool: Claude, free tier.
I’m building an inventory of my AI tool subscriptions.
Here is my list:
[paste tool names and monthly costs]
For each tool, I will provide its primary function and last active use date.
Identify:
- Likely free-tier alternatives
- Recent competitors that may have closed the capability gap
- Which tools I should prioritize for scoring first
Do not make cancellation recommendations yet. Return a concise bullet list.Time: 20–30 minutes for a complete inventory, including billing-statement review.
If this takes more than 45 minutes, stop. You are trying to score tools during the inventory pass.
Do not assess ROI yet.
If you cannot recall last active use, write “unknown” and continue.
Save scoring for Step 2.
Output: A complete list of every AI subscription with its monthly cost, primary function, and last active use date.
What it enables: A prioritized scoring pass. Start with the three to five tools that combine the highest monthly cost with the oldest last active use.
If more than five tools show no active use in 30+ days, score the three most expensive first. Do not score everything in one session; decision quality drops after the fourth evaluation.
Step 2: Score the Three Most Expensive Tools
Action: Run the 5-dimension scoring matrix against your three highest-cost AI subscriptions.
How to execute:
Score each tool across all five dimensions using the scoring guides in The 5-Dimension Scoring: The Matrix That Replaces the Judgment Call.
Record each dimension score and calculate the composite score.
Run one scoring prompt per tool, sequentially rather than in batches.
When uncertain between two scores, use the lower score. If a tool could be a 3 or 4, score it as a 3.
The Engine is designed to produce conservative composite scores. Kill decisions should be earned through evidence, not triggered by optimistic assumptions.
Tool: Use the scoring prompt from The 5-Dimension Scoring: The Matrix That Replaces the Judgment Call.
Time: 10–15 minutes per tool, including AI-assisted scoring and review.
If the process takes more than 20 minutes per tool, it has drifted into research mode. Do not pause the scoring pass to compare detailed features, validate every Capability Trajectory signal, or cross-check several replacement-cost estimates.
Score with the information you have.
Flag uncertain dimensions.
Complete a five-minute verification check after the full scoring pass.
Use the first pass to establish decision direction; improve precision in the review.
Output: A composite score from 5–25 and a Kill, Keep, or Upgrade designation for each of the three tools.
What to do at each threshold:
Kill (below 12): Start the 3-week migration assessment before cancelling. A Kill score means the tool does not currently justify its cost; it does not mean the tool is worthless. Confirm a viable replacement before the subscription ends.
Keep (12–18): Set a 60-day review date. Write the specific condition that would move the tool to Kill at the next review, such as a decline in Capability Trajectory or Time Recovered.
Upgrade (above 18): Choose one specific upgrade action for this month, such as moving to a higher tier, integrating the tool into a deeper workflow, or expanding use to more team members. An Upgrade score means value exists that you have not fully extracted.
Step 3: Execute Kill Decisions With the Migration Playbook
Action: For every tool that scores Kill, complete the 3-week migration sequence before cancelling.
The goal is not to prove the current tool is worthless. It is to confirm that a replacement, free alternative, different product, or manual process can handle the core workflow without creating a costly capability gap.
Week 1: Test the Replacement in Parallel
Run the replacement for the three highest-frequency use cases you currently handle with the Kill-designated tool.
Test five real examples.
Log every instance where the replacement underperforms.
Record the time and quality gap with specific evidence.
Useful note:
“The replacement required 25 minutes of editing versus 10 minutes with the current tool.”
Unusable note:
“The replacement isn’t as good.”
Do not test every edge case in Week 1. Test the core function.
Week 2: Calculate the Capability Gap
Review the Week 1 log and calculate the real cost of any gap between the current tool and its replacement.
Measure the additional time required.
Identify any lost or reduced capability.
Score the replacement using the same 5-dimension scoring matrix.
Compare the replacement composite score with the current tool’s score.
The composite score comparison determines the decision, not your preference for the familiar tool.
Week 3: Cancel With a Rollback Plan
If Week 2 confirms that the replacement is viable, cancel the subscription.
Before cancelling, document:
Data and outputs to export from the current tool
Workflows that need updating
The replacement tool or process
The rollback trigger: the specific condition that would justify reactivating the subscription
A rollback plan is not permission to second-guess the decision. It is insurance that removes hesitation from a justified cancellation.
Time: 3 hours total across three weeks for one Kill execution.
If Week 1 takes more than 2 hours, you are over-testing. Limit the test to five examples across the tool’s three most frequent use cases.
The Financial Case for Migration Discipline
For a Scaling-band operator with three Kill-designated tools averaging $89/month:
Monthly savings: $267
Annual savings: $3,204
Migration time: 3 hours
Migration-work value at $75/hour: $225
Return on migration work: roughly 14x in the first year
Step 4: Install the Monthly AI Spend Dashboard
Action: Set up a recurring monthly dashboard for your full AI tool stack.
Track five fields for each tool:
Tool name
Monthly cost
Current composite score
Month-over-month score change
Renewal decision due date
After the initial scoring is complete, the dashboard takes about 15 minutes to update each month.
The most important signal is the month-over-month change in Capability Trajectory. For example, a tool may hold a composite score of 16 for three months, then fall to 13 because its Capability Trajectory score drops from 3 to 1. The dashboard surfaces the decision before gut retention bias has time to rationalize another renewal.
No AI tool is required for the monthly update. Use a simple fill-in template. Use the Step 2 scoring prompt only when a specific tool needs re-evaluation, not every month for every tool.
Time: 15 minutes per month after setup.
Output: A running record of each tool’s score, trend, and next decision date.
If three or more tools decline in composite score at the same time, schedule a full-stack re-evaluation—not only a review of those tools. Multiple declines indicate that the AI market may be moving faster than your stack, and tools currently marked Keep may now have stronger replacements.
Three Operator Examples: How the Engine Changes AI Spend Decisions
Agency at $52,000/year
3 client retainers
7 AI subscriptions
Total AI spend: $380/month
The Tool Inventory identifies two subscriptions with no active use in 60+ days, costing $78/month combined.
The scoring pass covers the five most expensive tools:
1 Kill: $49/month
2 Keeps, each with a 60-day review condition
2 Upgrades for tools underused relative to their capability
The Kill decision saves $49/month. One Upgrade decision redirects $30/month into expanded use of the highest-scoring tool.
Net spend change: -$19/month
Stack performance: Measurably higher, because the highest-leverage tools are used more fully
Solo Consultant at $38,000/year
4 AI subscriptions
Total AI spend: $167/month
Three tools score Keep. One $29/month tool scores Kill with a composite score of 9:
Low Time Recovered
No measurable Revenue Enabled
A free alternative covers 90% of the function
The Kill decision saves $348 annually.
The scoring pass also identifies a $99/month tool with a Keep-with-conditions score of 15. Its Capability Trajectory has declined since the operator’s last informal evaluation.
The 60-day review date creates a structured expiration for the Keep decision for the first time.
Scaling Operator at $95,000/year
11 AI subscriptions
Total AI spend: $590/month
Full-stack audit completed in two sessions
Results:
2 Kills: $118/month combined
5 Keeps: 3 with 60-day conditions and 2 with 90-day conditions
4 Upgrades
The two Kill decisions free $118/month.
Two Upgrade tools are being used at only 30–40% of their capability. Deeper integration is estimated to create $400–$800/month in additional output value.
The Monthly AI Spend Dashboard is scheduled for the first Monday of every month. Within 90 days, one Keep-with-conditions tool moves to Kill after its Capability Trajectory score falls to 1. Monthly tracking catches the decline before the next renewal becomes automatic.
Checkpoint: The Engine is operational when:
Every active AI subscription has a composite score
Every Kill decision has a completed or active migration playbook
The Monthly AI Spend Dashboard is scheduled
This condition either exists or it does not.
The Kill decision is only as good as the migration playbook behind it. Without the 3-week assessment, Kill decisions create tool regret. With it, they can produce $3,204 in annual savings at a 14x ROI on migration work.
The Implementation Protocol handles the current stack. The Validation, Simulation, and Thinking section shows what to measure to confirm the system is working and how to spot Vendor Lock-In Risk Flags before they become a crisis.
How to Validate Your AI ROI Decision Engine
Building the Engine is month one. Knowing it’s working is month three.
Your AI Stack Cost Calculator
- Active AI subscriptions: _
- Total monthly AI spend: $_
- Tools with last active use over 30 days: _
- Estimated monthly cost of inactive tools: $_
- Monthly spend: $__
- Tools inactive 30+ days: _
- Inactive tool cost: $__
- Annual inactive cost: $__ x 12 = $__
- Engine audit time: 60-90 min
- Audit cost at $75/hr: $112
- Annual savings from Kill decisions: $__
- Return on audit: ____xPre-Filled Example: Survival Band
- Active AI subscriptions: 6
- Total monthly AI spend: $310
- Tools with last active use over 30 days: 2
- Estimated monthly cost of inactive tools: $78
- Estimated gut retention bias cost: $78–$156/monthAt $78-$156 monthly in inactive or low-ROI tools, the annual exposure is $936-$1,872 - paid to subscriptions that are not producing return. The Engine converts this from an estimate to a calculated Kill/Keep decision in 30 minutes of scoring.
Simulation: From Unmanaged AI Spend to a Governed Stack
Starting scenario:
Consultant revenue: $47,000/year
Active AI subscriptions: 8
Total AI spend: $420/month
Systematic measurement: None
Tools unused for 45+ days: 3
Discovery Phase
The Tool Inventory shows that $147 of the $420 monthly spend is tied to tools with no active use.
The 5-dimension scoring pass produces:
Three inactive tools: composite scores of 7, 9, and 11 — all Kill
Three most expensive active tools: composite scores of 13, 17, and 22 — Keep, Keep, and Upgrade
Resistance Phase
The $49/month tool scoring 7 has been in the stack for 11 months. It was originally purchased for a use case completed in Month 2, but continued running through gut retention bias.
During Week 1 of the migration assessment, a free alternative handles the core function at equivalent quality. The subscription is cancelled in Week 3, freeing $49/month.
Success Phase
Within 90 days:
Three Kill decisions actioned
$147/month saved
$1,764 saved annually
The tool scoring 22 receives a deeper integration into the proposal workflow
Estimated output improvement: 2 additional proposals per month at a $750 average value
Estimated additional revenue enabled: $1,500/month
The Monthly AI Spend Dashboard is running. Total stack spend falls to $273/month, down 35% from the starting point, while the remaining tools deliver measurably higher performance.
Tool: Claude, free tier.
I have 8 AI subscriptions totaling $420/month.
Here is the list with monthly costs:
[paste tool names and monthly costs]
Before I run full 5-dimension scoring, identify:
- Which tools most likely have free alternatives
- Which tools may now have their capability matched by tools I already pay for
- The three tools I should score first as potential Kill decisions
Use only the information provided. Do not recommend cancellation yet.
Return:
- A prioritized list of three tools
- A short reason for each priority
- Any information I need to verify during the scoring passTwo Futures: With and Without the Engine
Without the Engine: First 90 Days
Month 1
The stack continues without measurement.
$420/month is paid, including $147 for tools with no active use.
No Kill decisions are made because no composite score surfaces them.
Month 2
A Keep tool releases a feature that partly overlaps with the $49 Kill candidate.
The operator now pays for two tools performing the same function.
Month 3
An invoice review triggers a manual stack review.
Tools are cancelled emotionally rather than based on the 5-dimension matrix.
A tool that scores Keep (composite 14) is removed because it appears redundant.
Two actual Kill candidates remain because they have been in the stack long enough to feel indispensable.
$147/month in Kill-designated tools continues running.
Without the Engine: By Month 6
The emotional cancellation creates a reactive replacement project. Finding and onboarding a replacement takes 8–10 hours and creates a workflow gap, while the tool scoring 8 remains active.
By Month 6, the stack has drifted into a configuration the operator did not deliberately choose:
Three tools overlap
One critical workflow depends on a tool with a Capability Trajectory score of 1
No measurement system identifies the risk
The next billing cycle renews every tool automatically
$147/month in Kill-designated tools has run for six months, creating $882 in spend with no measurable return
With the Engine: First 90 Days
Month 1
The Tool Inventory is complete.
Three tools score Kill, totalling $147/month.
Migration playbooks begin for all three.
Two Kill decisions are confirmed and cancelled by Week 3.
$98/month is freed.
Month 2
The Monthly AI Spend Dashboard is running.
The remaining Kill decision is confirmed after its replacement passes the Week 2 capability assessment.
Another $49/month is freed.
Total monthly spend released: $147.
Month 3
The first Upgrade decision is actioned.
The tool scoring 22 receives deeper integration into the proposal workflow.
Improvement is visible within two weeks.
The stack is smaller, cheaper, and producing more leverage than in Month 1.
With the Engine: By Month 6
The $147/month released from Kill decisions is redirected to the highest-scoring Upgrade tool through a deeper tier or expanded use. As usage intensity grows, its Revenue Enabled score improves and its composite rises from 22 to 24 at the next review.
The dashboard also detects a Capability Trajectory decline in a Keep tool, from 3 to 1. Its migration playbook begins in Month 5 and finishes in Month 6, before the dependency becomes a crisis.
By Month 6, the operator has:
Completed six monthly reviews
Built a six-month trend line for every active tool
Eliminated surprise billing changes and reactive cancellations
Prevented tools from running beyond their productive life
Reduced stack cost by 35% from Month 1
Increased measurable output leverage from the tools that remain
What Good Looks Like at Each Stage
Day 14:
Tool Inventory complete - every subscription named, costed, and function-documented
Three most expensive tools scored through the 5-dimension matrix
At least one Kill decision identified with migration playbook initiated
Adjustment protocol if below: the inventory isn’t complete. Pull billing statements for all three months, not just the current one. Annual subscriptions frequently hide from monthly reviews.
Week 4:
First Kill decision actioned - migration complete, subscription cancelled, rollback plan documented
Monthly dashboard set up with all active tools scored
Keep-with-conditions tools have their 60-day review dates on the calendar
Adjustment protocol if below: the migration playbook stalled at Week 1. The most common cause: the operator found one instance where the replacement underperformed and stopped testing. The decision rule is composite score comparison, not perfection test. Re-run Week 2 with that benchmark explicitly in mind.
Week 8:
Monthly dashboard has two data points - enough to see direction of Capability Trajectory movement on any tool
At least one Upgrade decision actioned with a specific integration change made
All Kill decisions from the initial scoring are either complete or in active migration
Adjustment protocol if below: the Upgrade decisions haven’t been actioned. The most common cause: the Upgrade designation gets treated as “use this tool more” without a specific change. The action must be concrete: “add this workflow step,” “train one additional team member,” “move to higher tier for [specific feature].”
If It Does Not Work - Rollback and Retest
If a Kill decision produced a capability gap after cancellation:
Revert: reactivate the subscription if the rollback trigger was met. This is why the rollback plan exists - not as permission to second-guess every Kill, but as a structured reversal when the replacement genuinely underperforms.
Re-diagnose: identify which dimension was scored incorrectly. The most common error is overscoring the replacement on Replacement Cost (Dimension 3) - assuming a free alternative covers the function when it covers 60-70% of it. Rescore with the actual capability gap documented.
One-variable adjustment: correct the single dimension that was off. Recalculate the composite. If it moves above 12 with the corrected score, the tool was a Keep. Reactivate it. Note the scoring error for the next evaluation.
Retest timeline: 30 days after reactivation, rescore with tracked data rather than estimates.
Capability Trajectory Changes Fastest
Capability Trajectory changes faster than any other scoring dimension. A tool scoring 4 in Q1 2025 may score 2 by Q1 2026 as competitors improve.
That is why the Monthly AI Spend Dashboard exists. The AI market moves faster than an annual review cycle.
Decision Engine Failure Mode Analysis
The Engine can fail in four predictable ways. Each has an early signal and a defined recovery path.
Failure Mode 1: Sunk Cost Override
What goes wrong: A tool scores Kill with a composite of 8, but the operator has paid for it for 14 months and cannot cancel. It becomes a “Keep for one more month” decision that never ends.
Early signal:
A tool scores Kill in two consecutive monthly reviews.
No migration playbook has started.
Recovery: Set the rule before you run the Engine. Any tool that scores Kill in two consecutive monthly reviews automatically enters the migration playbook within five business days, regardless of how long it has been in the stack.
Timeline to correct: Start migration within one week of applying the two-consecutive-Kill rule.
Failure Mode 2: Discounted Annual Plan Trap
What goes wrong: A tool scores Kill, but eight months remain on a $588 annual plan. The operator keeps using it to “recover” the prepaid cost.
The financial logic is wrong. The prepaid fee is already sunk. Continuing to use a low-value tool creates eight more months of workflow friction, delays migration to a better option, and does not recover the prepaid amount.
Early signal:
The operator knows the tool scores Kill.
The renewal date becomes the decision date.
The Kill decision is acknowledged but deferred.
Recovery:
End the tool’s workflow integration now.
Keep the prepaid subscription dormant until renewal.
Cancel at renewal.
Do not wait to improve the workflow.
Timeline to correct: Immediate. The Kill decision is a workflow decision, not a billing decision.
Failure Mode 3: Capability Trajectory Score Inflation
What goes wrong: An operator gives a tool a 4 or 5 for Capability Trajectory because it is a familiar brand, not because they reviewed evidence. This pushes borderline Kill tools into the Keep range.
Early signal: The operator cannot name the tool’s most recent meaningful capability update.
Recovery: Make one evidence check mandatory before scoring Capability Trajectory:
Name the last meaningful update.
Record its release date.
Check the changelog or user community.
If the most recent meaningful update was more than 90 days ago in a fast-moving category, the score cannot exceed 3, regardless of brand recognition.
Timeline to correct: Apply the changelog requirement at the next monthly review.
Failure Mode 4: Dashboard Installed but Not Actioned
What goes wrong: The dashboard exists, tools have composite scores, and Keep decisions have review dates—but nothing happens when dates arrive. The dashboard becomes record-keeping rather than a decision instrument.
Early signal:
The dashboard contains three or more months of data.
No score has moved more than 1 point in either direction.
No new Kill or Upgrade decision has been made.
Static data in a dynamic market is the warning signal.
Recovery: Every monthly review must produce a decision, not merely an observation. Ask:
What changed?
What does the change require?
If nothing changed, why is no action required?
Record either the action or “no action required,” with a specific reason. This prevents the dashboard from becoming passive.
Timeline to correct: At the next monthly review, add a required “decision or no decision, and why” field to the dashboard.
What This Framework Trains You to See
Early Signal 1: Composite Score Compression
When a previously high-scoring tool (20+) falls into the 15–17 range across two or three monthly reviews, without a change in your usage, the decline is usually in Capability Trajectory. The tool is becoming less valuable as the market improves around it.
Action: Accelerate the next evaluation and start building a replacement shortlist now. Do not wait for the score to reach Kill.
Early Signal 2: Switching Risk Is Rising
A tool that starts with a Switching Risk score of 4 and falls to 2 is becoming more embedded in your workflows. This is a Vendor Lock-In Risk Flag.
Action: Run the Vendor Lock-In Risk Flags evaluation immediately. A tool you cannot easily leave has pricing power over your business, and that dependency compounds over time.
Early Signal 3: Revenue Enabled Is Disconnecting
A tool that initially scores high on Revenue Enabled can lose its justification when usage shifts toward administrative work.
Action:
Document the revenue-enabling use cases the tool no longer supports.
Identify whether another tool replaced those use cases.
Confirm whether the revenue activity was replaced or simply abandoned.
The Monthly AI Spend Dashboard separates a one-time audit from a governed stack.
Two data points produce a trend.
Three data points produce a decision.
Gut retention bias does not survive a visible trend line.
The Validation, Simulation, and Thinking section establishes what working looks like. The Capability Trajectory section shows how to measure the dimension that changes fastest.
Capability Trajectory: The Fastest-Changing Dimension
Two tools with identical composite scores are not equivalent if one is improving and the other is declining. Capability Trajectory changes the decision.
Capability Trajectory is the most forward-looking of the five dimensions and the easiest to score inaccurately. “This tool seems to be improving” is not enough. Feel-based scoring creates some of the Engine’s most expensive errors.
Score Capability Trajectory using three observable indicators.
Indicator 1: Meaningful Update Frequency
The clearest public signal of development investment is the cadence of meaningful capability updates, not minor interface changes.
Score 4–5: Meaningful capability updates every 4–6 weeks.
Score 1–2: No significant capability release in six months or more.
Check the tool’s official changelog, product blog, and user communities such as Reddit, Slack groups, or Discord servers. User discussions often reveal capability gaps more clearly than product marketing.
Indicator 2: Community Direction
A growing, active user community suggests that the tool is solving real problems at scale. A shrinking community, reflected in lower post volume, slower answers, and fewer new-user introductions, suggests users may be leaving faster than they arrive.
This check takes 10–15 minutes per tool.
Search the tool name on Reddit.
Filter results to the last 90 days.
Count relevant posts and assess the discussion.
Compare “how do I use this?” posts with “I’m switching to [alternative]” posts.
Indicator 3: Competitive Position
A tool’s position can change materially in 90–180 days. A writing tool differentiated by tone consistency in 2024 may face several competitors that match its core function in 2025.
When the feature that justified the original purchase becomes commoditized, reduce the Capability Trajectory score even if the tool itself has not declined.
Check:
G2
Product Hunt
Searches for “alternatives to [tool]”
If five free or lower-cost alternatives now exist that were unavailable at the last evaluation, the tool’s competitive position has weakened.
Capability Trajectory Review Cadence
A tool scoring 4 for Capability Trajectory in a Q1 evaluation should be re-evaluated on this dimension in Q2, not at annual renewal.
Review quarterly at minimum.
Review monthly for fast-moving categories such as AI writing, AI research, and AI outreach.
If a Keep tool’s Capability Trajectory falls from 3 to 1 between evaluations, rescore it immediately. A tool with a composite score of 15 and a Trajectory score of 3 falls to 13 if Trajectory drops to 1, which is still a Keep.
But if Time Recovered and Revenue Enabled have also declined, the composite may fall to 10: Kill. The Monthly AI Spend Dashboard catches this before it becomes a surprise.
Keep With an Active Exit Plan
A tool scoring 2 on Capability Trajectory but 4–5 on Time Recovered and Revenue Enabled is not automatically a Kill. It is a Keep with an active exit plan.
Active exit plan:
Research replacement options now.
Document workflow dependencies.
Set the specific Trajectory threshold that triggers migration.
Start migration if the next evaluation score falls to 1.
Passive monitoring means waiting to see what happens. That approach leads to expensive, reactive migrations after a tool has already deteriorated.
The Engine identifies tools at the Keep/Kill boundary early enough to make migration deliberate rather than disruptive.
Vendor Lock-In Risk Flags
Regardless of composite score, any of these six signals moves a tool from Keep to active exit planning:
Pricing increased more than 20% in the past 12 months without a corresponding capability improvement.
Your workflows depend on the tool’s proprietary output format, making past work difficult to reformat or rebuild.
The tool holds significant data that is not easily exportable, including client data, prompt libraries, trained models, or historical outputs.
The tool is the only access point for a critical client deliverable.
You avoid testing alternatives because switching feels complicated.
The vendor changed ownership, leadership, or investment terms in the past 12 months.
Any one flag requires two actions:
Score Switching Risk at 1, regardless of the previous assessment.
Begin the Migration Playbook even if the composite score remains in the Keep range.
Lock-in risk is separate from ROI. A tool can deliver strong ROI while creating a dependency that is too dangerous to leave unmanaged.
Running This System in Your Current Condition
Contraction: Protect Revenue, Cut With Evidence
When revenue is declining or inconsistent, the instinct is to cancel every AI tool that is not obviously essential. Do not cancel reactively. Without the 5-dimension scoring, operators often cut high-ROI tools because they are expensive, keep low-ROI tools because they are cheap, then rebuild the stack six months later at higher cost.
Minimum viable Engine during contraction:
Complete the Tool Inventory for every subscription within one hour.
Score the three highest-cost tools within the next two hours.
Start the Migration Playbook within 48 hours for every tool scoring below 12.
Protect the tool with the highest Revenue Enabled score, regardless of its cost. During contraction, do not cut infrastructure with the clearest link to revenue just to reduce subscription spend.
Watch for this failure signal: Kill decisions create workflow gaps that require hiring or outsourcing. This means Replacement Cost was scored inaccurately. Pause additional cancellations, rescore, and adjust.
Stability: Turn Reviews Into Decisions
Stable revenue is the best time to install the Monthly AI Spend Dashboard and complete the full-stack audit.
Score every tool.
Give each Keep decision a condition and review date.
Assign one specific integration action to every Upgrade decision.
The main risk at stability is leaving Upgrade decisions in draft form. “Use the tool more” produces no leverage.
Before the next billing cycle, specify:
The workflow to change
The feature to use
The team member responsible
The exact implementation action
Track the number of Keep-with-conditions tools that pass their review date without review. The target is zero overdue reviews.
Anti-Fragility Audit: Remove AI Stack Single Points of Failure
Three single points of failure create dangerous AI-stack fragility. Build redundancy before an outage, pricing change, or operator absence forces the issue.
SPOF 1: One Primary AI Model
If ChatGPT, Claude, or another model is the sole AI layer across proposals, research, client communication, and internal documentation, a 48-hour outage can disrupt all of those workflows at once.
Redundancy protocol:
Maintain active accounts on two AI models.
Keep prompt libraries available in both systems for your three highest-frequency workflows.
Ensure the backup model can handle critical client deliverables for 48–72 hours.
Once per quarter, run one real deliverable through the backup model.
Measure and document the quality gap.
SPOF 2: One Tool Owns a Critical Workflow
A tool that holds your full prompt library, is trained on business-specific data, or is the only path to a client-facing output concentrates outage and pricing risk in one vendor.
Redundancy protocol:
Export your prompt library monthly to a format you own, such as a PDF, Notion document, or GitHub repository.
Keep portable prompts ready for the backup model.
Transfer the workflow to the backup model when needed.
This takes about 20 minutes per month and removes single-vendor dependency from the most portable part of your AI stack.
SPOF 3: The System Exists Only in One Operator’s Head
If the scoring logic, thresholds, and Kill/Keep criteria exist only in the operator’s memory, a health event, vacation, or business transition creates a stack-management gap with no recovery path.
Redundancy protocol:
Document the Kill/Keep/Upgrade thresholds in a one-page reference.
Include the scoring guides in the Monthly AI Spend Dashboard template.
At Scaling band, have at least one team member run the scoring process under supervision before they need to do it alone.
Stress-Test the Stack
Test two scenarios before they become real problems.
Primary AI tool unavailable for 72 hours:
Can you deliver every active client commitment using backup tools and manual processes?
If not, identify the blocked workflow and build its redundancy protocol this month.
AI tool pricing rises 40%:
At Scaling band, $590/month in AI spend becomes an additional $236/month.
Does the tool’s Revenue Enabled score justify the increase?
If not, document the exit plan now, not when the pricing email arrives.
Expansion: Evaluate New Tools Before Renewal
Expansion is when tools enter the stack fastest. New tools are often bought reactively for a new use case, then remain after that use case ends.
Set one rule: Every tool added during expansion must be scored through the full matrix before its first renewal date, not at the next quarterly audit.
The first failure point is dashboard lag. New tools get added faster than they are documented and evaluated, leaving the inventory incomplete.
Do not overvalue Capability Trajectory for newly added tools. New products often look promising because they are early in their development curve, but they have not yet proven Time Recovered or Revenue Enabled in your business.
Guardrail for tools added in the past 90 days:
Track usage from day one.
Wait for a 3-month minimum usage window before scoring Time Recovered and Revenue Enabled.
Use actual 90-day data, not early optimism, for the full evaluation.
The AI Kill/Keep Decision Engine in the AI-First Operating System
The Automation Audit identifies which business tasks are worth automating before you evaluate tool performance. Use this when you have not prioritized AI opportunities.
The Five Numbers shows where AI tool costs and returns belong in core business metrics. Use this when AI ROI is not tracked financially.
The $55K Premature Automation Disaster shows the cost of deploying AI without measurement in place. Use this when tools are running without validated returns.
I Think I’m Paying for Tools AI Already Replaced - The Stack Redesign Map helps you decide which tools to replace, augment, or retain. Use this when your stack has overlapping capabilities.The Stack Redesign Map uses that data to determine which traditional SaaS tools AI has made redundant. Both systems are required to make the complete stack decision.
Your AI ROI Fix Starts Now
What you’ll be able to say at Week 8:
“Every active AI subscription has a composite score. I know exactly which ones are Kill, Keep, or Upgrade - and I have the data behind each decision.”
“The monthly dashboard is running. I catch Capability Trajectory declines before they become expensive surprises.”
“The Kill decisions I’ve made this quarter freed $147/month I’ve redirected to the tool scoring 22 on the matrix. The stack is smaller, cheaper, and producing more.”
30 minutes today:
Pull your last three months of billing statements. List every AI subscription.
Add up the total. That number is your baseline - the starting point for every decision the Engine will make.
This week:
Run the 5-dimension scoring on your three most expensive tools. Calculate composite scores.
Identify the first Kill decision. Initiate the migration playbook.
Before next month:
Monthly dashboard installed. All tools scored.
First Kill actioned. One Upgrade decision with a specific integration change made.
AI Kill/Keep Decision Engine Progress Milestones:
Milestone 1 - Inventory Complete: Every active AI subscription named, costed, and function-documented. Last active use date recorded for every tool. Total monthly stack spend calculated.
Milestone 2 - Scoring Complete: Every tool scored on all five dimensions with composite scores calculated. Kill/Keep/Upgrade designation assigned to every subscription. At least one Kill decision identified.
Milestone 3 - First Kill Actioned: Migration playbook complete for the first Kill-designated tool. Subscription cancelled with rollback plan documented. Monthly savings from the Kill calculated.
Milestone 4 - Dashboard Running: Monthly AI spend dashboard installed and updated for the first time. All composite scores recorded.
Keep-with-conditions tools have review dates set. Upgrade tools have specific integration actions assigned.
Milestone 5 - Trend Line Visible: Dashboard has at least two data points per tool. At least one Capability Trajectory change detected and actioned. Stack is smaller and higher-performing than at Milestone 1.
If you take one thing from each section:
The gut retention bias doesn’t feel like bias - it feels like caution. The scoring matrix is the only thing that separates caution from inertia.
A composite score from 5 to 25 replaces the judgment call. Below 12 is Kill. Above 18 is Upgrade. The number makes the decision before emotion does.
The Kill decision is only as good as the migration playbook behind it. Without the 3-week assessment, Kill decisions produce tool regret. With it, they produce $3,204 annual savings at a 14x ROI on the migration work.
The monthly dashboard is what separates a one-time audit from a governed stack. Two data points produce a trend. Three produce a decision. The gut retention bias can’t survive a trend line.
Capability Trajectory is the dimension that changes fastest. A Keep decision made with a Trajectory of 4 may be a Kill decision 90 days later if the market has moved. The quarterly check is not optional.
But if you remember only one thing:
Operators at $30K-$150K/year aren’t overspending on AI tools because they’re careless - they’re overspending because they have no kill threshold. The Engine installs the threshold. Every tool either earns its composite score or it doesn’t renew.
AI Kill/Keep Decision Engine Checklist
Pull billing statements and run every tool through this system.
☐ List every AI subscription with monthly cost and last active use date
☐ Score your three most expensive tools on all five dimensions
☐ Assign each tool a Kill, Keep, or Upgrade designation from composite score
☐ Initiate the 3-week Migration Playbook for every tool scoring Kill
☐ Install the Monthly AI Spend Dashboard and set 60-day review dates
A complete Engine run takes 60–90 minutes. Every month without it continues paying the gut retention tax on tools that haven’t earned their seat.
FAQ: AI Kill/Keep Decision Engine
Q: How many AI subscriptions do I need before this system is worth running?
A: The minimum viable threshold at Survival band is three active AI subscriptions. Below that, the time investment in scoring doesn’t justify the return. At Scaling band with eight or more tools, the Engine runs monthly as a standard governance practice because gut-feel management at that stack size produces consistent misallocation.
Q: How long does the full stack audit actually take?
A: For a three-to-five tool evaluation, 30–45 minutes. For a complete stack audit of ten or more tools, plan two 60-minute sessions rather than one long block — decision fatigue after the fourth evaluation produces Keep decisions that should be Kill. The monthly dashboard update after initial setup runs 15 minutes.
Q: What if I score a tool and genuinely can’t tell which rating to give a dimension?
A: Score the lower number. The system is designed to produce conservative composite scores, which means Kill decisions are earned rather than triggered by optimism. When you’re uncertain between a 3 and a 4, score 3. An incorrectly estimated Time Recovered score is the single most common cause of reactivating a cancelled tool.
Q: I have tools on annual plans. Do I still execute Kill decisions on them?
A: Yes — but the Kill decision is a workflow decision, not a billing decision. Stop integrating the tool into active work immediately, let the subscription run dormant until renewal, then cancel at renewal.
Q: What makes Capability Trajectory the hardest dimension to score accurately?
A: Most operators default to brand familiarity rather than product evidence. A well-known tool gets a 4 or 5 without a changelog check. The fix is one mandatory step before scoring this dimension: name the last meaningful capability update and the date.
Q: What is gut retention bias and how does it cost me money specifically?
A: Gut retention bias is the pull toward keeping tools you know, regardless of whether they’re still delivering return. At Survival band, one Kill-designated tool at $49/month staying in the stack for twelve months costs $588 annually with zero measurable return.
Q: What are the Vendor Lock-In Risk Flags and when do they trigger action?
A: Six signals escalate a tool from Keep to active exit planning regardless of composite score: pricing increased more than 20% without capability improvement; workflows built around proprietary output formats; significant data held only within the platform; the tool is the single point of access for a critical client deliverable; you’ve avoided testing alternatives because switching.
Q: What happens if I cancel a tool and discover the replacement doesn’t cover everything?
A: The rollback plan handles this. Before any cancellation, document what data needs to be exported, which workflows need updating, and the specific condition that would justify reactivation.
Q: How do I know the dashboard is working rather than just being a record-keeping exercise?
A: The dashboard must produce a decision at every monthly review, not just an observation. If three or more months of data show no composite score changing by more than one point and no new Kill or Upgrade decisions have been made, the system has gone passive.
Q: How does the AI Kill/Keep Decision Engine connect to other systems in the operating framework?
A: The Automation Audit identifies which tasks to automate. The Engine validates whether the tools handling those tasks are worth their ongoing cost — the two run in sequence. The Five Numbers framework is where AI tool ROI belongs as a cost structure line item and return multiplier.
⚑ Found a Mistake or Broken Flow?
Spotted a math error, unclear framework, or broken link? Use this form to flag it — helps me keep the articles accurate and useful. Report a problem →
› More to Explore: Quick Navigation · AI For Operators
➜ Help Another Founder, Earn a Free Month
If the AI Kill/Keep Decision Engine just showed you which tools to cut and which to expand, share it with one founder stuck paying for subscriptions they can’t justify.
When you refer 2 people using your personal link, you’ll automatically get 1 free month of premium as a thank-you.
Get your personal referral link and see your progress here: Referrals
Get The AI Kill/Keep Decision Engine Toolkit
You’ve read the system. Now implement it.
Premium gives you:
Ready-to-use PDF toolkit—every template, diagnostic, and formula pre-filled, zero setup, immediate use
Plug-and-play AI diagnosis sessions—drop into Claude, Gemini or ChatGPT, answer a few questions, save hours of guessing, get your exact next move
Audio key points—concentrated frameworks you can absorb in minutes, implement while you move
Unrestricted access to the complete library—every system, every update
What this prevents: Paying $882+ to Kill-designated tools that run six months unscored.
What this costs: $12/month.
Download everything today. Implement this week. Cancel anytime, keep the downloads.
Already upgraded? Scroll down to download the PDF, audio, and your AI session.



