Gartner predicts that by 2028, 70% of enterprises will abandon agentic AI built by vendor forward-deployed engineering teams, “trapped by soaring costs and unable to evolve it on their own.”[1] I run a forward-deployed engineering firm, so I take that sentence personally. It describes the moment I think most engagements plan for least: the forward-deployed engineering handover, when the vendor’s engineers leave and the client’s people have to run, pay for and change what was built.
A forward-deployed engineering handover is the transfer of a system, and of the ability to change it, from the vendor’s embedded engineers to the client’s staff: code, accounts, evaluations, runbooks and working knowledge. Gartner names two causes: soaring costs and a system the client cannot evolve. This post turns both into checks a client can run before the vendor leaves: a twelve-item ownership test, and a number I have not seen anyone publish, the time to first unaided change.
What did Gartner actually predict?
The 70% figure, reported on 30 September 2026, is a forecast about 2028, not a measured failure rate, and none of the coverage I read explains how Gartner arrived at it.[1] It does name a cause. In The Register’s summary, “FDE engagements can fail when customers do not acquire the knowledge and control needed to maintain and develop the resulting systems after the vendor leaves.”[2]
Gartner’s Mukul Saha defines success the same way: “Success is measured not by implementation completion, but by the enterprise’s ability to manage, optimise, and scale the technology”.[1] A second prediction targets vendors: through 2028, less than 20% of FDE engagements will turn recurring customer needs into capabilities in the vendor’s core product, “exposing a growing risk of ‘FDE washing’”, its term for consulting services marketed as FDE.[1]
Saha’s remedy is practical: “FDE success starts with getting the engagement structure right, from scope and incentives to governance, ownership, and exit.” He adds: “The best-scoped FDE engagements have clear policies on governance, business value delivery, IP ownership, project co-ownership, knowledge transfer, and an exit strategy from day one.”[3] I read that as a specification, and the rest of this post rewrites it as checks.
Demand for forward-deployed engineers is climbing. Lightcast found that US job postings for forward-deployed engineers “rose from approximately 200 in 2024 to around 1,200 in 2025”. “More than 5,200 postings appeared in the first seven months of 2026 alone.”[4] Indeed data reported by Business Insider shows the same curve: “By April 2026, postings had surged to 5,230% above January 2025 levels”.[5] More embedded hiring suggests more handovers ahead.
Everest Group, writing about Palantir, states the tension well: “The intent is autonomy. The reality is often a long-term relationship, not because of dependency, but because Palantir’s product roadmap keeps offering new capabilities”.[6] That describes one company, not lock-in, and a long relationship is healthy when the client could leave and chooses to stay. The trouble starts when firms copy half the model. a16z quotes observers warning that “if you only copy the embedded-engineer part, you end up with thousands of bespoke deployments that are impossible to maintain or upgrade.”[7]
Why systems die after the builder leaves
None of these forces are new. They get sharper when the system is an agent and the people who understand it work for someone else.
Maintenance is the real bill
Sculley and colleagues wrote in 2015 that “we find it is common to incur massive ongoing maintenance costs in real-world ML systems.”[8] Agents add prompts, tool definitions, model versions and eval data to the list of things that decay when nobody tends them.
Knowledge sits with a few people
A system’s truck factor is the smallest number of people whose departure would leave part of it with nobody able to work on it. Avelino and colleagues estimated it with an automated method for 133 popular GitHub projects and found that “the majority of our target systems (65%) have TF <= 2”.[9] A vendor-built agent often starts with a client-side truck factor of zero: everyone who understands it is on the vendor’s payroll until the handover changes that.
Coding agents add a twist: Mehra and colleagues argue in a July 2026 arXiv preprint that “changes the agent executes that the developer cannot fully understand accrue over time.”[10] It is an argument, not a measurement, but no handover document can transfer understanding the vendor’s own engineers never had.
Nobody is watching after launch
RAND found in 2024 that “Organizations that quickly move from prototype to prototype often find that they are completely blind to failures that arise after the AI model has been completed and deployed.”[11] In a survey of developers running agents in production, Pan and colleagues report that “Reliability (consistent correct behavior over time) remains the top development challenge”.[12] A client that cannot run the evals cannot see reliability move.
Costs and models move under you
Cost has its own mechanism. Gartner has also said, via The Register, that “routing a task to an agentic reasoning model increases inference costs at least fivefold”.[13] Routing rules are configuration. If the client cannot see spend per task or edit the rule, nobody on their side can act.
Models also retire on someone else’s schedule. Anthropic’s deprecation policy commits, for customers with active deployments, to “providing at least 60 days’ notice before model retirement for publicly released models.”[14] A team that can swap a model and re-run its evals treats that as a sprint. A team that has never deployed the system calls the vendor back.
DORA’s 2025 report calls AI “an amplifier, magnifying an organization’s existing strengths and weaknesses.”[15] A handover works the same way: a team that can change a system improves it, and a team that cannot watches it decay. The simulator plays out eighteen invented months after a vendor leaves.
Handed over at exit
The vendor has just left
Pick what was handed over, then press Play.
0 absorbed0 needed the vendorcost trap clearcannot evolve clear
- M2Operations asks for a new refusal ruleevolve
Needs: ✓ repo✗ evals
Arrives in month 2.
- M4Model retirement notice, 60 daysevolve
Needs: ✗ model watcher✗ evals
Arrives in month 4.
- M7Volume doubles; hard tasks route to a reasoning modelcost
Needs: ✗ cost dashboard
Arrives in month 7.
- M10The engineer who knew the system resignsevolve
Needs: ✗ second engineer✗ runbook
Arrives in month 10.
- M13A new rule: log every automated decisionevolve
Needs: ✓ repo✗ second engineer
Arrives in month 13.
- M16Answers drift after an upstream data changeevolve
Needs: ✗ evals✗ runbook
Arrives in month 16.
The repository alone absorbs nothing, because almost every event needs someone who can change the system and a way to prove the change is safe. The failures also sort into Gartner’s two causes: a few arrive as a bill, most as a change nobody can make.
The ownership test for a forward-deployed engineering handover
The ownership test is a set of yes-or-no acceptance items the client must pass before the vendor leaves. We write it into the statement of work on day one, so both sides know what done means and the vendor’s incentive points at the client’s independence. Twelve items sit in five groups.
| Group | What must be true | Closes |
|---|---|---|
| Code and IP | Client owns the repo, cloud account, model keys, billing and prompts | costcannot evolve |
| Evals | Eval set in the client repo, re-run by a client engineer | cannot evolve |
| Operations | Rehearsed incident, cost dashboard on client billing, model fallback | costcannot evolve |
| Knowledge | Two client people per job; an unaided change shipped | cannot evolve |
| Exit | Dated exit plan from day one; vendor components licensed beyond exit | costcannot evolve |
How the test runs
- Agree the items and an exit date in the statement of work, before any code.
- The vendor produces evidence as the work goes, not in the final week.
- A client engineer who did not build the system runs the remove-the-vendor drill; the vendor observes and logs any help.
- Score each item yes or no. Partial means no, with a dated fix.
- The vendor leaves when every item passes, or when the client accepts named gaps in writing.
3/12
Readiness
Fail. The vendor is still required: the eval set and a second person are missing.
Gartner’s two failure modes
1 of 4 guarding items in place
3 of 10 guarding items in place
By group
- Code and IP3/3
- Operations0/3
- Knowledge0/2
- Exit0/2
3 of 12. Fail. The vendor is still required: the eval set and a second person are missing. Weakest group: Evals. Cost trap open; cannot-evolve open.
The gates are the items I will not trade: the repository, the accounts, the eval set and a second client-side person. The rest can pass with a dated fix; without those four, the client cannot start fixing anything.
Knowledge: count people, not documents
For each job the system needs (deploy, roll back, debug a bad answer, change a prompt or tool, swap the model), count the client staff who can do it unaided. The smallest count is the client-side truck factor, and the test asks for at least two.
| Person | Deploy2 | Roll back3 | Debug1 | Change1 | Swap model1 | If they leave |
|---|---|---|---|---|---|---|
| Vendor FDEbuilt itonly one: debug, change, swap model | TF 0 | |||||
| Engineer Aclient platform | TF 1 | |||||
| Engineer Bclient platform | TF 1 | |||||
| Ops leadruns the queue | TF 1 | |||||
| Analystowns the eval set | TF 1 |
Truck factor
1
One resignation from stranded: only the vendor FDE can debug a bad answer, change a prompt or tool and swap the model. Without the vendor it drops to 0.
For scale: Avelino and colleagues found that 65% of 133 popular GitHub projects had a truck factor of two or less, estimated with an automated method. This figure counts capabilities instead, a simplification.
The vendor switch is the point: with the vendor on call, the matrix looks covered. The truck factor that counts is the one without them.
The remove-the-vendor drill
The drill turns items into evidence. A client engineer who did not build the system does three things while we watch and do not type:
- Ships a change to a prompt, tool, eval or config, through the pipeline to production.
- Re-runs the full eval suite and explains one failing case.
- Swaps the model for the named fallback, as if a retirement notice had landed, and shows the evals still pass.
The third step sounds dramatic and is not. Under a 60-day notice policy like Anthropic’s, a model swap is a scheduled event,[14] and rehearsing it makes it routine. Our guide to what forward-deployed engineering actually means covers the wider drills, such as restoring a backup; these three are specific to agents.
Time to first unaided change
Time to first unaided change is the number of days from the handover date until a client engineer ships a production change (a prompt, a tool, an eval or a config) with no vendor contact. It borrows from DORA, which defines change lead time as “The amount of time it takes for a change to go from committed to version control to deployed in production.”[16] DORA measures how fast any change flows. This measures when the first change flows without the vendor.
I like it because it is the part of a forward-deployed engineering handover that is hardest to fake. A handover document can look complete while the system sits frozen, and models, prices and data move too often for a frozen agent to stay healthy.
How to measure it
- Write the handover date into the contract. The clock starts there.
- Tag every production deploy with its author and any vendor help: a chat reply, a call, a pull-request review.
- Stop the clock at the first deploy that client staff authored, that changes behaviour, and that had no vendor contact.
- If the vendor helped, that change does not count and the clock keeps running.
- Report the number at the exit review, next to the ownership test.
Time to first unaided change: 21 days. Change lead time for that change: 2 days.
We keep the plan as a file in the client’s repository, adapted from the example handover plan on our offering page; targets are agreed per engagement.
# transfer-acceptance.yaml · the ownership test for an agentic system.
# Ownership moves when every gate passes with the client at the keyboard.
system: claims-intake-agent # illustrative name
receiving_team: client-platform # the client's engineers
observer: verne-fde # logs any help; does not type
handover_date: 2027-01-15 # the unaided-change clock starts here
exit_date: 2027-02-26 # vendor leaves if the gates pass
ownership:
code_and_ip:
repository: client-org/claims-intake # full history, not a zip
cloud_account: client # model keys and billing too
prompts_and_tools: versioned_in_repo # nothing lives only in a console
vendor_components: licence-schedule.md # what the client will not own
evals:
suite: evals/ # cases, gold labels, scoring
rerun_by: receiving_team
operations:
runbook: ops/runbook.md
cost_dashboard: client_billing # spend per task, with alerts
models:
- current: primary-model
fallback: fallback-model # evals must pass before exit
watch: provider_deprecation_notices
knowledge:
min_people_per_capability: 2 # deploy, roll back, debug, change, swap
# "Remove the vendor": run by someone who did not build it.
drills:
- id: ship-a-change
task: Change a prompt or tool, run the evals, deploy to production
evidence: [pull_request, eval_report, deploy_record]
pass_if: deployed_by in receiving_team and vendor_contacts == 0
- id: rerun-evals
task: Re-run the full suite and explain one failing case
evidence: [eval_report, written_explanation]
pass_if: report_committed_by in receiving_team
- id: swap-model
task: Move to the fallback as if a retirement notice had landed
evidence: [eval_report_on_fallback, rollback_test]
pass_if: evals_pass_on_fallback and rollback_tested
metrics:
time_to_first_unaided_change:
starts: handover_date
# Stops at the first production deploy of a prompt, tool, eval or
# config change by receiving_team, with no vendor contact on it.
stops: first_unaided_production_change
target_days: 30 # an example; agreed per engagement
support_boundary:
verne_on_call: optional, as agreed in the contract
escalation: client-platform-leadThe observer logs help instead of refusing it, because an honest record of where the runbook fell short beats a clean pass. I know of no published benchmark for this metric, so a missed target is a finding, not a penalty.
What to put in the contract
Most of a forward-deployed engineering handover is decided in the statement of work, before any code exists. The clause most often glossed over is intellectual property. MinterEllison notes that AI “deliverables typically incorporate underlying proprietary components”, such as “pre-existing models, algorithms, software code or datasets”, “which the customer does not own”.[17] That is fine as long as the contract lists those components and grants a licence that survives exit.
Gartner’s advice is to have an exit strategy “from day one”.[3] I would add a model-change clause: who watches for retirement notices, who tests the replacement and who pays for the migration. The table compares a typical statement of work with the ownership-first version.
| Clause | Typical SOW | Ownership-first SOW |
|---|---|---|
| Acceptance | Works at go-live | Ownership test passes |
| Repo and accounts | Vendor’s, transferred on request | Client’s, from the first commit |
| IP | Client owns “the deliverables” | Vendor components listed and licensed beyond exit |
| Evaluations | Results in a slide | Eval set in the client repo, re-run by client staff |
| Model changes | Not mentioned | Named fallback, watcher, migration cost owner |
| Knowledge | Training sessions and a document | Truck factor of two, shown in a drill |
| Exit | A notice period | Dated plan and a first-unaided-change target |
| Support after exit | Retainer by default | Optional, with a written boundary |
The typical column is a composite of common patterns, not any particular contract. None of this makes a vendor unnecessary; a long relationship can be right, as the Palantir example shows.[6] The difference is choice: support should be a service the client buys, not a toll it pays because nobody else can change the system. The same goes for the approval step and the audit trail around an agent: both belong to the client.
Honest limits
Gartner’s 70% is a prediction, and the coverage publishes no method behind it, so treat it as a warning rather than a measurement.[1] The ownership test is our proposal. It draws on research into maintenance cost, knowledge concentration and delivery metrics, but it is not a validated standard, and the grouping, the gates and the targets are judgement calls.
The larger gap is outcome data. We have no published results yet showing that engagements which pass the test run longer or cost less than those that do not. We are collecting that evidence, and I will publish it when there is enough to mean something, including if it shows the test matters less than I think.
Frequently asked questions
Why do vendor-built AI agents get abandoned?
Gartner attributes its prediction to two causes: soaring costs, and enterprises being unable to evolve the systems on their own. Gartner also says engagements can fail when customers do not acquire the knowledge and control to maintain the system after the vendor leaves.
What is a forward-deployed engineering handover?
It is the transfer of a system, and the ability to change it, from a vendor's embedded engineers to the client's own staff. That covers code, cloud accounts, model keys, evaluations, runbooks and the people who can deploy and debug it.
What should an FDE exit plan include?
A dated exit, the list of vendor-owned components and their licence terms, the ownership test items, a named fallback for each model and a target for time to first unaided change. Gartner advises having an exit strategy from day one.
How do you measure whether an AI handover worked?
Count the days from the handover date until a client engineer ships a production change to a prompt, tool, eval or config with no vendor contact. If that number keeps growing, the system is frozen, whatever the handover document says.
What truck factor should an AI system have after handover?
Verne's ownership test asks for at least two client people who can deploy, roll back, debug, change and swap the model. A study of 133 popular GitHub projects found 65% had a truck factor of two or less, so knowledge concentration is common even in open projects.
Is forward-deployed engineering the same as consulting?
Gartner draws a line between them. It warns of "FDE washing", consulting services marketed as FDE, and predicts that through 2028 fewer than 20% of FDE engagements will turn recurring customer needs into capabilities in the vendor's core product.
Sources
- Three-phase solution to avoid agentic AI FDE catastrophe (opens in a new tab) Gartner's 70% and under-20% predictions, and Mukul Saha's definition of success. Predictions, not observed rates.
- 7 in 10 enterprises expected to abandon vendor-built agentic AI by 2028 (opens in a new tab) Gartner's view that engagements can fail when customers lack the knowledge and control to maintain the system.
- Gartner predicts 70% of enterprises will abandon vendor-built agentic AI solutions by 2028 (opens in a new tab) Saha on engagement structure: governance, IP ownership, knowledge transfer and an exit strategy from day one.
- Forward deployed engineer: the AI job (opens in a new tab) US job postings for the role, 2024 to July 2026.
- Job postings for this tech role have grown more than 700% in the last year (opens in a new tab) Indeed posting index for the role, as of April 2026.
- Palantir: inside the category of one (opens in a new tab) About Palantir only; not evidence of lock-in.
- The Palantirization of Everything (opens in a new tab) Quotes observers warning against copying only the embedded-engineer part of the model.
- Hidden Technical Debt in Machine Learning Systems (opens in a new tab)
- A Novel Approach for Estimating Truck Factors (opens in a new tab) Truck factors of 133 popular GitHub projects, estimated with an automated approach.
- Agents That Teach (opens in a new tab) An argument about knowledge debt from agent-executed changes, not a measurement.
- The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (opens in a new tab)
- Measuring Agents in Production (opens in a new tab) Survey of developers who run agents in production.
- Agentic AI costs set to balloon fivefold by 2028 (opens in a new tab)
- Model deprecations (opens in a new tab) Vendor policy for publicly released models and customers with active deployments.
- 2025 State of AI-assisted Software Development (opens in a new tab)
- Software delivery performance metrics (opens in a new tab) Source of the change lead time definition.
- Implementing AI in 2026? Here's how to get contracts right (opens in a new tab)