Fifteen hours a week. That is what your team spends re-keying customs and shipment data from emails, PDFs and spreadsheets into a legacy system. You manage operations at a 14-person freight-forwarding and customs-clearance company in Ljubljana. A software vendor has just shown you an AI agent that claims to do this automatically. The demo ran on clean test data. You have three weeks before the budget approval window closes, and a quarterly customs-compliance report in front of you.
The decision is not simply automate or don't. The decision is which of three moves to make now: commission the build, postpone it with evidence, or run a workflow audit first. We recommend the third — run a one-day readiness check and let its numbers choose. The wrong choice costs more than budget; it can put errors into a compliance report you are personally accountable for.
Caution here is not paranoia. The published estimates on automation failure are messy enough that they should not be boiled down to one number. EY is reported as seeing 30–50% of initial RPA projects fail, attributing failures to company methodologies and misunderstanding of the technology rather than the technology itself. An ABBYY survey of 400 senior executives in 2020 found 38% say RPA projects fail because they are too complex, and roughly 30% say the intended automation processes were not understood or the underlying tools were not well understood. A journal article published on 5 October 2025 asserts a high (~50%) failure rate for RPA initiatives. The sources disagree on the cause: EY's reported position is that companies misuse the technology, while ABBYY respondents blame project complexity and insufficient process understanding; these are different populations and different measures, and should not be merged into one statistic.
The four-question readiness filter
The framework below turns "not ready" from a feeling into a documented decision. It has four questions. All four must pass before automation is worth building. Any single red flag routes to postpone and audit.
| Question | One-day measure | Red flag | Decision route |
|---|---|---|---|
| 1. Stability | List agreed rulebook or system changes in the next 12–24 months. For customs: the UCC three-year retention baseline and the EU Data Hub timeline (voluntary 2030, compulsory 2032). | A material rulebook or system change is already agreed or in flight. | Postpone; re-run the filter after the change lands. |
| 2. Exception rate | Pull a defined period of real incoming emails, PDFs and spreadsheets; count total items and the items needing human interpretation or a different rule; divide exceptions by total. | Exception handling and manual checks would consume more than roughly a third of the hours automation is supposed to save. | Postpone; audit the process to reduce exceptions first. |
| 3. Human-judgement content | List every step where a person currently interprets, decides, or approves; require the vendor to show how those steps stay human. | Any compliance-relevant or irreversible step depends on judgement with no human approval gate. | Postpone; require verification and approval gates in the design. |
| 4. Data quality | On the same sample, count fields with missing, inconsistent, or multi-format values; compute the clean-hit rate. | Clean-hit rate is too low for the legacy system's import rules. | Postpone; clean the data or redesign the input first. |
One red flag is enough to stop the purchase. That is the point of the filter: it prevents a clean demo from deciding for you.
Question 1: Is the rulebook stable?
For a customs data-entry process, stability has two parts: the legal baseline you must preserve, and the rulebook changes already in flight.
The baseline is Article 51 of the Union Customs Code (Regulation (EU) No 952/2013). It requires the person concerned to keep the documents and information covered by Article 15(1) for at least three years, "by any means accessible by and acceptable to the customs authorities." For goods released for free circulation or declared for export, that period runs from the end of the year in which the customs declaration was accepted — not from the transaction date. Any automation you build must preserve that retention rule, including the access and acceptability requirement.
The changes in flight are equally concrete. In May 2023 the European Commission proposed a fundamental reform of the EU customs framework; on 26 March 2026 the Council and the European Parliament reached a political agreement on it. The agreed reform introduces a new decentralised EU agency, an EU customs data hub, enhanced simplifications for the most trusted traders, and a new EU-wide handling fee. According to the customs reform calendar circulated by the Commission in December 2025, voluntary use of the EU Data Hub by all traders together with the Trust & Check regime is set for 2030 (previously 2032), and compulsory use for all traders in 2032 (previously 2037). These are politically agreed reform plans and planning targets, not legally binding milestones, and they may change before formal adoption and implementation.
If your integration would hard-code today's declaration flow, ask the vendor what breaks when the Data Hub becomes compulsory. A stable rulebook is not one that never changes; it is one where the changes you know about have enough lead time to plan for.
Questions 2 and 4: Measure the process in one day
Do not trust a demo. Pull your real inputs instead. This five-step sampling protocol is a screening test, not a full process audit, but it produces the two numbers your decision needs.
- Pull a defined period of real incoming emails, PDFs and spreadsheets — the last four weeks is usually enough.
- Count the total items. That is your denominator.
- Count the exceptions: items that need human interpretation or a different rule. A shipment split across two emails, a customs procedure you see once a quarter, a supplier who formats weights differently every time.
- On the same sample, measure field-level data quality. For each record, flag fields that are missing, inconsistent, or stored in multiple formats — dates saved as both 12/03/2026 and 2026-03-12, for example.
- Compute two ratios:
exception rate = exceptions ÷ total items
clean-hit rate = clean records ÷ total records
Suppose the sample holds 200 items, 45 of them exceptions, and 130 records with every required field in one format. The exception rate is 22.5% and the clean-hit rate is 65%. Those two numbers go into the decision record.
Now the honest part: what exception rate is too high?
Here is ours, plainly labelled as a professional judgement. If exception handling and manual checks would consume more than roughly a third of the hours automation is supposed to save, the build is not paying for itself. For a process that currently takes 15 hours a week: if eliminating the core data entry still leaves you spending more than about five hours a week on exceptions and output checking, you have not recovered the time you need to justify the build and its maintenance.
Total-cost decision worksheet
Record these before you commission anything:
- One-off costs: build, legacy-system integration, testing, data cleaning, staff training.
- Recurring costs: software, human review, and maintenance when rules or document formats change.
- Recovered time: the fully loaded hourly cost of the hours you would recover, and the expected weekly hours recovered after review.
The build proceeds only when the expected benefits cover both one-off and recurring costs within your company's stated payback period. Otherwise, the decision is postpone or audit.
Question 3: Where does human judgement still matter?
Compliance-sensitive processes fail in the places where a person currently interprets a document, decides whether an exception applies, or approves an irreversible action. List every one of those steps before the vendor visits again.
On the narrow question of whether the law requires a signature: within the returned text of Article 51, no requirement for a human signature on customs records appears. That is a fact about this article only; it does not establish that no human involvement is required elsewhere in the UCC. Do not let a vendor turn that limited finding into a general claim that the whole process can run without a person.
For every step on your list, require the vendor to show how the step stays human. If the answer is "the model handles it", that is not an architecture; that is a risk you will discover during the compliance report.
Before you commission a build: this filter does not determine legal sign-off requirements outside Article 51. List every compliance-relevant decision in the process, and obtain written confirmation from the responsible customs-compliance adviser or authority stating which steps may be automated, which require human approval, and what audit trail must be retained before commissioning a build.
Send the vendor this test before you commit
What the vendor needs before quoting
- The legacy-system name and its import method.
- A de-identified, representative sample of your real emails, PDFs and spreadsheets.
- Weekly input volume and peak periods.
- Required output fields and their mappings.
- Exception categories and how they are handled today.
- The named humans who approve compliance-relevant actions.
- Required audit-trail and record-retention outputs.
- Known rule or document-format changes in the next 12–24 months.
A quote issued without these inputs is only a preliminary estimate.
Before any commitment, send the vendor this test. It is written to be copied into an email.
A serious automation for this process should be a supervised agent system: a system in which AI models work inside bounded stages, and their output is treated as a claim that must pass checks before it counts as done. A verification gate is a deterministic check — code that passes or fails on explicit rules — placed before a human approves an irreversible or compliance-relevant action.
Our published AI-automation doctrine is the standard to compare against: "We never trust model output blindly. Every automation ships with deterministic verification checks, budget ceilings, and human approval gates for irreversible actions. We build systems where AI output is treated as a claim that must pass checks before it counts as done."
The demo must do four things:
- Run on a de-identified sample of your real emails and PDFs, not the vendor's curated screenshots.
- Show what happens to exceptions: which ones route to a queue, which need a human, and how the system marks them.
- Name the verification steps and human approval gates for irreversible or compliance-relevant actions.
- Explain what breaks when rules or document formats change, and how the system detects those changes.
We publish an example of this architecture in practice — a content-agent platform with 14 stages and deterministic validators — for reference if you want evidence before a call.
The case for automation when the filter passes
None of this is an argument against automation. A stable, high-volume, low-exception, well-structured process is exactly the right place to automate. When the four questions pass, the repetitive work is known, the rules are explicit, and a deterministic system can check its own output before anything irreversible happens.
The standard a build should meet is the same regardless of vendor: verification checks, budget ceilings, and human approval gates for irreversible actions. That is an engineering requirement, not a sales line. Our AI automation service is built around this approach — mapping a company's workflows and building supervised systems with verification and human approval built in.
Write the "not yet" decision as evidence
If your filter returns a red flag, write the decision down. A one-page record is enough for a director:
- Process: customs and shipment data entry.
- Sample: 200 items from a defined four-week period; state the dates.
- Exception rate: 22.5%.
- Data-quality score: 65% clean records.
- Stability: EU customs reform politically agreed on 26 March 2026; EU Data Hub voluntary in 2030, compulsory in 2032.
- Conditions that would flip the decision: exception rate below the break-even point, clean-hit rate above a specified level, or the Data Hub timeline slipping further.
That document turns a "no" into a business case. When your director asks why you did not buy the tool, you answer with sample size, exception rate and the reform timeline — evidence, not intuition.
What to do next: audit the process, not the product
The next step after the filter is a workflow audit that starts from your actual process, not from a product. That is how our AI-automation engagements begin: mapping a company's workflows and building supervised AI systems with verification and human approval built in. If your filter passed, the audit verifies your numbers before you commit. If it failed, the audit identifies what has to change before automation becomes viable.
We publish indicative price ranges for reference, but exact costs follow the audit. Book a strategy call to start that review.
Sources
- 01Reducing the High Failure Rate (50%) of RPA Implementation Projects: A Real-World Application Using Design Science Researchdoi.org
- 02RPA projects fail because of complexity, misunderstandingciodive.com
- 03Modernising the EU customs union - Consilium - European Councilconsilium.europa.eu
- 04EU Customs Reform - Taxation and Customs Union - Europa.eutaxation-customs.ec.europa.eu
- 05LIMITE ENdata.consilium.europa.eu
- 06Regulation (EU) No 952/2013 of the European Parliament and of the ...legislation.gov.uk
- 07[PDF] Regulation (EU) No 952/2013 of the European Parliament ... - EUR-Lexeur-lex.europa.eu
- 08Why RPA Implementation Projects Failcmswire.com
- 09What can we learn from RPA failures? - Raconteurraconteur.net
- 10Get ready for robotseyfs.ie
- 11State of Process Mining and Robotic Process Automation 2020digital.abbyy.com
- 12RPA projects fail because of complexity and misunderstanding | Supply Chain Divesupplychaindive.com