The first workflow a leadership team automates teaches the organisation what AI adoption means. Choose a task with visible value, manageable consequences and a clear owner, and the team can learn from real work. Choose the most politically attractive demonstration, and it may learn to hide the difficult cases.
My recommendation is to start by examining exceptions. A workflow that looks repetitive from a meeting room can contain ambiguous customer commitments, old data and unwritten judgement. Those details determine whether AI assistance will reduce work or move it to somebody else.
Describe a task small enough to observe
“Use AI in customer service” is a department-level ambition. “Prepare a draft reply to delivery-status enquiries using approved order information, with a person sending it” describes a workflow. The second version identifies inputs, an output, a reviewer and a boundary.
In a hypothetical Norwegian distribution company, three candidates are available: drafting delivery updates, deciding credit limits and summarising supplier manuals. I would not rank them by how impressive the demonstration looks. I would compare the consequence of an error, the availability of reliable source material and the work required to check the result.
Credit decisions affect people and commercial exposure. Manual summaries may be useful but require accurate treatment of technical warnings. Delivery drafts may be more bounded, provided they do not invent promises or expose another customer's order. None is safe merely because it produces text.
Examine the awkward cases before scoring the average
Take a representative sample of work, using information you are allowed to process in the evaluation. Include the normal cases and the cases employees find difficult: incomplete order references, conflicting dates, complaints, mixed Norwegian and English terminology, and requests the business should refuse or escalate.
Ask the employees who handle these cases to explain what makes them difficult. They may know that a familiar product name refers to two versions, or that a customer agreement overrides the standard procedure. That knowledge belongs in the process design, not in a private workaround used by the best reviewer.
Where personal data are involved, GDPR Article 5 includes purpose limitation, data minimisation and accuracy requirements; Article 6 concerns a lawful basis. A trial is still processing. Use genuinely non-personal or appropriately prepared test material where possible, and assess the actual use before connecting live data.
Use a selection sheet with reasons
For each candidate, record the following in plain language. A rating is only useful if another person can challenge its reasoning.
- Value: What completed outcome improves, and how often does the task occur?
- Source quality: Is there an authoritative answer or record against which the result can be checked?
- Consequence: What could an incorrect answer or action cause, and who would be affected?
- Reviewability: Can the person checking the output recognise the important errors within the available time?
- Reversibility: Can the work be paused, corrected or returned to the existing process?
- Ownership: Who controls the workflow and has time to fix it when the pilot exposes weaknesses?
I would reject a first candidate if there is no credible owner or no practical way to identify serious errors. A large expected benefit does not repair those gaps. Consider simplifying the workflow or using ordinary rules-based software first.
Define the smallest useful pilot
For the delivery example, limit the pilot to a defined enquiry type, approved sources and draft-only output. Record today's handling time, error patterns and escalations before the change. Decide what the assistant must do when information is missing. An honest request for clarification can be a successful result.
Evaluate complete cases in the languages customers actually use. Where relevant, include Bokmål, Nynorsk and English, rather than assuming a good English demonstration establishes Norwegian performance. If spoken input is involved, test the relevant speech variation separately.
The first pilot should answer a specific question, such as whether verified drafting reduces total handling time without increasing incorrect delivery promises. The production-readiness guide then helps determine whether the result is ready for ordinary operations.
Sources and scope
Sources checked on 11 October 2026. The selection sheet is my proposed management tool, and the distribution company is hypothetical. A favourable score does not establish legal compliance or replace a use-case-specific assessment. For Norwegian integration questions, see the guide to internal AI agents and GDPR.

