A successful pilot often has advantages that normal operations will not: an enthusiastic specialist, selected inputs, rapid access to the supplier and time to inspect every surprising result. Production starts when those advantages disappear and customers still expect the service to work.
For Norwegian management teams, I would make the handover to ordinary operations a separate decision. The question is not only whether the AI can perform the task. It is whether the organisation can deliver the service consistently with the people, contracts and controls it actually has.
Test the service with its best supporter absent
Consider a hypothetical Norwegian service company whose pilot champion checks every AI-prepared customer reply. The results look strong. During a week of ordinary staff rotation, other employees cannot tell which source is current, exceptions accumulate and replies wait for the champion to return.
That pilot has demonstrated one person's ability to supervise a tool. It has not yet demonstrated an operating process. I would ask a second trained employee to run a bounded trial using the documented instructions and normal escalation route. Include an absence and an interruption in the exercise rather than hoping they will be easy to handle later.
Agree on a release decision before the final demonstration
My proposed release review has five questions:
- Quality: Does the system meet the defined task requirements on a representative set of cases, including difficult Norwegian-language examples and cases where it should decline or escalate?
- Capacity: Can the team perform the required checking at normal and peak demand? Who handles the unresolved queue?
- Control: Are access rights, data use, supplier arrangements and change responsibilities documented and implemented for this deployment?
- Continuity: Can the service be paused and the fallback operated without losing, duplicating or misrouting work?
- Ownership: Is there an accountable operational owner, a support route and a budget for maintaining the service after the project ends?
The questions are deliberately about the whole service. An average accuracy score cannot answer them. Separate serious failure types from minor wording errors, and define acceptance criteria according to consequence. Passing a test set is evidence within that test's scope, not proof of universal reliability.
Version the process that people depend on
A working AI service includes more than its model. Instructions, connected documents, permissions, retrieval settings and review rules can all affect the result. Record the configuration used during acceptance testing so a later change can be understood and, where possible, reversed.
The NIST generative AI risk-management profile provides voluntary guidance on identifying and managing generative AI risks. My operating recommendation is to keep a small, relevant regression set: representative cases that are rerun when material components change. It should evolve as the business discovers new failure modes, with suitable protection for any personal or confidential data.
The point is not to freeze the service forever. It is to prevent a supplier update or a revised knowledge document from silently changing what the company has agreed to deliver.
Start production with a controlled scope
Expand the number of users, case types and permitted actions separately where practical. If all three change at once, management may struggle to explain a deterioration. Choose review intervals that fit the service's impact and rate of use; a fast-moving customer workflow may need attention much sooner than a monthly meeting.
Keep a visible route for staff to report uncertainty. If reporting a poor output is treated as resistance to change, the organisation loses the evidence it needs to improve. Explain who can pause the process and when that authority should be used.
Before live personal data enter the system, assess relevant privacy obligations, including whether GDPR Article 35 requires a data protection impact assessment. An internal release meeting cannot waive a legal requirement.
A launch needs an owner after the applause
Record the accepted scope, unresolved limitations, responsible owner and next review. If a condition is unmet, make the choice explicit: reduce scope, postpone release or stop. Do not convert a pilot into a permanent exception because its project budget has run out.
Use the incident response playbook to practise failure and the benefits guide to test whether the service improves work. Production readiness is the organisation's ability to keep doing both.
Sources and scope
Sources checked on 11 October 2026. The readiness review and operating example are my recommendations and a hypothetical scenario. NIST guidance is not a Norwegian certification or a substitute for applicable legal requirements.

