Different tasks need different capabilities

Sorting messages, forecasting demand and drafting a paragraph are different tasks. A system suited to one is not automatically suited to another. Generative AI creates content such as text or images, while other AI applications may classify information or estimate a likely outcome. Avoid treating a convincing demonstration in one area as evidence of broad competence.

Describe the input, expected output and permitted action for your own task. If the system suggests a response, that is different from allowing it to send the response. Keeping those boundaries explicit makes the discussion about usefulness far more concrete than asking whether the company should use AI in general.

AI and automation are not the same thing

Automation concerns work being carried out with reduced direct intervention. It can use simple explicit rules without AI. A reminder sent on an agreed date does not require a system to interpret meaning. AI may be useful where interpretation or pattern recognition is part of the task, but that flexibility introduces uncertainty about individual outputs.

The two approaches can be combined: a model proposes a category, a person confirms it and a rule routes the item. Compare this with a simpler rule or an improved form. The most suitable design may use AI for only one limited step.

An illustrative customer support draft

Imagine a fictional outdoor equipment shop receiving questions about product care. The team considers an AI tool that drafts replies from approved care instructions. A helpful pilot would include a straightforward question, an ambiguous product name and a request not covered by the instructions. The desired behaviour in the last case may be to ask for clarification or refer the question to a colleague. The draft should not invent a care procedure to appear helpful.

This example describes a test design, not a claimed result. A fluent answer is only useful if the facts, scope and proposed action fit the customer’s actual situation.

Three test cases for a reply draft

Before validation, define what would count as correct. These cases provide a starting point for the fictional outdoor shop. They do not establish an accuracy rate for all customer questions. Add further typical and difficult cases from the real task and repeat the checks after changes.

For classification, also distinguish the consequences of a wrong routing decision from an additional manual review. Record error types separately: a good average would not make an individual harmful care instruction acceptable.

  • Unambiguous product, answer in the approved instructions: the draft must match the right model and the relevant passage. Reject invented additional care advice.
  • Ambiguous product name: ask for the model or identifier. Reject an unsupported choice of product variant.
  • Question outside the documents: identify the missing basis and refer to a responsible person. Reject an invented procedure or source.

Confident wording is not evidence

A generated statement can sound specific while being unsupported. Treat names, references, calculations and factual claims as things to verify against suitable sources. When a system works with company documents, check whether it has retrieved the relevant passage and interpreted it correctly. Providing documents does not make every answer reliable by itself.

Good knowledge management helps because current, clearly owned material is easier to use than contradictory drafts. Human reviewers also need time and subject knowledge. A review step becomes weak when the person is expected to approve more material than they can meaningfully inspect or has no way to check the underlying information.

Choose information and access deliberately

Decide what information may enter the tool and which people or systems can access the result. Different services and configurations have different arrangements, so inspect the actual terms and settings rather than assuming all AI tools behave alike. Use fictional or appropriately prepared examples when early testing does not require live customer information.

If the system can take actions, limit its permissions to what the task needs. A tool that drafts a proposal has different responsibilities from one that changes an order. Clarify who can stop it, investigate an unexpected action and restore the affected work when something goes wrong.

Assess the complete work, including review

A quick first draft does not automatically reduce total effort. Someone may spend time checking facts, correcting tone, supplying context and handling exceptional cases. Include those activities in a cost-benefit analysis.

Compare the outcome with the existing method and a simpler alternative. Ask whether quality is at least adequate and whether any saved time can be used meaningfully. Also consider how the arrangement will be maintained when source material or working practices change. A bounded use that is easy to inspect may provide more practical value than a broad system whose outputs require extensive correction and explanation.

Start with a reversible, observable use

Choose a task where errors can be detected before they cause consequential action. Define the owner, the source material, the review criteria and the point at which a person takes over. Keep examples of both useful and unsuitable outputs. After the pilot, decide whether to continue, narrow the scope or stop. Revisit the evaluation when the tool or task changes; an earlier successful demonstration does not settle every future use. The aim is a dependable working arrangement in which people understand what the system contributes and retain responsibility for decisions that require context, judgment or commitments to others.

  • Case and source document: [identifier and version, without unnecessary customer data].
  • Observed error or supported answer: [specific passage]. Checked by: [responsible role].
  • Decision: [accept/correct/escalate]. Actual review and correction effort: [time].
  • Stop rule for this pilot: do not send a draft containing unsupported care instructions. A knowledgeable person handles the case; investigate the cause before further trials.

Common questions

Does AI understand the business like an experienced colleague?

Do not assume that. A system may produce useful outputs without sharing the organisation’s practical understanding or knowing an unstated commitment. Supply relevant context and judge the result against the task rather than its conversational confidence.

Do we need AI for every digital improvement?

No. Clear forms, reliable search, shared records and simple rules can resolve many problems. Start with the difficulty you want to address, then compare approaches. Adding AI is not itself evidence that the work has improved.

Who should own an AI pilot?

Someone who understands the task and can judge the result should share responsibility with whoever manages the technical setup. Users need a clear contact for errors and questions, and the organisation needs authority to change or stop the arrangement.

Sources and further reading