Anyone trying ChatGPT, Claude or Microsoft Copilot for the first time is usually impressed. The text flows, the answer comes at once, and it sounds as though someone knows what they are talking about. That is exactly the difficulty: a language model sounds equally sure whether it is right or not.
That is no argument against using it. It is an argument for knowing beforehand what a language model does well and what it does not, and for building its use around that. This article describes the limits that come up again and again in conversations with businesses, without jargon and without alarm.
It invents when it does not know
A language model guesses which word fits best next. That works surprisingly well, but it has a consequence: if the model lacks the information, it invents something that sounds plausible. An item number that does not exist. A clause that was never written. A delivery condition that might suit your supplier but is not yours. The technical term is hallucination. The model is not lying; it simply does not know that it does not know. For your business this means: every detail you do not know yourself has to be checked before it goes into a quote or an invoice.
Arithmetic and counting are not its strength
A language model is not a calculator. It has learnt how numbers appear in texts, not how to calculate with them. Small sums it usually gets right; in longer calculations, totals over many line items or unit conversions, errors creep in, quietly. Good assistants get round this by handing the arithmetic to a program and only putting the result into words. Where it is not built that way, every total needs rechecking, most easily in the ERP or in the spreadsheet where the figures already sit.
It does not know what happened recently
A model is trained up to a certain point. What happened afterwards it does not know, unless someone hands it over: the new price list, the supplier that has recently stopped delivering, the regulation that changed. Without that, it answers from a state that may be out of date, and does not say so.
It reads rules and contracts like prose
In contracts, standards and internal rules, every word counts. A language model summarises well and finds passages again, but it does not reliably check whether a clause applies or an exception holds. A summary of a contract is a way in for the person who then reads it, not a substitute. Where legal questions are involved, the answer belongs with your lawyer or tax adviser anyway.
It does not know your business
The model has never seen your customers, your items or your workflows. Ask it for a customer’s payment terms and it guesses. Only once the right documents or access to the ERP are handed to it does it answer from your data. How that works, and why your filing decides the outcome, is described in How AI reads your documents. Without that connection, an assistant is a very well-read stranger.
It cannot say why
Ask an experienced member of staff why they held back an invoice and you get a reason. Ask the model and you get a justification that sounds good but is not necessarily the one that led to the answer. For decisions you will later have to explain, to a customer, an auditor or your own team, that is not enough. A result becomes traceable only when the source is attached and a person has checked it.
Responsibility stays with people
From all of the above follows the most important point: an AI cannot take responsibility. If the quote was wrong, the invoice matched to the wrong order or the reply to the customer out of place, that is your business, not the model. That is not a weakness of the technology but a property of tools. The drill is not to blame for the hole in the wrong place either.
What follows for its use
The limits are no reason to keep AI out of the business. They are a set of building instructions:
- Proposal instead of decision. The AI matches the invoice, drafts the text, pre-sorts the mailbox. A person decides.
- Sign-off before anything goes outside. No text to customers, no purchase order, no payment without a look from someone who answers for it.
- Check figures or have them calculated. Totals come from the system, not from the text.
- Show the source. Every answer from documents names the document it came from.
- Hand over current data. Price lists, stock levels and rules come from the ERP, not from the model’s memory.
- Start small. With a task whose result is easy to check. How to recognise one is described in Choosing an AI pilot project.
Anyone who has built an application with AI tools themselves knows a further kind of limit: not the model’s, but that of what it built. More on that in Built it yourself with AI.
Questions to ask of every AI proposal
- Where does this detail come from, and can I open the source?
- Is the figure calculated or guessed?
- Is it current?
- Who checked it before it goes out?
- Could I explain to the customer how this decision came about?
Build these five questions into the workflow and most of the limits are defused, leaving the strengths to use: reading, sorting, summarising, drafting, finding.
