A demo has to impress. Daily use has to work

A demo has one job: work once, on a chosen example, in front of an audience that wants to see the “wow” moment. Nobody asks what happens when the document is unreadable, the model responds late, or the input looks nothing like the example prepared for the show.

A system in daily use has to answer those questions every day, for every user, without exceptions. The difference isn’t in the model itself — the same language model sits behind both the demo and the live system. The difference is in what you build around it.

Monitoring and observability

In a demo, nobody checks whether the answer was correct — you can see it with your own eyes, because there’s only one example. In daily use, you need a mechanism that measures answer quality at scale: how often the model refuses, how often the result needs a human to step in, where the edge cases keep showing up. Without that, you don’t know the system is degrading until it fails visibly, which is too late.

Cost at scale

A single query to a language model is cheap when you’re showing it off once. Multiplied by thousands of documents, requests, or transactions a day, it becomes a real line item that has to be designed for deliberately — starting with which tasks genuinely need a model at all, down to how much context you actually feed it.

Model drift and errors

Language models sometimes get things wrong, and they sometimes change behavior between provider versions. A system in daily use has to assume this from day one: a way to catch a bad answer before it reaches a customer or an accounting ledger, and a plan for what happens when the model isn’t confident.

Who’s accountable for the decision

This question never comes up in a demo, and it’s central once a system goes live. If the AI recommends something wrong, who’s accountable — the model provider, the company that built the system, or nobody? The answer has to be clear from the start, not worked out after something breaks.

Integrating with what a company already has

A demo usually runs in isolation — on a prepared dataset, with no connection to a real accounting system, CRM, or customer database. After go-live it’s different: AI has to exchange data with the systems a company already runs, without expecting the team to replace them just because it’s adopting AI. That’s extra work that simply doesn’t exist in a demo, and it can take longer than the model itself once live.

Our rule: AI proposes, code decides

Across the implementations we run, we hold to one rule: AI reads, classifies, and proposes — anything touching money or a critical decision gets calculated by deterministic code. A language model is excellent at making sense of unstructured input — a document, a message, a photo of a shopping basket. It isn’t the tool you hand the final tax calculation or a payment authorization to.

In our view, this separation of roles is one of the main reasons most AI projects stall at the pilot stage: there’s no clear line between what the model proposes and what the system decides.

The proof: three products, one common thread

Rather than talk about daily-use AI in the abstract, we can point to our own products — all three run AI in the layer that operates every day, not just in demos.

In Qkwit, our AI-powered accounting product, the model reads and extracts data from accounting documents, and the result feeds into a deterministic engine that calculates taxes and social security contributions. The model understands the document — the code calculates the liability.

In Taniej po Lek, our basket-based medication price comparison service running in 11 countries, AI recognizes basket contents from a photo or from pasted text and suggests cheaper substitutes — while a deterministic engine always calculates the basket total.

In Brokik, our rental management platform available in 25 countries, AI adapts an existing lease document to a plain-language instruction, and a setup assistant drafts property and tenant records that the user approves. That’s a daily feature, not a conference demo.

What connects all three: the AI is working there every day, on real data, with monitoring and a plan for when it’s wrong — not just once, for a demo.

Why this gap is still common

Many of the AI projects we come across stall at the working-prototype stage. The reason usually isn’t the model — it’s what’s missing around it: no monitoring, no clear ownership of decisions, no plan for cost at scale, no answer to “what do we do when the model is wrong.”

These are engineering and organizational questions, not questions about whether the model is smart enough. We see this especially where the team behind an implementation knows the model well but is less sure how to design the system around it — task queuing, error handling, cost limits. Those elements, not model quality, usually decide whether a project stays a demo forever or reaches daily use.

What this means for your company

If you’re weighing an AI implementation, it’s worth asking from day one not “will this work once” but “what happens when this runs every day, on a thousand cases, with a different person on the other end.” That question leads to a very different set of design decisions — about monitoring, cost, the split between AI and code, and who owns the outcome. We describe how we approach this on our AI implementation page.

Frequently asked questions

What’s the difference between an AI demo and a system in daily use?

A demo shows one chosen case. A system in daily use has to handle every case, every day, with quality monitoring and a plan for errors — no exceptions and no pre-rehearsed scenario.

Can AI make financial decisions on its own?

Not in our implementations. AI reads, classifies, and proposes; deterministic code handles anything involving money or a critical decision — that’s the rule we apply across all of our own products.

Why do so few AI projects reach daily use?

Usually what’s missing isn’t the model — it’s monitoring, a plan for cost at scale, and a clear owner for the outcome. These are typically engineering gaps, not model limitations.

How do you check whether a vendor actually has experience running AI in daily use?

Ask about specific, running products rather than completed demo projects — and ask how monitoring and the split of responsibility between AI and the system actually works there.

Does AI in daily use cost more than a demo?

Yes, because cost scales with volume — which is why a well-designed implementation decides upfront which tasks genuinely need a language model and which don’t.