Data Readiness for AI: What SMEs Really Need — and What They Don't

The most common reason businesses delay AI is a belief that their data is not ready. Usually the data is good enough — it is just scattered, unlabeled, and mixed with material that should have been archived years ago.

Leitspur2 min read
SCATTEREDCLEAN+ LABELREADY
Figure · Data

AI vendors talk about data as if every company had a warehouse, a lake, and a governance team. SMEs have SharePoint folders, a fileserver, an ERP, inboxes, and a handful of spreadsheets that quietly run the business. The good news: for most practical AI workflows, that is enough raw material.

What AI workflows actually consume

Assistants and automations rarely need all company data. They need the narrow slice behind one workflow: the price list behind quoting, the SOPs behind support answers, the contracts behind compliance checks. Readiness is a per-workflow question, not a company-wide score.

  • Current — the latest version is identifiable, and the outdated ones are archived.
  • Findable — the material lives in a known place with sensible names.
  • Permissioned — it is clear who may see what, so the tool can respect the same lines.
  • Extractable — text can be read from the files (scans need OCR, tables need structure).
  • Owned — someone answers for keeping the source true.

The cleanup that pays for itself

A focused cleanup of one workflow's sources typically takes days, not months. Archive stale versions, agree on naming, fix the two or three broken export formats, and write down where the truth lives. Every hour spent here removes confusion for people and machines alike.

What you can safely postpone

A central data platform, a company-wide taxonomy, perfect master data — none of these are prerequisites for a first working system. They may become worthwhile later; letting them gate the first project mostly guarantees the first project never starts.

  1. Pick the workflow and list the documents and fields it truly needs.
  2. Clean and label that slice; archive what is stale.
  3. Wire the slice to the tool with permissions intact.
  4. Assign an owner who keeps the source current after launch.

Readiness shows up in the answers

Once a system is live, its wrong and unanswerable questions map your real data gaps — the same sensor effect we describe for knowledge assistants. Fix the sources the gaps point to, and quality compounds month over month.

Frequently asked questions

Do we need a data warehouse before starting with AI?
No. Most practical AI workflows read from existing documents and systems. Central platforms become relevant later, when several systems need the same governed data.
Our documents are PDFs and scans — is that a problem?
Usually manageable. Digital PDFs extract well; scans need OCR. The bigger risk is outdated content, not file format.
How do we handle sensitive data?
Scope the workflow so the tool only touches what it needs, mirror your existing permissions, and agree processing terms with any external provider before connecting real data.
Contact

Let's build something useful.

Tell us where the hours go. We reply within two working days — with a first read on the smallest system that would win them back.