“We have a terabyte of data. Let’s connect it to AI.” Storage volume tells you little about the work involved. Videos, scanned documents, spreadsheets and compressed archives need different treatment. There is no universal conversion from a terabyte on disk to a number of billable tokens.
Start with the business question
Consider a fictional industrial maintenance team. A technician needs the procedure that applies to a particular equipment model. The answer depends on the reference, version and sometimes the site. A similar but obsolete procedure can be more harmful than no answer.
The first workshop should establish the expected questions, authoritative documents, users and unacceptable errors. That defines a representative corpus. You do not need to reorganise the entire information system before testing whether an initial use case has value.
Prepare data to retrieve the right evidence
Retrieval-augmented generation retrieves information before producing an answer. Depending on the corpus, exact-term search and meaning-based search can work together. Selected material then reaches the model. The technical choice needs testing; a vector database is not mandatory for every use case.
For the maintenance example, we would propose the following work:
- Inventory: identify locations, formats, owners and access restrictions.
- Extract: make text usable and check the handling of scans, tables and reference numbers.
- Qualify: distinguish versions, remove unwanted duplication and isolate outdated material.
- Structure: preserve equipment references, dates, status and provenance alongside the content.
- Evaluate: check that business questions retrieve the right evidence, including cases in which no answer should be provided.
An index helps retrieve knowledge. It does not automatically correct an inaccurate document or replace its business owner. The corpus needs to evolve without losing track of its versions. A document that cannot be extracted reliably should remain a visible exception to resolve rather than silently disappear from the project.
Separate preparation from usage costs
We recommend budgeting three groups of work. The first is preparation: connectors, extraction, checks, indexing and exception handling. The second is queries: retrieval, possible reranking, model calls and generation. The third is operations: updates, testing, monitoring and maintenance.
An architecture that restricts the information sent with each request can control context volume, but still incurs preparation costs. Conversely, exploring data on demand may avoid some initial processing while increasing work at query time. The right trade-off depends on the actual task.
What the engagement should deliver
Mintera proposes to define the need, map the corpus, establish governance rules and design an appropriate architecture. The initial scope tests these choices with your teams before expanding the system. It also identifies who will maintain the knowledge after the first deployment.
Expected deliverables are specific: a diagnostic, preparation rules, a reference set of questions, evaluation results and an expansion roadmap. We recommend tracking the relevance of retrieved documents, evidence-backed answers, justified refusals, processing time and cost per completed task.
Measure completion from the user’s point of view, including any corrections they must make. A short generation time means little if a specialist must rebuild the answer afterwards. The objective is a useful system whose quality and operating effort can be assessed.
