What if the AI came to your documents instead of the other way round?
Some organisations have quietly opted out of AI entirely. Not because they aren't interested — because the material is confidential, and sending it somewhere else isn't a conversation they want to have.
There are organisations doing exactly the kind of work AI is good at — reading, comparing, summarising, remembering — that have quietly opted out. Not because they aren't interested. Because the material is confidential, and sending it somewhere else isn't a conversation they want to have.
We keep meeting the same situation. Someone inside the firm has already worked out three or four genuinely valuable things an AI assistant could do with the archive. The ideas exist, they are specific, and they are good. What stops them is a short conversation with a partner, a compliance lead or a client that ends in "we're not uploading that."
So the work carries on the old way: someone senior spends an afternoon looking for the version of a clause they know they wrote in 2021.
The blocker is almost never interest.
It is where the documents go.
This note sets out what we have found looking at that problem: why the middle option is usually the right one, what a private system would actually consist of, which organisations it fits, and — just as usefully — which it does not. We are selecting the first organisations to build this with. If the pattern above is your situation, the last section is the part to read.
To be clear about the other side
Cloud AI is not automatically unsafe or non-compliant. The major providers offer enterprise agreements, encryption and access controls, and contract not to train on business data — Google, for instance, states that it will not use your data to train or fine-tune models without your prior permission, and Anthropic's commercial terms take the same position. For a great many organisations that is entirely sufficient, and we would tell you so.
The argument here is narrower: in some environments, moving less data is valuable in itself — because of a regulator, a client contract, a national-security clause, or simply because it removes an objection that would otherwise stall the project for a year.
A system that answers questions without your files crossing the line
"Private" and "offline" are not the same thing
Most conversations collapse this into cloud versus no-cloud. There's a middle option, and it's the one we think fits most organisations.
| Setup | Where your documents are processed | Internet | Best suited to |
|---|---|---|---|
| Public AI services | An external provider's infrastructure | Required | Most general business work, where the material isn't especially sensitive |
| Private, on your side of the line | Hardware you own or control | Connected, but controlled | Professional firms and industrial businesses with confidential archives — the case this page is about |
| Fully disconnected | Hardware you own, with no external network | None | Classified, defence-adjacent or safety-critical environments where isolation is mandated |
Full disconnection sounds like the strongest answer and is usually the wrong one. It costs you updates, backups, remote support and integration with the systems your team already uses. We'd only recommend it where something external requires it.
The hardware is the easy part. The knowledge is the work.
Anyone can put a model on a machine in an afternoon. What determines whether people use the thing in month six is how well the archive was understood at the start. This is roughly the sequence we'd expect to follow.
Find the questions worth answering
A short round of interviews with the people who actually do the work, to identify the searches that cost the most time today and the ones nobody attempts because they'd take too long.
Map the archive
Where the material lives, what condition it's in, what's duplicated, what's authoritative and what should never have been kept. Almost every organisation discovers something uncomfortable here, and it's useful on its own.
Agree the rules before the build
Who is allowed to see what, which material is out of scope entirely, what the system must cite, and what it must refuse to answer. Written down and signed off, not implied.
Install and configure
Hardware sized to the archive and the number of users, models selected for the work rather than for benchmarks, and the retrieval layer that lets the system find the right passage before it answers.
Make it citable
Every answer points back at the source document. Without this, professionals correctly refuse to rely on it — and they should.
Pilot with one team
One group, one set of real questions, measured against how the work is done now. This is the point at which we'd expect to find out whether the idea holds.
Train, then hand over the habit
Short sessions on how to ask well, where the system is reliable, and — more importantly — where it isn't and what to check.
Run it
New documents ingested, models updated, retrieval tuned as the questions change, access reviewed, and someone accountable when it stops working.
What you would get out of it
The objection disappears
Projects that stall at "we can't put that in an external tool" become buildable. For many firms this is the entire value: it unblocks a year of postponed ideas.
Institutional memory becomes searchable
Thirty years of matters, projects or batches stop being an archive nobody opens and start being something a junior can interrogate in a minute.
Senior time goes back to judgment
The expensive people stop doing retrieval. They still make the decision, sign the drawing, advise the client — but they arrive at it with everything relevant in front of them.
Costs behave like infrastructure
A capital item plus a support line, rather than per-seat pricing that grows every time someone new needs access to the archive.
You keep the audit trail
Because it's your system, you can answer what was asked, what was retrieved and what was shown — to a regulator, a client or an insurer.
Accountability stays where it belongs
This is the model we argue for generally: the system holds the context, the professional holds the responsibility. Nothing here is designed to make a decision on your behalf.
Six questions we would ask in the first ten minutes
Nothing is sent anywhere — this runs in your browser and is just a way of thinking about it.
How confidential is the material your team works with day to day?
How much documentation has the organisation accumulated?
How much of the work is finding, comparing and summarising documents?
Would faster answers from your own archive change how the work gets done?
How does the organisation feel about that material going to an external AI service?
Is there a budget line for tools that make senior people faster?
The pattern we're looking for is an organisation where good AI use cases already exist and confidentiality is what's stopping them — not an organisation that's simply curious about AI.
The kinds of organisations this is aimed at
Different work, same underlying shape: a large confidential archive, and expensive people spending hours inside it.
Law firmsInterrogating the matter file
Contracts, correspondence, litigation bundles and precedent, none of which the client expects to be uploaded anywhere.
"Find every change-of-control clause in the acquisition documents, summarise how they differ, and cite the source."
Accounting, tax and payrollClient history without the digging
Years of filings, correspondence and advice held on behalf of hundreds of clients, spread across systems and inboxes.
"What changed between this client's last two sets of accounts, and where did we advise on the VAT treatment?"
Engineering practicesInstitutional memory, made usable
Drawings, calculations, standards, specifications and inspection reports from projects going back decades.
"Which of our projects used this connection detail, and what did the calculations assume?"
ManufacturingThe answer on the shop floor
Manuals, maintenance logs, procedures, incident reports and supplier documentation that nobody can search under time pressure.
"Why does this machine throw this fault, what did we do last time, and what does the manufacturer say?"
Chemicals, labs and pharmaConfidential and safety-critical at once
Safety data sheets, procedures, batch records, formulation and internal research — commercially and physically consequential.
"Pull every procedure and prior record involving this compound at this concentration."
Clinics and healthcareAdministration, not diagnosis
Patient documentation, policies, insurance paperwork and internal guidance. Scope would stay firmly on the administrative side, with clinical boundaries set explicitly.
"Find the current policy on this procedure and the paperwork it requires."
Architecture and constructionProject memory per job
Requirements, permits, drawings, specifications, contractor correspondence and change records for every project in the archive.
"What did we agree with the client about this specification, and when did it change?"
Financial advice and insuranceClient files under obligation
Portfolios, statements, policies, claims and suitability documentation held under regulatory duties.
"Summarise this client's position and everything we've recommended, with the documents behind it."
Defence and aerospace supply chainWhere isolation may be required
Restricted technical documentation, tender material and supplier information, sometimes under contractual terms that rule external services out entirely.
"Which of our submissions covered this requirement, and what did we commit to?"
Public sector and local governmentRecords at scale
Planning applications, tenders, legal correspondence, regulations and council documents, with transparency duties attached.
"Find the precedent decisions relevant to this application and what was cited."
Honest exclusions
Who we'd talk out of it
If you're a café, a gym, a salon, a small retailer, an ecommerce store or a freelance practice, this is almost certainly the wrong purchase. You don't have the archive to justify it, and you'd be paying for infrastructure to solve a problem you don't have.
We'd also push back where:
- the documents are in such poor condition that the honest first project is cleaning them up, not installing anything
- an enterprise agreement with a mainstream AI provider would satisfy the actual obligation — which it often does
- nobody internally is accountable for the archive, in which case the system will rot within a year
- the real requirement is a decision-making system rather than a retrieval one; that's a different, much heavier conversation
Saying this in advance is cheaper for both of us than discovering it in month three.
The commercial shape
Not a box with software on it. A build, then a service — and the second half is the one that decides whether it still works in year two.
- Discovery and archive mapping
- Access rules and scope agreed in writing
- Hardware specified and installed
- Models selected and retrieval built
- Interfaces your team will actually open
- Pilot, measurement and training
- New documents ingested as they arrive
- Model and software updates
- Retrieval quality tuned as questions change
- Monitoring, backups and access reviews
- New workflows as people find uses
- Support with someone accountable
Cost depends on archive size, user count, hardware and how much structuring the material needs first — which is why we scope it against a specific archive rather than publishing a number. The ongoing half is the point: a private AI system nobody maintains degrades quietly. Documents stop being added, retrieval drifts, people go back to searching manually and conclude the technology didn't work.
What we are not claiming
We have not run this as a managed service. There is no case study here because there isn't one yet, and we would rather have a slower conversation than an inflated one.
What we do bring is the surrounding practice: designing how people work alongside AI systems, deciding what a system is allowed to do, and making its output verifiable by the professional who has to sign their name to it. The retrieval layer is well-understood engineering. The part that decides whether anyone still uses it in month six is the part we work on.
Open questions we would like your view on
- Is the constraint you face contractual, regulatory, or a matter of client perception? Each implies a different design.
- Would you want this on your own hardware, or in infrastructure you rent but control?
- Which single question, answered reliably, would justify the project on its own?
- Who inside the organisation would own it after we leave?
If this describes a problem you have, we want to hear about it.
No pitch attached. We are building a picture of who this genuinely fits, and a short conversation is worth more than a form fill. If we conclude it isn't right for you, we will say so — see the exclusions above.
Mention "private AI" and the archive you have in mind — the sector, roughly how far back it goes, and the one question you would want answered from it.
Read next & further reading
Method · This note describes a service we are scoping, not one we have delivered. Nothing here is a case study. Provider terms change; the primary sources above should be checked against your own contract before any decision rests on them.