On-premises AI
Self-hosted AI for manufacturers
A shop's real assets are its drawings, its tooling knowledge, and what it knows about pricing a job. None of that belongs in a chatbot — which is exactly why it should run on your own hardware instead.
Manufacturing is one of the few industries where the case for AI and the case against it are made with the same sentence: everything we do is in those files.
The files are why AI would help. An RFQ packet is forty pages of PDF that somebody has to read carefully at 7am. The answer to "have we made this before?" is buried in eight years of job history that only one estimator can navigate. Setup knowledge lives in a machinist's head and walks out of the building when he retires.
The files are also why shops stop. Customer prints arrive under NDA. Defence work carries handling terms written long before anyone imagined pasting a drawing into a web form. And the pricing history — the single most valuable dataset in the building — is the last thing you'd hand to a service you don't control.
Both things are true at once. The resolution isn't to skip AI. It's to keep the model inside the fence.
What "inside the fence" means in practice
Open-weight models — published AI models you can download and run yourself — install on a computer you own. In a shop that's typically one workstation-class machine sitting alongside your existing server, connected to your network and nothing else.
Your ERP, your job folders, and your document scanner talk to it over the LAN. Prints are read on that machine. Answers come back on that machine. There is no outbound connection to disable, because there was never one to begin with — and if you want it physically air-gapped after installation, that's a supported configuration rather than an exotic request.
The compliance question changes shape entirely when nothing is transmitted. You are no longer arguing about a vendor's terms of service. You are pointing at a box on your own network.
The five jobs worth doing first
Not everything in a shop benefits from AI, and the demos that circulate — generating toolpaths, predicting machine failure — tend to be the least practical starting points for a business under fifty people. These five are where the return actually is.
1. Reading RFQ packets
An incoming quote request is a mixed bag: a PDF drawing, a spreadsheet of quantities, an email with delivery terms, sometimes a scanned markup. Somebody reads all of it and re-types the essentials into your system.
A local model reads the packet and pulls out part numbers, quantities, materials, tolerances, finish callouts, and required dates into a structured record — flagging anything ambiguous rather than guessing. The estimator reviews a filled-in form instead of building one from scratch. On a shop quoting several packets a week, that's the difference between quoting same-day and quoting Thursday.
2. "Have we made this before?"
Most shops have years of history that is technically searchable and practically not. You can find a job number if you know it; you cannot ask "what similar brackets have we run in 304, and what did they actually cost us?"
Pointing a local model at your own job history, quotes, and travelers makes that question answerable in plain language — with links back to the source records, so the estimator verifies rather than trusts. This is the single most requested capability we hear from shops, and the one that most obviously must stay in-house.
3. Drafting the quote
Once the packet is read and the history is retrieved, a first-draft quote can be assembled from your own pricing logic and past jobs. Not sent — drafted. The estimator adjusts and approves.
The gain is not accuracy, it's latency. Quotes that go out the same day win work that quotes going out next week do not, and every shop owner already knows this.
4. Capturing what's in people's heads
Setup sheets, process notes, fixture photos, the annotated printout taped inside a cabinet door. A local model can make that body of material askable — "how did we fixture this last time?" — which matters most in shops where one or two people hold the knowledge and everyone else waits for them.
This is also succession insurance. It is not a substitute for training people, and we'd be suspicious of anyone selling it that way.
5. Shop-floor paperwork
Delivery notes, certs, inspection records, receiving documents. Photographed or scanned, read, and filed as structured data instead of a PDF nobody will ever find again. Unglamorous, and usually the quickest thing to get running.
What we'd steer you away from at the start. Generating machining strategy, predictive maintenance from sensor data, and automated scheduling all sound compelling and all require either data you don't have cleanly or a tolerance for error you shouldn't accept on the floor. Start where the material is text and paperwork, where a human reviews the output, and where being wrong costs a minute rather than a spindle.
On ITAR, CMMC, and customer NDAs
We're going to be careful here, because this is the area where AI vendors make claims they can't support.
ITAR, CMMC, and similar frameworks govern where controlled material may be stored and processed, and who may access it. Whether any particular cloud AI service satisfies your obligations depends on that service, your contract, your enclave, and your own compliance posture. That's a question for your compliance advisor — not something a consultancy should answer with a marketing page, and not something we'll claim to certify.
What we can say plainly is structural: running the model on hardware you control removes transmission from the analysis. If the material never leaves your network, the question "is this vendor an approved processor?" doesn't need answering, because there is no vendor in the path. Many shops under these obligations choose on-premises for exactly that reason — it makes the conversation shorter.
The same logic applies, less formally, to ordinary customer NDAs. Most were written to prohibit disclosure to third parties. A reasonable reading is that uploading a print to an outside service is disclosure to a third party. Keeping it in-house sidesteps the argument.
What it takes to stand up
For a shop of ten to fifty people, this is a smaller project than it sounds:
- Hardware — usually one workstation-class machine with a modern GPU. Not a rack, not a data centre, and it can live wherever your existing server does.
- Installation — models installed locally, chosen for the work you actually do rather than benchmark scores.
- Connection — links to your ERP, job folders, and scanner over your own network, with access limited to the people who need it.
- Handover — documentation and training pitched at whoever handles your IT, whether that's a person on staff or an outside firm on call.
Timeline is typically two to four weeks from scoping to a working system. The full cost picture, including the expenses people don't anticipate, is in what on-premises AI actually costs.
The split most shops end up with
Very few shops need everything in-house, and pretending otherwise costs money. The usual outcome is a boundary: general admin, scheduling, and marketing run on ordinary cloud services where they're cheap and easy; drawings, job history, and pricing never leave the building.
Deciding where that line goes is most of the work — and it's worth doing before anyone buys hardware.
Related reading: AI without sending your data to the cloud is the general primer on how on-premises AI works. What on-premises AI actually costs covers hardware sizing and the hidden expenses.
M Kiln AI builds AI systems for small and medium businesses in Rochester, New York — a region that has made precision parts for a century — and remotely across the United States. We make no compliance certifications; we build systems that keep your material on your own hardware, and we'll tell you when you don't need us.