BoltProof · How it's set up
One headless computer in your office runs the AI. Staff connect from their own laptops and phones. A NAS keeps a copy of everything. Here is the full picture, then a live demo.
A BoltProof system is one small computer in your office, running without a screen, that every member of staff can reach from their own laptop or phone. The models, your documents and your chat history stay on that computer and its NAS backup — the only network traffic is encrypted Tailscale access for your authorized devices. This page shows how the parts fit together, what installation involves, and what you will see on the first day.
Solid boxes are things you own. The dashed box on the right is optional. Staff devices can be anywhere; the server and NAS stay in your office.
Staff use whatever they already have: Windows, Mac, iPhone, Android. They install one app (Tailscale) and open a web address. No software licences per seat, no new hardware for staff.
A private network between your devices and the server. Traffic is encrypted end to end. There's no port forwarding and no public-facing IP — the server isn't reachable from the open internet, only from devices your Tailscale network has authorized. You can see who is connected and remove a device in one click.
A computer with no monitor or keyboard. It runs the language models (Ollama), the chat interface (Open WebUI), the automation engine (n8n) and the document index. the OS is hardened: disk encryption, firewall, automatic security patches, no sharing services.
Two mirrored drives, so one can fail without losing data. Every night it takes a full copy of the server. Snapshots let you roll back to yesterday, last week or last month. Priced in your proposal.
An encrypted copy of the NAS sent to a location you choose: a second office, a home NAS, or a storage provider in the UAE. Protects against fire, flood and theft. You hold the encryption key.
Quoted per request, no retainer. Ask us to apply updates, swap in newer models or fix a problem remotely, whenever you need it.
A chat window that looks like the tools your staff already know. Drafting letters, summarising a long email thread, rewriting a paragraph, translating to Arabic. Every conversation is stored on your server and nowhere else.
Ask a question across your own files: contracts, policies, patient leaflets, price lists, tenancy agreements. The answer quotes the source document and page so you can check it. New files added to the shared folder are indexed automatically.
An enquiry lands in your shared inbox. n8n reads it, the local model drafts a reply using your price list and standard wording, and the draft appears in a folder or a chat for approval. A person clicks send. Nothing goes out without a human.
Other automations we set up on request: meeting notes to a summary and action list, invoices to a spreadsheet line, a daily digest of new documents, a WhatsApp enquiry to a CRM record.
Hardware fails, offices flood, laptops walk. With the backup add-on, the NAS holds last night's full copy of the server and a chain of snapshots. We put replacement hardware in place, restore from the NAS, reconnect Tailscale, and your staff carry on with the same address, the same chat history and the same documents. Typical downtime is measured in hours. Without the NAS, expect one to three working days to rebuild.
Already live: try the assistant yourself on synthetic documents at demo.boltproof.com — no signup, 10 free messages an hour.
We share our screen, you send us two or three of your own non-confidential documents in advance, and you watch the system answer questions about them — then generate a draft. No slides. You can ask anything.
Try the live sandbox on synthetic documents, no signup required, or compare the three packages / use the configurator to see which one fits your staff count and document volume.
Try the live sandbox demo See packages Open the configurator
The plain-English explanations above are the whole story for most people. If you want to validate the architecture yourself, here is the detail.
Unified memory (RAM and GPU memory in one pool) is what makes running large open-weight models on a desktop machine practical — a 64 to 512 GB unified-memory machine can hold and run models that would otherwise need a rack of GPUs. The machine's GPU and neural-engine cores accelerate inference. We size the hardware to the model sizes and concurrent-user count you need — see the how we work.
Models run through Ollama, an open-source runtime that loads open-weight models (Qwen, Llama, Gemma, DeepSeek and others) directly on the server. Open WebUI provides the chat interface on top. No model weights or inference requests leave the machine for local models — inference happens entirely on that hardware. We choose specific models based on your memory budget and task mix, and swap them for better releases when you ask us to.
This is retrieval-augmented generation (RAG): your documents are split into chunks, each chunk is converted to a vector embedding using a local embedding model (nomic-embed), and those vectors are stored in a local index. When you ask a question, the system finds the most relevant chunks by vector similarity, feeds them to the language model as context, and the model answers using that context — citing which file and page the answer came from. New files dropped into a watched folder are indexed automatically. Nothing about this process sends your documents anywhere; the index and the embeddings live on the same machine as the model.
Open WebUI gives each staff member their own login. Access to specific document folders can be restricted per user or group — for example, so reception can't see HR files. Every login is tied to a Tailscale-authenticated device, so there's no shared password floating around. We set this up at delivery based on who should see what; ask us if you need more granular controls than the default.
Tailscale creates a private mesh network (built on WireGuard) between your staff's devices and the server. There is no port forwarding and no public-facing IP or exposed service — the server is simply unreachable from the open internet. Each device authenticates individually, traffic is encrypted end-to-end, and you can revoke a single device's access instantly from the Tailscale admin console without touching anything else.
The optional NAS add-on takes a nightly backup of the server (documents, chat history, model configuration and automations) onto two mirrored drives, with versioned snapshots so a file overwritten today can be recovered from yesterday's snapshot. An optional encrypted off-site copy protects against fire, flood or theft; you hold the encryption key. Without the NAS add-on, you're relying on whatever backup discipline you already have — we'll say so plainly if that's a gap worth closing.
Every deployment gets OS hardening (full-disk encryption, firewall enabled, automatic security updates, no unnecessary sharing services, a separate admin account) as standard. The fuller checklist — 2FA enforcement, access reviews, ransomware-resistant backup design, incident response planning — is documented on the Security & hardening page.
If you ask us for remote support, we connect over the same Tailscale network your staff use — there's no separate remote-access tool or exposed port added for us. We can only reach the server while your Tailscale network authorizes our device, and you can revoke that at any time. Updates are staged with a NAS snapshot taken first, so any update can be rolled back if something breaks.
We pick models based on three things: how much unified memory the server has, how many people will be using it at once, and what the work actually is (drafting and Q&A need different strengths than translation or code review). Smaller machines run models in the 8–30B parameter range; larger ones run 70B to 400B-class models, usually quantised to fit memory while keeping quality close to full precision. We test the first real answer against your own documents before handover, not a generic benchmark.
n8n handles automations and connects to external systems through whatever export or API they expose — reading an inbox via IMAP, watching a folder for new invoices, calling a practice-management system's API if it has one. If a system only exports CSV or PDF, we build the workflow around that instead. Tell us what you already run (accounting software, CRM, practice management) and we'll say plainly whether it connects before you order.
By default, nothing is configured to call an external AI API. If you choose to add a cloud model (GPT, Claude, Gemini) for non-sensitive work — general drafting, marketing copy, research with no client data — that's a separate, clearly labelled option in Open WebUI that you switch on and pay for yourself with your own API key. Staff can see which model, local or cloud, they're using for every conversation. The local vs frontier vs hybrid comparison covers how to decide what should go where.
Prices are in AED. Hardware is new unless marked certified pre-owned. Model names and sizes change as better open-weight releases appear; we tell you what is installed at handover.