Builds

PrivacyThink Chats With Your Documents Without a Single Cloud Call

· Builds

Every chat-with-PDF tool ships your files to someone else's servers. I built the opposite in Tauri and Rust — local embeddings, llama.cpp, zero network calls.

PrivacyThink Chats With Your Documents Without a Single Cloud Call

Last updated: October 3, 2026 · 4-minute read

Read the privacy policy of any "chat with your documents" product and find the sentence about where your files go. It is always there: uploads are processed on the vendor's servers, retained for some window, sometimes used for improvement. For a grocery list, fine. For a contract, a medical report, a salary slip or a client agreement — that trade is not yours to make on the tool's behalf. PrivacyThink is the counter-design: a desktop app, built in Tauri and Rust, that runs embedding and generation entirely on your machine. Zero network calls in the retrieval path. It lives at privacythink.netlify.app, and this is how it works.

The architecture is the privacy policy

The design rule that shapes everything: privacy must be structural, not promissory. A privacy policy is a promise; a missing network stack is a fact. PrivacyThink's pipeline is local end to end — documents are parsed on device, chunked, embedded with a local embedding model, indexed in a local vector store, and answered by a GGUF quantized model through llama.cpp. There is no upload step to secure because there is no upload step.

Tauri earns its place here. The Rust shell gives the app a small footprint and a security model that treats the webview as untrusted, which suits an app whose entire pitch is confinement. The result is a binary measured in megabytes, not the 200-megabyte Electron tax — the same reasoning behind the stack this site's desktop tools use, taken to its logical endpoint for a privacy product.

What local actually costs

Honesty requires the other half of the ledger. Local models, at the quantizations that fit consumer hardware, are less capable than frontier cloud models. The sweet spot on a modern laptop is the 7B-to-8B class of models — the sizing math is the same one I laid out in what actually fits in 8 GB. Retrieval quality compensates more than people expect: because the model receives only the relevant chunks, a modest local model with a good index often produces more grounded answers than a giant model skimming the whole document through a context window.

The second cost is hardware variance. A user on an 8 GB integrated-graphics laptop and a user on a 32 GB discrete-GPU desktop have different experiences, and the app has to degrade gracefully rather than promise one number. That is a support burden cloud tools never carry — their hardware bill arrives monthly, invisible to the user.

The retrieval pipeline, step by step

Worth walking through, because every stage is a privacy decision. Parsing happens on device with no network — a document that never parses never leaves. Chunking splits the text at semantic boundaries rather than fixed character counts, which noticeably improves answer grounding: a clause and its context land in the same chunk. Embedding runs a local model to turn chunks into vectors — the same technology class cloud RAG uses, minus the API key. Retrieval scores the question against the index and hands the model only the top-scoring chunks. Generation is llama.cpp on a GGUF quant, producing an answer that can only cite what the retrieval step supplied.

The threat model that falls out of this is easy to state: there is no vendor to trust, no API key to leak, no usage log on someone else's server, and no terms-of-service change that can retroactively alter where your documents live. What remains is the local attack surface — disk encryption, OS user separation, the usual desktop hygiene — which is the same surface every offline app has always had, and orders of magnitude smaller than "a third party retains your uploads".

Who this is actually for

The pattern I kept hearing while building: professionals who refuse cloud AI for real work, not from ideology but from liability. A lawyer's drafts, an auditor's working papers, a doctor's patient summaries, an HR manager's complaints file. For that group, the question is never "is the local model as clever as GPT-6" — it is "can this file leave this machine at all", and the only acceptable answer is no. The same logic drives the tools in my local AI and privacy collection, including for everyday scenarios that do not involve NDAs.

The secondary audience is the one I belong to: developers who want the pipeline. Local embeddings, chunking strategy, vector-store choice and llama.cpp integration are all inspectable here — a working reference, not a black box.

Try it

The site is privacythink.netlify.app, and the documentation repository is github.com/Bilal140202/Privacythinkdoc, MIT-licensed. Start with one sensitive document you would never paste into a web app, and let the workflow argue for itself. If local AI is new territory, the privacy-first rationale and the 8 GB model guide are the on-ramps. The rest of what I ship is in the projects section.

---

Not affiliated with the llama.cpp project or any cloud AI vendor mentioned. Sources: the PrivacyThink site, its documentation repository, and my own build notes.

ansaribilal.com — technology, tested in public.