Ask Claude anything about your Google Drive.
Corpusly builds a private, searchable index of your Drive on your own machine and connects it to Claude, so your assistant can finally search your real documents and quote them with citations.
Corpusly is the private search layer between your Google Drive and your AI assistant, not another chatbot or document-storage service.
When you ask a question, Corpusly shares only the relevant passages with the AI assistant you've configured.
See it in action
Watch Claude answer from your Google Drive
How it looks in Claude Desktop: ask in plain language, and Corpusly searches your Drive on your device and hands back passages cited to your real files.
Illustrative demo: a simulation of Claude Desktop using Corpusly, not a screen recording.
Works with your AI assistant
Where your data lives
What stays local, and what reaches your assistant
Corpusly reads your Drive, does all the indexing on your machine, and shares only the passages you ask about with the assistant you choose. There's no Corpusly server in the middle.
Stays on your device
Corpusly builds and stores your file text, embeddings and search index locally.
Read from Google Drive
Corpusly streams files read-only, on demand, and never copies them to a Corpusly server.
Shared with your assistant
Only the passages relevant to your question go to the AI assistant you choose.
No Corpusly server
There's no vendor backend holding your files, so there's nothing to breach or subpoena.
How it works
From your Drive to a grounded answer
Corpusly sits on your machine between Google Drive and your AI assistant. It reads your files, indexes them and keeps them in sync, so Claude can search your real documents on demand.
Connect your Drive
Sign in with Google and grant read-only access. Corpusly runs on your machine and never asks for more than permission to read.
Index on your device
Corpusly streams your files and embeds them on your own hardware into a private vector index, then keeps that index in sync automatically.
Ask your assistant
Claude searches your Drive through Corpusly and answers with citations to your actual documents.
Who it's for
Built for people with a lot to remember
If your work lives in Google Drive and you'd rather not hand it to another cloud, Corpusly is for you.
Researchers & writers
Ask across years of notes, papers and drafts, and get answers cited to the exact source.
Consultants & operators
Pull the number, clause or decision out of a sprawling client Drive in seconds.
Developers & technical teams
Ground Claude in your specs, runbooks and source without shipping any of it to a cloud.
Privacy-conscious professionals
Search sensitive records while the index and embeddings stay on your own machine.
Features
A search engine for your own documents
What it takes to make your Google Drive answerable: private by default, and current without any work from you.
Local-first by design
Corpusly indexes and embeds your documents on your own machine, using an open-source embedding model. There's no Corpusly cloud, because there's no Corpusly server, so your Drive and index are never uploaded. When you search, only the matching passages go to the AI assistant you've connected.
Search everything, by meaning
Google Docs, Office files, PDFs, images and more, including scanned files read with built-in OCR. Corpusly understands each file by meaning rather than by keyword, so Claude finds the right passage even when the wording is different.
Scales to enormous Drives
Corpusly catalogs thousands of files in seconds, then embeds each document the moment you first search it. You're productive immediately, without waiting for an overnight index to finish.
Always up to date
Incremental sync follows Google Drive's change feed in the background, so new and edited files show up on their own. You never have to re-index by hand.
Grounded, cited answers
Claude answers with real passages pulled straight from your files, so you can open the source and check any response yourself.
Source-available core
The engine behind Corpusly is source-available, so you can read exactly what it does with your files and run it yourself. No black box sits between your Drive and your assistant.
Built on an open standard
Corpusly speaks the Model Context Protocol (MCP), so it works with Claude Desktop, Claude Code, OpenClaw and any other MCP-compatible client.
Source-available
Read the code that reads your Drive.
Anyone can claim their software is private. Corpusly's core is source-available, so you don't have to take our word for it. The engine that connects to your Drive, builds your index and answers every search is open for you to read and run yourself, and we're publishing the source shortly.
That core is free. A full-featured desktop app is on the way for anyone who'd rather skip the command line, and while that app is a separate, closed-source product, the part that actually handles your documents stays the part you can read.
Source publishing soon.
- Source-available and free. Read the core, run it yourself, and see exactly what it does with your files.
- All the indexing and searching happens on your machine, so there's no server doing anything you can't see.
- A one-click desktop app is coming soon, giving the same engine a friendly interface so you never open a terminal.
Already using Google's Drive connector?
It reads your Drive. Corpusly understands it.
Google's connector is a good way to pull up a file you can already name, and if that's your workflow, it's a fine choice. Corpusly is for the rest of your Drive: the formats it won't read, the questions you can't phrase as keywords, and the files you'd rather never left your machine.
| Capability | Google's Drive connector | Corpusly |
|---|---|---|
| Reads Docs, Sheets, Slides, PDFs, Office and images | Yes | Yes |
| Reads text, Markdown, CSV, JSON, notebooks and source code | No | Yes |
| Reads email archives (.eml, .mbox) | No | Yes |
| Reads EPUB, Rich Text and HTML | No | Yes |
| Searches your Drive by meaning | Keyword only | Yes |
| Access it asks for in your Drive | Read and write | Read-only |
| Works with assistants other than Claude | No | Yes |
| Source you can read and verify | No | Yes |
The capability rows are compared against the Google Drive connector's own published tool definitions as of 16 July 2026. “Reads” means extracting a file's text for the assistant to use; the connector can still download other file types as raw data. Google's connector also offers file-creation tools, so it requests write access to your Drive, while Corpusly only ever asks to read. Google runs its connector as a hosted service you cannot inspect; Corpusly's core is source-available, so you can read and run it yourself.
Privacy
Your Drive is never copied to a Corpusly cloud.
Corpusly is local-first. It streams your Drive, computes embeddings on your own hardware, and stores the index in a private database on your machine. There's no Corpusly cloud, because there's no Corpusly server. When you ask a question, Corpusly shares only the relevant passages with the AI assistant you've connected. Nothing else leaves your machine.
That's the difference from a Drive connector, which hands your assistant a whole file to answer from. Corpusly narrows what travels to the passages that actually answer your question. Those passages are still handled by your assistant's provider under its own terms. What changes is how much of your Drive is ever exposed, and that Corpusly only ever asks to read it.
- An open-source embedding model runs locally on your machine, with no calls to a Corpusly server
- There's no cloud component, so there's no server to breach and no vendor backend
- Read-only Drive access means Corpusly can never modify or delete your files
- Your Google sign-in token stays on your device, in your operating system's keychain where available
- Source files are streamed on demand, never copied or mirrored to disk
Privacy architecture
Three ways to give AI your documents
Most document-AI tools copy your files into their cloud, and Drive connectors extract your files server-side before handing them to the assistant. Corpusly keeps the whole pipeline on your machine.
Typical hosted document AI
Google's Drive connector for Claude
Corpusly
Relevant passages are provided to the AI assistant you choose when answering a question.
Sources
Bring your whole knowledge base
Scanned contracts, photographed whiteboards, decade-old .doc files, whole Gmail archives exported to .mbox: Corpusly reads them all, with OCR running on your own device. You don't have to export or convert anything first.
See every supported format
Google Workspace
- Docs
- Sheets
- Slides
- Drawings
Microsoft Office
- .docx / .doc
- .xlsx / .xls
- .pptx / .ppt
- Templates & macro-enabled (.docm, .xlsm, .pptm)
PDF & images
- Scanned PDF (OCR)
- PNG / JPG (OCR)
- SVG (vector text)
OpenDocument
- .odt / .ott
- .ods / .ots
- .odp / .otp
- .odg drawings
Ebooks & rich text
- EPUB
- Rich Text (.rtf)
- HTML
Email & mailboxes
- .eml messages
- .mbox archives
- Attachments indexed
Text & data
- Text
- Markdown
- CSV
- JSON
- XML / YAML
Source code
- Most languages
- Jupyter / Colab notebooks
The founder
A local-first product, from a local-first engineer.
Jeff Mixon is a senior software engineer who designs distributed backends, ships full-stack products end to end, and trains, fine-tunes and serves machine-learning systems. He works at both ends of the stack, from TypeScript microservices down to Linux kernel drivers and software-defined radio.
That's why Corpusly is local-first. Jeff has spent years building end-to-end-encrypted, offline-first software with on-device machine learning, and the same instincts keep your Drive and your search index on your own machine rather than on someone else's server.
FAQ
Questions, answered
Does my data go to the cloud?
Your files, index and embeddings never go to a Corpusly cloud. There is no Corpusly server, and all indexing and embedding happens on your own machine. The one exception: when you ask a question, Corpusly sends the relevant passages to the AI assistant you've connected (such as Claude) so it can answer, and that provider's privacy terms apply to them.
Does Claude receive content from my documents?
When Claude searches your Drive through Corpusly, Corpusly returns only the passages relevant to your question, and those are sent to Claude through your existing Claude connection so it can answer. Corpusly doesn't control how your AI provider processes or retains that content, so its terms and privacy settings still apply.
Which AI assistants are supported?
Claude Desktop, Claude Code and OpenClaw today, over the open Model Context Protocol (MCP). ChatGPT support is coming soon. Because Corpusly is a standard MCP server, any MCP-compatible client can connect to it.
What about Google's Drive connector for Claude?
It's a good fit for a lot of work, and if it covers yours, use it. Google publishes a Drive connector in Claude's directory, and it's genuinely handy for pulling up a document you can already name. Corpusly solves a different problem, and three differences are worth knowing. Formats: the connector's own tool definitions list the file types it can turn into text, namely Google Docs, Sheets and Slides, PDFs, Word, Excel, PowerPoint, OpenDocument and images. Plain text, Markdown, CSV, JSON, source code, notebooks, email archives, EPUB, Rich Text and HTML aren't on that list, and neither are Google Drawings or SVG; Corpusly reads all of them. Search: the connector searches Drive by keyword, so you need to guess the words in the file. Corpusly searches by meaning, so it finds the right passage when the wording is different. Access: Google's connector includes tools for creating and copying files, so it asks to write to your Drive. Corpusly only ever asks to read. That said, Corpusly still sends the passages it retrieves to Claude, where Anthropic's terms apply. What's different is that your index and everything you didn't ask about stay on your machine.
Why not just use Google Drive search?
Drive search is excellent at finding a file when you know roughly what it's called. It's not built to answer a question from what's inside your files. Corpusly reads the contents, understands them semantically, and lets your assistant answer from them with citations you can open. You can ask what a contract says instead of hunting for the contract.
Why not just use Gemini?
Corpusly is a different approach, not a rival to any one assistant. It lets you keep the assistant you already use, indexes your Drive locally instead of relying on a hosted service, reads more file types (including scanned PDFs and images via OCR), gives you transparent citations you can open, and keeps the retrieval layer under your control. If Google's built-in features fit your needs, they're a fine choice too.
What access to my Drive does Corpusly need?
Read-only. You sign in with Google and grant read access; Corpusly can never modify, move or delete anything in your Drive.
How large a Drive can it handle?
Very large. Corpusly catalogs your entire Drive in seconds, then embeds documents the moment you first search them, so you're searchable right away without waiting for a full index to finish.
What file types can Corpusly index?
Most of what real work lives in: Google Docs, Sheets, Slides and Drawings; Microsoft Word, Excel and PowerPoint, including the legacy .doc, .xls and .ppt formats and template and macro-enabled files like .docm, .xlsm and .pptm; OpenDocument (.odt, .ods, .odp, .odg) and its templates; PDFs; images; SVG; EPUB; Rich Text (.rtf); HTML; email messages and mailbox archives (.eml and .mbox); Jupyter and Colab notebooks; and plain text, Markdown, CSV, XML, YAML, source code and JSON. Corpusly reads scanned PDFs and images with built-in OCR, and skips anything it can't extract text from rather than uploading it.
Can Corpusly read scanned PDFs and images?
Yes, both. Corpusly runs optical character recognition (OCR) on scanned, image-only PDFs and on image files such as screenshots, photographed receipts or contracts, and whiteboard snapshots, so the text locked inside them is indexed and searchable just like a native document. The OCR runs entirely on your own device, so your images and scans are never uploaded.
Can Corpusly search my email?
If your email lives in your Drive as a file, yes. Corpusly indexes both single messages (.eml) and entire mailbox archives (.mbox), including the Gmail archives Google Takeout produces. It makes the Subject, From, To and Date headers and the message body searchable, and it extracts attachments inline, so a PDF or contract attached to an email is indexed too. Corpusly reads these files from your Drive on your own machine; it doesn't connect to your live inbox or mail server.
What is MCP?
The Model Context Protocol is an open standard that lets AI assistants connect to external tools and data. Corpusly exposes your Drive to Claude as an MCP server, so your assistant can search it directly.
Is Corpusly open source?
The engine at the core of Corpusly, the part that indexes your Drive and runs your searches, is source-available. Anyone can read the code, run it, and check what it does; we're publishing it shortly. It's under a fair-source license rather than a classic open-source one, but none of what it does with your files is hidden. The embedding model it uses is open source and runs on your own machine. The desktop app that wraps the core in a one-click interface is a separate, closed-source product.
What will Corpusly cost?
The core engine is free and source-available, yours to run at no cost. The full-featured desktop app is still in early access. Request an invite and we'll share its pricing as we roll out.
Do I need to be technical to use it?
Right now the core runs from the command line, so if you're comfortable in a terminal you can set it up today. For everyone else, a full-featured desktop app is on the way. It does the setup for you, so you never have to touch a command line. Request early access to get it as soon as it's ready.
Early access
Make your Drive searchable from Claude.
Corpusly is launching soon. Request early access and get launch updates.
No spam. Just a launch invite. Unsubscribe anytime.
