Private RAG for Google Drive · Early access

Ask Claude anything about your Google Drive.

Corpusly builds a private, searchable index of your Drive on your own machine and connects it to Claude, so your assistant can finally search your real documents and quote them with citations.

Corpusly is the private search layer between your Google Drive and your AI assistant, not another chatbot or document-storage service.

When you ask a question, Corpusly shares only the relevant passages with the AI assistant you've configured.

Local indexSource-availableNo Corpusly cloudRead-only accessExcerpts only
Claude
Connected to Google Drive

What did we decide about pricing in the Q3 planning doc?

search_drivesearched 1,240 files · local search

Three tiers — Free, Pro at $12/mo and Team at $20/seat — with annual billing 20% off. Team was gated on shipping SSO first.

Q3 Planning — Pricing.gdoccited

What's our liability cap in the Acme contract?

search_drivesearched 38 contracts · local search

Capped at 12 months of fees paid — with a carve-out making data-breach and IP claims uncapped.

MSA — Acme Corp.pdf · p. 8cited

What churn number did we show the board last quarter?

search_drivesearched 1,240 files · local search

Net revenue retention was 112% and logo churn was 4.1% for Q2 — up from 3.6% the quarter before.

Board Deck — Q2.pptx · slide 14cited
Reply to Claude…

See it in action

Watch Claude answer from your Google Drive

How it looks in Claude Desktop: ask in plain language, and Corpusly searches your Drive on your device and hands back passages cited to your real files.

Illustrative demo: a simulation of Claude Desktop using Corpusly, not a screen recording.

Works with your AI assistant

Claude DesktopClaude CodeSoonChatGPTOpenClawAny MCP client

Where your data lives

What stays local, and what reaches your assistant

Corpusly reads your Drive, does all the indexing on your machine, and shares only the passages you ask about with the assistant you choose. There's no Corpusly server in the middle.

read-onlyrelevant passagesGoogle Driveyour filesCorpuslyon your deviceYour AI assistantasks & cites

Stays on your device

Corpusly builds and stores your file text, embeddings and search index locally.

Read from Google Drive

Corpusly streams files read-only, on demand, and never copies them to a Corpusly server.

Shared with your assistant

Only the passages relevant to your question go to the AI assistant you choose.

No Corpusly server

There's no vendor backend holding your files, so there's nothing to breach or subpoena.

How it works

From your Drive to a grounded answer

Corpusly sits on your machine between Google Drive and your AI assistant. It reads your files, indexes them and keeps them in sync, so Claude can search your real documents on demand.

1

Connect your Drive

Sign in with Google and grant read-only access. Corpusly runs on your machine and never asks for more than permission to read.

2

Index on your device

Corpusly streams your files and embeds them on your own hardware into a private vector index, then keeps that index in sync automatically.

3

Ask your assistant

Claude searches your Drive through Corpusly and answers with citations to your actual documents.

Who it's for

Built for people with a lot to remember

If your work lives in Google Drive and you'd rather not hand it to another cloud, Corpusly is for you.

Researchers & writers

Ask across years of notes, papers and drafts, and get answers cited to the exact source.

Consultants & operators

Pull the number, clause or decision out of a sprawling client Drive in seconds.

Developers & technical teams

Ground Claude in your specs, runbooks and source without shipping any of it to a cloud.

Privacy-conscious professionals

Search sensitive records while the index and embeddings stay on your own machine.

Features

A search engine for your own documents

What it takes to make your Google Drive answerable: private by default, and current without any work from you.

Local-first by design

Corpusly indexes and embeds your documents on your own machine, using an open-source embedding model. There's no Corpusly cloud, because there's no Corpusly server, so your Drive and index are never uploaded. When you search, only the matching passages go to the AI assistant you've connected.

Search everything, by meaning

Google Docs, Office files, PDFs, images and more, including scanned files read with built-in OCR. Corpusly understands each file by meaning rather than by keyword, so Claude finds the right passage even when the wording is different.

Scales to enormous Drives

Corpusly catalogs thousands of files in seconds, then embeds each document the moment you first search it. You're productive immediately, without waiting for an overnight index to finish.

Always up to date

Incremental sync follows Google Drive's change feed in the background, so new and edited files show up on their own. You never have to re-index by hand.

Grounded, cited answers

Claude answers with real passages pulled straight from your files, so you can open the source and check any response yourself.

Source-available core

The engine behind Corpusly is source-available, so you can read exactly what it does with your files and run it yourself. No black box sits between your Drive and your assistant.

Built on an open standard

Corpusly speaks the Model Context Protocol (MCP), so it works with Claude Desktop, Claude Code, OpenClaw and any other MCP-compatible client.

Source-available

Read the code that reads your Drive.

Anyone can claim their software is private. Corpusly's core is source-available, so you don't have to take our word for it. The engine that connects to your Drive, builds your index and answers every search is open for you to read and run yourself, and we're publishing the source shortly.

That core is free. A full-featured desktop app is on the way for anyone who'd rather skip the command line, and while that app is a separate, closed-source product, the part that actually handles your documents stays the part you can read.

Source publishing soon.

  • Source-available and free. Read the core, run it yourself, and see exactly what it does with your files.
  • All the indexing and searching happens on your machine, so there's no server doing anything you can't see.
  • A one-click desktop app is coming soon, giving the same engine a friendly interface so you never open a terminal.

Already using Google's Drive connector?

It reads your Drive. Corpusly understands it.

Google's connector is a good way to pull up a file you can already name, and if that's your workflow, it's a fine choice. Corpusly is for the rest of your Drive: the formats it won't read, the questions you can't phrase as keywords, and the files you'd rather never left your machine.

CapabilityGoogle's Drive connectorCorpusly
Reads Docs, Sheets, Slides, PDFs, Office and imagesYesYes
Reads text, Markdown, CSV, JSON, notebooks and source codeNoYes
Reads email archives (.eml, .mbox)NoYes
Reads EPUB, Rich Text and HTMLNoYes
Searches your Drive by meaningKeyword onlyYes
Access it asks for in your DriveRead and writeRead-only
Works with assistants other than ClaudeNoYes
Source you can read and verifyNoYes

The capability rows are compared against the Google Drive connector's own published tool definitions as of 16 July 2026. “Reads” means extracting a file's text for the assistant to use; the connector can still download other file types as raw data. Google's connector also offers file-creation tools, so it requests write access to your Drive, while Corpusly only ever asks to read. Google runs its connector as a hosted service you cannot inspect; Corpusly's core is source-available, so you can read and run it yourself.

See the full file-type-by-file-type comparison →

Privacy

Your Drive is never copied to a Corpusly cloud.

Corpusly is local-first. It streams your Drive, computes embeddings on your own hardware, and stores the index in a private database on your machine. There's no Corpusly cloud, because there's no Corpusly server. When you ask a question, Corpusly shares only the relevant passages with the AI assistant you've connected. Nothing else leaves your machine.

That's the difference from a Drive connector, which hands your assistant a whole file to answer from. Corpusly narrows what travels to the passages that actually answer your question. Those passages are still handled by your assistant's provider under its own terms. What changes is how much of your Drive is ever exposed, and that Corpusly only ever asks to read it.

Security & data-flow details

  • An open-source embedding model runs locally on your machine, with no calls to a Corpusly server
  • There's no cloud component, so there's no server to breach and no vendor backend
  • Read-only Drive access means Corpusly can never modify or delete your files
  • Your Google sign-in token stays on your device, in your operating system's keychain where available
  • Source files are streamed on demand, never copied or mirrored to disk

Privacy architecture

Three ways to give AI your documents

Most document-AI tools copy your files into their cloud, and Drive connectors extract your files server-side before handing them to the assistant. Corpusly keeps the whole pipeline on your machine.

Typical hosted document AI

Upload files
Cloud parsing
Cloud embeddings
Hosted vector database

Google's Drive connector for Claude

Read from Drive
Extracted on Google's servers
Whole file text to Claude

Corpusly

Read from Drive
Local parsing
Local embeddings
Local index

Relevant passages are provided to the AI assistant you choose when answering a question.

Sources

Bring your whole knowledge base

Scanned contracts, photographed whiteboards, decade-old .doc files, whole Gmail archives exported to .mbox: Corpusly reads them all, with OCR running on your own device. You don't have to export or convert anything first.

Google Workspace
Microsoft Office
PDF & images
OpenDocument
Ebooks & rich text
Email & mailboxes
Text & data
Source code
See every supported format

Google Workspace

  • Docs
  • Sheets
  • Slides
  • Drawings

Microsoft Office

  • .docx / .doc
  • .xlsx / .xls
  • .pptx / .ppt
  • Templates & macro-enabled (.docm, .xlsm, .pptm)

PDF & images

  • PDF
  • Scanned PDF (OCR)
  • PNG / JPG (OCR)
  • SVG (vector text)

OpenDocument

  • .odt / .ott
  • .ods / .ots
  • .odp / .otp
  • .odg drawings

Ebooks & rich text

  • EPUB
  • Rich Text (.rtf)
  • HTML

Email & mailboxes

  • .eml messages
  • .mbox archives
  • Attachments indexed

Text & data

  • Text
  • Markdown
  • CSV
  • JSON
  • XML / YAML

Source code

  • Most languages
  • Jupyter / Colab notebooks
Jeff Mixon

Jeff Mixon

Founder of Corpusly

Engineer · entrepreneur · explorer

The founder

A local-first product, from a local-first engineer.

Jeff Mixon is a senior software engineer who designs distributed backends, ships full-stack products end to end, and trains, fine-tunes and serves machine-learning systems. He works at both ends of the stack, from TypeScript microservices down to Linux kernel drivers and software-defined radio.

That's why Corpusly is local-first. Jeff has spent years building end-to-end-encrypted, offline-first software with on-device machine learning, and the same instincts keep your Drive and your search index on your own machine rather than on someone else's server.

FAQ

Questions, answered

Does my data go to the cloud?

Your files, index and embeddings never go to a Corpusly cloud. There is no Corpusly server, and all indexing and embedding happens on your own machine. The one exception: when you ask a question, Corpusly sends the relevant passages to the AI assistant you've connected (such as Claude) so it can answer, and that provider's privacy terms apply to them.

Does Claude receive content from my documents?

When Claude searches your Drive through Corpusly, Corpusly returns only the passages relevant to your question, and those are sent to Claude through your existing Claude connection so it can answer. Corpusly doesn't control how your AI provider processes or retains that content, so its terms and privacy settings still apply.

Which AI assistants are supported?

Claude Desktop, Claude Code and OpenClaw today, over the open Model Context Protocol (MCP). ChatGPT support is coming soon. Because Corpusly is a standard MCP server, any MCP-compatible client can connect to it.

What about Google's Drive connector for Claude?

It's a good fit for a lot of work, and if it covers yours, use it. Google publishes a Drive connector in Claude's directory, and it's genuinely handy for pulling up a document you can already name. Corpusly solves a different problem, and three differences are worth knowing. Formats: the connector's own tool definitions list the file types it can turn into text, namely Google Docs, Sheets and Slides, PDFs, Word, Excel, PowerPoint, OpenDocument and images. Plain text, Markdown, CSV, JSON, source code, notebooks, email archives, EPUB, Rich Text and HTML aren't on that list, and neither are Google Drawings or SVG; Corpusly reads all of them. Search: the connector searches Drive by keyword, so you need to guess the words in the file. Corpusly searches by meaning, so it finds the right passage when the wording is different. Access: Google's connector includes tools for creating and copying files, so it asks to write to your Drive. Corpusly only ever asks to read. That said, Corpusly still sends the passages it retrieves to Claude, where Anthropic's terms apply. What's different is that your index and everything you didn't ask about stay on your machine.

Why not just use Google Drive search?

Drive search is excellent at finding a file when you know roughly what it's called. It's not built to answer a question from what's inside your files. Corpusly reads the contents, understands them semantically, and lets your assistant answer from them with citations you can open. You can ask what a contract says instead of hunting for the contract.

Why not just use Gemini?

Corpusly is a different approach, not a rival to any one assistant. It lets you keep the assistant you already use, indexes your Drive locally instead of relying on a hosted service, reads more file types (including scanned PDFs and images via OCR), gives you transparent citations you can open, and keeps the retrieval layer under your control. If Google's built-in features fit your needs, they're a fine choice too.

What access to my Drive does Corpusly need?

Read-only. You sign in with Google and grant read access; Corpusly can never modify, move or delete anything in your Drive.

How large a Drive can it handle?

Very large. Corpusly catalogs your entire Drive in seconds, then embeds documents the moment you first search them, so you're searchable right away without waiting for a full index to finish.

What file types can Corpusly index?

Most of what real work lives in: Google Docs, Sheets, Slides and Drawings; Microsoft Word, Excel and PowerPoint, including the legacy .doc, .xls and .ppt formats and template and macro-enabled files like .docm, .xlsm and .pptm; OpenDocument (.odt, .ods, .odp, .odg) and its templates; PDFs; images; SVG; EPUB; Rich Text (.rtf); HTML; email messages and mailbox archives (.eml and .mbox); Jupyter and Colab notebooks; and plain text, Markdown, CSV, XML, YAML, source code and JSON. Corpusly reads scanned PDFs and images with built-in OCR, and skips anything it can't extract text from rather than uploading it.

Can Corpusly read scanned PDFs and images?

Yes, both. Corpusly runs optical character recognition (OCR) on scanned, image-only PDFs and on image files such as screenshots, photographed receipts or contracts, and whiteboard snapshots, so the text locked inside them is indexed and searchable just like a native document. The OCR runs entirely on your own device, so your images and scans are never uploaded.

Can Corpusly search my email?

If your email lives in your Drive as a file, yes. Corpusly indexes both single messages (.eml) and entire mailbox archives (.mbox), including the Gmail archives Google Takeout produces. It makes the Subject, From, To and Date headers and the message body searchable, and it extracts attachments inline, so a PDF or contract attached to an email is indexed too. Corpusly reads these files from your Drive on your own machine; it doesn't connect to your live inbox or mail server.

What is MCP?

The Model Context Protocol is an open standard that lets AI assistants connect to external tools and data. Corpusly exposes your Drive to Claude as an MCP server, so your assistant can search it directly.

Is Corpusly open source?

The engine at the core of Corpusly, the part that indexes your Drive and runs your searches, is source-available. Anyone can read the code, run it, and check what it does; we're publishing it shortly. It's under a fair-source license rather than a classic open-source one, but none of what it does with your files is hidden. The embedding model it uses is open source and runs on your own machine. The desktop app that wraps the core in a one-click interface is a separate, closed-source product.

What will Corpusly cost?

The core engine is free and source-available, yours to run at no cost. The full-featured desktop app is still in early access. Request an invite and we'll share its pricing as we roll out.

Do I need to be technical to use it?

Right now the core runs from the command line, so if you're comfortable in a terminal you can set it up today. For everyone else, a full-featured desktop app is on the way. It does the setup for you, so you never have to touch a command line. Request early access to get it as soon as it's ready.

Early access

Make your Drive searchable from Claude.

Corpusly is launching soon. Request early access and get launch updates.

No spam. Just a launch invite. Unsubscribe anytime.

Request early access