Privacy Policy

Corpus Retriever (Chrome extension)
Effective date: 30 July 2026
Contact: github.com/corpus-hub/corpus-studio/issues

Corpus Retriever does not collect your data. There is no analytics, no tracking, no account, and no server operated by the developer. Nothing you do in this extension is reported to us, because there is nowhere for it to be reported to.

What the extension stores

One value: a randomly generated identifier, created the first time the extension needs one. You are never asked for it, and it is not derived from anything about you or your browser.

It is stored using Chrome’s own extension storage, which means it lives in your browser profile. It is never transmitted to the developer.

It exists because two open-access services — Unpaywall and the PubMed Central ID converter — require a contact address with each API request and reject requests that omit one. They use it for rate limiting and to reach whoever is generating unusual load. The extension satisfies that requirement without asking you for a real address: it sends something of the form word.word12@corpus-hub.github.io, where the first part is random and the domain points at this project.

It is deliberately not a fingerprint of your browser. A fingerprint is built from a small enough set of values to be worked backwards, and it would identify you across unrelated websites. A random identifier tells those two services which installation is asking and nothing whatsoever about the person using it.

It stays the same between requests on purpose. A fresh one each time would be indistinguishable from evading a rate limit, and would leave those services unable to slow down one installation without blocking everybody. It is sent only to those two services, and only as part of a request you started.

Removing the extension removes it.

What the extension transmits, and to whom

When you paste an identifier and click Download, the extension sends that identifier to academic sources in order to locate the paper:

These requests contain the identifier you entered. Requests to Unpaywall and PubMed Central also contain the random identifier described above. Each of these services has its own privacy policy, and their handling of the request is governed by it rather than by this one.

Requests to publisher websites are made from your browser using your existing session with that publisher, exactly as if you had opened the page yourself. This is how the extension retrieves papers your institution or account already entitles you to read. The extension does not log in for you, does not store credentials, and cannot obtain a paper you could not obtain by hand.

What the extension does not do

Downloaded files

Retrieved PDFs are handed to Chrome’s own download manager and saved to your Downloads folder. The extension does not keep a copy, does not index them, and does not report what you downloaded to anyone.

The optional desktop companion

The extension can optionally be driven by Corpus Studio, a desktop research application by the same developer, over Chrome’s native messaging channel. This is a connection to a program on your own computer, not to the internet. If that application is not installed, the channel is never used. Papers requested this way are saved to your Downloads folder in the same way.

Children

This extension is a research tool and is not directed at children.

Changes to this policy

If the extension’s data handling ever changes, this policy will be updated and its effective date changed before the change takes effect, and the change will be disclosed to users.

Contact

Questions about this policy, or requests relating to data, can be raised at github.com/corpus-hub/corpus-studio/issues.