Skip to content

Latest commit

 

History

33 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

zot

A CLI for querying and maintaining your local Zotero library. It offers hybrid semantic search (BM25 keyword plus vector embeddings, with an optional reranker) and write commands to add papers by DOI/arXiv/ISBN/PubMed ID or PDF, edit metadata, attach files, and trash items.

zot talks to the Zotero local HTTP API (the desktop app's built-in server at http://localhost:23119), so the library never leaves the machine and no Zotero web API key is required for reading. For semantic search it builds a local index (embeddings plus a Tantivy full-text index) on disk.

Requirements

  • Rust toolchain (cargo), installed via rustup.
  • Zotero desktop running, with the local API enabled under Settings → Advanced → "Allow other applications on this computer to communicate with Zotero". The app must be open when you run zot.
  • First use of --rerank downloads a BGE reranker model of about 1 GB (cached by fastembed). The default embedding model is small and downloads automatically.

Install

From a clone of this repo:

git clone <repo-url> zot
cd zot
cargo install --path .

This builds an optimized release binary and places it on your PATH at ~/.cargo/bin/zot, so make sure ~/.cargo/bin is on your PATH.

Update an existing install

Pull the latest changes and reinstall with --force, which is required because cargo otherwise refuses to overwrite an already installed package:

cd zot
git pull
cargo install --path . --force

After updating it is worth refreshing the local index with zot index so it picks up any indexing fixes. If you suspect stale data, do a full rebuild with zot index --force.

Build without installing

cargo build --release       # binary at target/release/zot
cargo run --release -- <args>

Quick start

# 1. Build the local semantic index (incremental; re-run anytime to sync)
zot index

# 2. Semantic search
zot search "diffusion models for point clouds"

# 3. Live keyword search straight from Zotero (no index needed)
zot find "kalman filter" --everything

# 4. Add a paper by DOI, arXiv ID, ISBN or PubMed ID (also refreshes the index)
zot add 10.1145/361598.361623 --tag to-read
zot add arXiv:2401.12345 --collection "Large Language Models"
zot add --pdf ~/Downloads/paper.pdf     # metadata recognized from the PDF

Commands

Command Description
zot index Build or update the local search index (incremental). --force for a full rebuild, --status for index stats and an extraction-status breakdown.
zot index issues List every item whose fulltext extraction has a problem (key, title, status, reason).
zot search <query> Hybrid semantic search (BM25 plus vector) over the local index. Warns if the index is out of sync with Zotero; --no-sync-check skips the check.
zot find <query> Live keyword search via the Zotero local API, always in sync and requiring no index.
zot get <key> Full metadata for an item.
zot fulltext <key> Stored fulltext for an item, from the local index. --start, --end, and --max-chars return a slice.
zot pdf <key> Local file path of an item's PDF attachment.
zot tags List tags in the library; --contains filters.
zot authors List authors and creators in the library; --contains filters.
zot collections [ref] List the collection tree with direct and subtree item counts. --flat drops the indentation, --tree-ids shows connector IDs, and a key, exact name, or tree-view ID limits the listing to one subtree. --create NAME creates a collection instead, under --parent <ref> or at the top level (web API plus sync). --rm <ref> permanently deletes one and every subcollection under it, after a typed confirmation (--force skips it); the items in it are unfiled, not deleted.
zot unfiled List top-level items that are in no collection; --count prints the bare number of them.
zot export <key>... Export items as a bibliography rendered by Zotero. --collection <ref> exports a whole collection instead, --format bibtex|ris|csljson picks the format, --output PATH writes a file, --raw keeps the private BibTeX fields.
zot add [id] [--pdf f] Add a paper by DOI/arXiv/ISBN/PubMed identifier and/or PDF, locally via Zotero's connector API.
zot edit <key> Update item metadata, tags, and collection membership (web API plus sync).
zot attach <key> <file> Attach a file to an existing item (web API plus sync).
zot rm <key>... Move items to the Zotero trash (web API plus sync, restorable).
zot config One-time setup of the Zotero web API key for the write commands.

Add --json to any command for machine-readable output, ready to pipe into jq. For scripts, note that progress and log lines go to stderr while only the result JSON goes to stdout, so do not merge the streams with 2>&1 before parsing.

search options

zot search "graph neural networks" \
  --tag "to-read" \
  --creator "Hamilton" \
  --type journalArticle \
  --collection ABCD1234 \
  --limit 20 \
  --rerank            # apply BGE reranker for higher precision (slower)

Before searching, zot makes one cheap call to Zotero to check whether the local index is still in sync, diffing item versions the same way zot index does. If the library has changed since the last zot index, it prints a note:

Note: index may be out of date -- 4 new/updated, 0 removed since last sync. Run `zot index` to update.

If Zotero is not reachable, the note instead says freshness could not be verified, and the search still runs against the local index. In --json mode the message is carried as a note field instead of printed. Skip the check with --no-sync-check (or ZOT_NO_SYNC_CHECK=1) for a fully offline, slightly faster search.

export options

zot export 6V9QI2NL VZUBFXFP                    # BibTeX on stdout
zot export --collection "EC3-2026" --output refs.bib   # a collection to a file
zot export --collection C39                     # by connector tree-view ID
zot export KEY --format csljson                 # or ris
zot search "protein folding" --export bibtex    # the search hits as a bibliography

Zotero renders the bibliography itself, so the output is exactly what its BibTeX, RIS and CSL JSON translators produce. --collection takes the same reference as everywhere else in the CLI, a key, an exact name, or a connector tree-view ID, and exports the collection's direct members without walking subcollections.

BibTeX arrives with fields that describe your library rather than the work, and they are stripped by default:

Field Why
file absolute path into your Zotero storage directory
annote the item's child notes
abstract a paragraph per entry, rarely wanted in a .bib
keywords your tags, not the publisher's

--raw keeps all of them. note and urldate are left alone: Zotero puts the Extra field in note, which routinely holds the arXiv ID or a version number, and urldate comes from the access date, which biblatex pairs with url on @online and @misc entries. RIS and CSL JSON carry none of these, so --raw changes nothing there and both are always passed through as Zotero rendered them, except that CSL JSON is re-serialized (and so pretty-printed with two-space indent) because merging several pages into one array requires parsing it.

Some items have no translator output at all: a standalone note or attachment, for instance. A rendered entry's citation key (jumperHighlyAccurateProtein2021) has no relation to the Zotero item key, so zot counts entries per request and, when a batch comes back short, re-requests its keys one at a time to find exactly which item was dropped. Each one is named on stderr:

Warning: no bibtex entry for "note: Curriculum" (VETJB3WE) -- the Zotero translator produced no entry for it.

Warnings always go to stderr, so stdout stays exactly the bibliography, and a --output file stays exactly the bibliography. Under --json stdout is one document instead: {format, requested, exported, path, content}. path and content are both always present with exactly one of them non-null, so the shape does not vary between runs: content holds the bibliography as a string, and --output sets path instead and leaves content null, since the file is then the payload. A dropped array appears alongside them only when something was dropped. zot search --export nests that same object under an export key alongside the search results, so one --json document still carries both.

find options

zot find "transformer" \
  --tag survey --creator "Vaswani" --type conferencePaper \
  --collection ABCD1234 \
  --sort dateAdded --desc \
  --everything \       # search all fields (default: title/creator/year)
  --limit 25

index

zot index            # incremental sync (only changed/new items)
zot index --force    # full rebuild from scratch
zot index --status   # item/chunk/vector counts, extraction-status breakdown, data dir
zot index issues     # list items with extraction problems and the reason

Fulltext is extracted from local Zotero data only, in this order per item: a PDF attachment (child or standalone), a locally stored HTML snapshot, or the note body for top-level notes. Nothing is ever fetched from the network; to make a URL-only item searchable, attach a snapshot in Zotero (drag the browser address-bar icon onto the item) and re-run zot index.

Every item carries a persisted extraction status shown by --status and detailed by issues:

Status Meaning
ok fulltext extracted and indexed
partial some PDF pages failed; the rest is indexed
suspicious extraction reported success but yielded implausibly little text (e.g. a scanned PDF with no text layer)
failed the file could not be processed (malformed, password-locked)
no-attachment nothing local to extract from

Items with failed, partial, or suspicious status are retried automatically on the next zot index run.

Adding papers (zot add)

zot add writes through the local connector API, the same endpoints the browser connector uses, so a single-collection add needs no account, no API key, and no sync, just the running Zotero app. Filing into several collections is the exception: the connector takes one collection, and the rest go through the web API (see below).

zot add 10.1038/nature14539                 # DOI (also doi.org URLs)
zot add arXiv:2401.12345                    # arXiv ID (also arxiv.org URLs)
zot add 978-0-262-03561-3                   # ISBN-10 or ISBN-13 (also ISBN:...,
                                            # hyphens and spaces optional)
zot add 23193287                            # PubMed ID (also PMID:... / PubMed:...)
zot add --pdf paper.pdf                     # PDF only: Zotero's recognizer
                                            # creates the metadata item
zot add 10.1000/xyz --pdf paper.pdf         # PDF + identifier (see below)
zot add ... --collection KMHNIPDA           # collection key, exact name, or
                                            # tree-view ID (default: library root)
zot add ... --collection Papers --collection "To Read"   # file in several
zot add ... --no-collection                 # library root on purpose, no warning
zot add ... --tag agents --tag to-read      # tags on the new item
zot add ... --force                         # skip the duplicate guard
zot add ... --no-index                      # skip the automatic index refresh

Behavior worth knowing:

  • Identifier forms: a bare value is read as a DOI when it starts with 10. and contains a /, as an arXiv ID when it looks like 2401.12345, as an ISBN when it normalises to 10 or 13 digits and its check digit matches, and as a PubMed ID when it is a plain number of up to 9 digits with no leading zero. The check digit is what keeps a 13-digit order number or timestamp from resolving as a book; a number that fails it is rejected rather than guessed at. The ISBN:, PMID: and PubMed: prefixes (case-insensitive) say which one you mean and give a type-specific error when the value is malformed.
  • Where the metadata comes from: DOIs use doi.org content negotiation and arXiv IDs use arxiv.org's BibTeX export, both as before. A PubMed ID makes one NCBI efetch call in MEDLINE form; when the record carries a DOI the publisher's BibTeX is used instead, since it is richer, and otherwise the MEDLINE record is mapped to BibTeX directly. An ISBN goes to OpenLibrary (openlibrary.org/api/books), which answers with author names inline in a single request, and is mapped to a BibTeX @book. There is no Google Books fallback: its keyless endpoint answers 429 for everyone sharing the anonymous project, so a second lookup would only delay the error.
  • Duplicate guard: before adding, the identifier is checked against the library (DOI, URL, extra fields). If it matches, the add is refused and reports the existing item's key; --force overrides. An ISBN is checked differently, because Zotero stores it hyphenated and its quicksearch does not match across the hyphens: the books are listed and their ISBN fields compared digit by digit, which catches a book whether it is stored hyphenated or not and whether its field holds one ISBN or the print and electronic ones together. One gap remains: a PubMed ID resolved through its DOI leaves an item carrying that DOI and no PMID, so re-adding the same PMID is not caught (re-adding its DOI is).
  • PDF recognition: with --pdf, the file is saved and Zotero's metadata recognizer creates the parent item, waiting until recognition finishes. An identifier given alongside the PDF is used only for the duplicate check and as a metadata fallback when recognition fails, so verify that the recognized metadata matches. If recognition fails entirely, the PDF is kept as a standalone attachment and, when an identifier was given, the metadata is imported separately; join them with zot attach or in the Zotero UI.
  • Several collections: --collection is repeatable and every value is resolved before anything is written, so a typo in the second one fails before the item exists. The connector saves into one collection only, so the first is filed on the spot and the rest are added afterwards through the web API, which needs a configured key (zot config set-key) and is checked before the add rather than after it. A library root (L1, or a group library's root) is only accepted as the sole value, since the collections after the first are written by collection key and a root has none; omitting --collection targets L1 anyway.
  • The web API only sees the item once Zotero has synced it up, so filing into more than one collection waits for that sync, polling every 2s for up to 60s and reporting progress on stderr. If the sync does not arrive in time the add still succeeds and exits 0, reporting the item as partially filed: it names the collections it is in, the ones it is not, and the exact zot edit KEY --add-collection ... to run once the sync catches up.
  • No collection: the item goes to the library root and becomes an unfiled item, which every collection-based view then misses, so the add warns on stderr and names the zot edit KEY --add-collection ... that files it. --no-collection is the opt-out for a deliberate root add and silences the warning; it cannot be combined with --collection, and that contradiction is rejected before anything is written. When the add produced only a standalone attachment (a PDF whose metadata Zotero could not recognize), the warning says so instead: that key is a stray to reparent with zot attach <item-key> <file>, which is what zot unfiled reports about it too. The warning is stderr only, so --json stdout stays a single document, where the same fact reads as collections: null; that key is always serialised, so a strict consumer can test it.
  • Index refresh: after a successful add, the search index updates incrementally so the paper is immediately findable via zot search.
  • No local delete: the connector API cannot remove items, so a mistaken add must be undone with zot rm (web API) or in the Zotero UI.
  • Adding by plain URL is on the roadmap, see BACKLOG.md; it needs Zotero's web translators, which the connector API does not expose.

Auditing filing (zot unfiled)

zot unfiled            # key, type and title for every unfiled top-level item
zot unfiled --count    # just the number, for a scripted filing-drift check

A pure local read that needs no API key and no index: an item is unfiled when it sits at top level and belongs to no collection.

Top-level attachments and notes are counted and listed apart from the rest. A standalone attachment with no parent is not a paper waiting to be filed, it is a file that wants reparenting or deleting, so acting on the main list never touches it. Under --json the two lists are items and stray, alongside the count and stray_count totals; --count sets both lists to null, which is distinct from the [] that means there are none. Human --count prints the unfiled count alone, so N=$(zot unfiled --count) is a number; --count --json is the way to a script that also wants stray_count.

Editing the library (zot edit, zot attach, zot rm)

Zotero's local API is read-only, so everything that modifies existing items goes through the Zotero web API (api.zotero.org) and reaches the local library on the next sync, usually within seconds when auto-sync is on. This requires Zotero sync and a one-time key setup:

# Create a key with write access at https://www.zotero.org/settings/keys
zot config set-key <API-KEY>     # stored + validated once; rerun to rotate
zot config show                  # config path, masked key, user ID

The ZOTERO_API_KEY env var overrides the stored key when set.

zot edit A1B2C3D4 --set date=2024 --set "publicationTitle=Nature"
zot edit A1B2C3D4 --add-tag reviewed --rm-tag to-read
zot edit A1B2C3D4 --add-collection ABCD1234 --rm-collection "To Read"
zot edit A1B2C3D4 --patch '{"creators":[{"creatorType":"author","firstName":"Ada","lastName":"Lovelace"}]}'
zot attach A1B2C3D4 paper.pdf --title "Preprint PDF"
zot rm A1B2C3D4 E5F6G7H8         # moves to trash (restorable in the UI)

--set uses Zotero field names (title, date, DOI, abstractNote, publicationTitle, and so on), and unknown fields are rejected by the API. Edits use optimistic concurrency: the item is read once, the write is guarded by that version, and a conflicting change landing in between aborts the write untouched so the command can be re-run against the current state.

--add-collection and --rm-collection are repeatable and take anything zot collections accepts: a collection key, an exact name, or a connector tree-view ID. Every value is resolved before anything is written, so a typo fails without a partial change, and the library root (L1) is rejected since "no collection" is not a collection. The output reports membership before and after (collections: [Inbox (ABCD1234)] -> [Inbox (ABCD1234), Read (EFGH5678)]). A request that changes nothing, adding a collection the item is already in, writes nothing and still exits 0, reporting No change with the item's unchanged version. A filing loop can therefore re-run over items it has already filed without special-casing them.

zot rm never deletes permanently, items go to the Zotero trash.

An item created locally moments ago, for example via zot add, must sync up before edit, attach, or rm can see it. If you get "not found on api.zotero.org", sync Zotero and retry.

Creating collections (zot collections --create)

zot collections --create "Reading list"                    # at the top level
zot collections --create "2026" --parent ABCD1234          # under a collection, by key
zot collections --create "2026" --parent "Reading list"    # parent by exact name
zot collections --create "2026" --parent C42               # parent by connector tree-view ID

--parent takes anything the rest of the tool takes: a collection key, an exact name, or a tree-view ID. Omit it to create at the top level; the library root L1 names the same place but only resolves while Zotero is running, since tree-view IDs come from the connector. The parent is resolved before the write, so an unknown one fails without creating anything. --parent is only meaningful with --create and is rejected without it, as is the listing positional together with --create. This is a web API write and needs the same key as zot edit.

A name a sibling under the same parent already carries is refused, naming the existing key. Zotero itself accepts the duplicate (verified against api.zotero.org), and every later zot collections NAME or --add-collection NAME would then be ambiguous. The comparison is exact, so two siblings differing only by case stay individually addressable; creating one prints a warning on stderr. The check reads the local library while the write goes upstream, so a sibling created seconds ago and not yet synced down is invisible to it and the duplicate goes through.

The new collection reaches the local library on the next sync. The command waits up to 20 seconds for it and reports whether it arrived (synced_local under --json); a collection that has not arrived yet exists upstream regardless and appears in zot collections once Zotero syncs, so re-running --create would make a second one rather than retry the first.

The output is the created collection alone, so --create cannot be combined with the listing flags --flat and --tree-ids.

Deleting a collection (zot collections --rm)

zot collections --rm ABCD1234            # by key
zot collections --rm "Reading list"      # by exact name
zot collections --rm C42                 # by connector tree-view ID
zot collections --rm ABCD1234 --force    # no prompt, for scripts

This is permanent. Zotero has no trash for collections, so unlike zot rm, which moves items to a trash you can empty or restore from in the Zotero UI, a deleted collection cannot be brought back from anywhere.

The items are not touched. Deleting a collection removes the grouping, not its contents: every item in it stays in the library, keeps any other collection it was filed in, and turns up in zot unfiled if that was its only one.

The delete cascades. Deleting a collection deletes every subcollection beneath it, in one request, because api.zotero.org removes the descendants itself. That is the case worth being careful about, and it is why the command prints the whole subtree before asking.

Before writing anything, it reports the collection's name and key, every descendant collection that goes with it, how many items are filed in that subtree, and how many of those would end up in no collection at all. Then it asks for confirmation on stderr, so --json stdout stays a single document:

  • a collection with no subcollections takes yes;
  • a collection with subcollections takes its own name, typed out, since the danger there is not knowing what hangs below the name you passed;
  • anything else aborts without writing.

--force skips the prompt. Without it, a non-interactive stdin (a pipe, a cron job, an agent) is refused rather than prompted or silently allowed, and the refusal names what would have been removed.

The summary is built from the local library while the cascade happens on api.zotero.org. Before asking anything, the command therefore reads every collection in the subtree from the server and compares its meta.numCollections with the children it is about to list. A subcollection created elsewhere and not yet synced down makes those counts disagree: the command refuses and tells you to sync Zotero, instead of prompting about a subtree it cannot show in full. A subtree of more than 50 collections is refused too, since checking it would mean 50 requests before the prompt; delete it in parts, or pass --force once you have checked it in the Zotero UI. The DELETE itself carries the collection's version, which rejects a concurrent write landing between that version being read and the delete going out. The command then waits up to 20 seconds for the removal to sync down and reports whether it arrived (synced_local under --json); until it does, zot collections still lists the tree even though it is gone upstream.

--rm is rejected together with --create, --parent, the listing positional, --flat and --tree-ids.

Where data lives

The local index is stored in the platform data directory, ~/Library/Application Support/zot/ on macOS and ~/.local/share/zot/ on Linux:

  ├── meta.json      # index metadata (model, sync state)
  ├── tantivy/       # BM25 full-text index
  └── vectors.bin    # embedding vectors

The web API key lives in the platform config directory, the same directory on macOS and ~/.config/zot/config.json on Linux.

To reset the index completely, delete the data directory or run zot index --force.

Troubleshooting

  • "Could not reach Zotero. Is it running?" means the Zotero desktop app is not open or the local API is disabled. Open Zotero and enable the setting under Settings → Advanced.
  • Search that returns nothing or looks stale usually means a stale index: run zot index to sync, or zot index --force for a clean rebuild.

License

MIT (see Cargo.toml).

About

A command-line interface (CLI) for managing and searching a personal Zotero library. It is useful for letting an agent help to manage a library.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages