返回 Skill 列表
extension
分类: 开发与工程无需 API Key

paper-cli

使用`paper` CLI工具的指南——一个具有AI驱动向量搜索功能的本地学术论文管理系统。每当用户想要管理学术论文、创建知识库、将文件(PDF、TXT、MD、TEX等)添加到知识库、进行语义论文搜索、配置嵌入模型,或者管理文献元数据和笔记时,请使用此技能。当用户提到“paper” CLI、研究用的知识库、文献管理,或是想要查询他们的论文收藏时也应触发此技能。即使用户只是说类似于“添加这个PDF”、“添加这个文本文件”或“搜索我的论文”这样的内容,在使用paper-manager的项目中,此技能也应该被激活。

person作者: jakexiaohubgithub

paper CLI — Academic Paper Management

paper is a local-first CLI tool for managing academic papers with AI-powered semantic search. It stores data in SQLite databases and uses FAISS vector stores for similarity search.

Core Concepts

Scope: User vs Project

Every resource (config, knowledge base, literature) lives in one of two scopes:

  • Project scope (default): stored in ./.paper-manager/ — tied to the current directory, good for project-specific paper collections
  • User scope (--user flag): stored in ~/.paper-manager/ — shared across all projects, good for a personal paper library

Project-level config overrides user-level config. When looking up a knowledge base or literature, the tool checks project scope first, then falls back to user scope.

Data Hierarchy

Config (embedding models)
  └── Knowledge Base (has a name, description, and embedding model)
        └── Literature (a paper with metadata, source file, and vector embeddings)
              └── Notes (key-value pairs for personal annotations)

You must configure an embedding model before creating a knowledge base, and you must create a knowledge base before adding literature.

Troubleshooting: Installation & Setup

The user should already have paper installed and configured. Only use this section if something is missing.

Installation (if paper command is not found)

npm install -g paper-manager    # or: pnpm add -g paper-manager

Requires Node.js >= 24. Verify with paper --version.

Config setup (if commands fail due to missing config)

# If data directory is missing or incomplete, initialize it first
paper config init [--user]

# Set up an embedding model (OpenAI-compatible provider)
paper config set embeddingModels '{"my-model": {"provider": "openai", "model": "text-embedding-3-small", "apiKey": "sk-...", "dimensions": 1536}}' --user

# Set a default model so you don't have to specify it every time
paper config set defaultEmbeddingModelId my-model --user

The embedding model config fields:

  • provider: currently only "openai" (works with any OpenAI-compatible API)
  • model: the model name (e.g., "text-embedding-3-small")
  • baseUrl: optional custom API endpoint
  • apiKey: API key for the provider
  • dimensions: embedding vector dimensions (e.g., 1536)
  • batchSize: optional max number of texts per embedding API request (set this if your provider limits batch size)

Command Reference

paper config — Configuration Management

paper config init [--user]              # Initialize data directory (config, db, subdirs)
paper config list [--user]              # Show all config (merged or user-only)
paper config get <key> [--user]         # Get a specific config value
paper config set <key> <value> [--user] # Set a config value (value is parsed as JSON, falls back to string)
paper config remove <key> [--user]      # Remove a config key

Config keys:

  • embeddingModels — a JSON object of { [modelId]: { provider, model, baseUrl?, apiKey, dimensions, batchSize? } }
  • defaultEmbeddingModelId — which model ID to use when none is specified
  • email — email address for Unpaywall API identification (required for lit add --doi)

paper kb — Knowledge Base Management

# Create a knowledge base (requires an embedding model in config)
paper kb create <name> -d <description> [-e <embedding-model-id>] [--user]

# List knowledge bases
paper kb list [--user | --all] [--json] [--jq <expression>]

# Update knowledge base metadata (name and/or description)
paper kb update <id> [-n <name>] [-d <description>]

# Remove a knowledge base and ALL its data (literatures, PDFs, vectors)
paper kb remove <id>

# Semantic search across a knowledge base
paper kb query <id> <query-text> [-k <top-k>] [--json] [--jq <expression>]   # default top-k is 5

The <id> for knowledge bases is a UUID assigned at creation time. Use paper kb list to find it.

paper lit — Literature Management

# Add a paper (extracts content, splits text, creates embeddings)
# Supports PDF, TXT, MD, TEX, and other text-based formats
# For PDFs, automatically extracts metadata (title, author, keywords, DOI, etc.)
# If opendataloader-pdf is available, also converts PDF to Markdown automatically
paper lit add <kb-id> <file-path> [-t <title>] [-f]
# Title defaults to PDF metadata title, then filename if not specified
# Rejects duplicate DOI in the same knowledge base; use -f/--force to override

# Add Open Access paper by DOI via Unpaywall API (downloads PDF automatically)
# Requires email config: paper config set email "you@example.com"
# Errors if paper is not Open Access — use file mode instead for non-OA papers
paper lit add <kb-id> --doi <doi> [-t <title>] [-f]

# Convert an existing literature PDF to Markdown (requires opendataloader-pdf)
paper lit convert <lit-id>

# List papers in a knowledge base (shows associated files: PDF, MD, etc.)
paper lit list <kb-id> [--json] [--jq <expression>]

# Search papers in a KB by metadata (at least one filter required)
paper lit search <kb-id> [-t <title>] [-a <author>] [-k <keyword>] [--doi <doi>] [--json] [--jq <expression>]
# Filters use substring matching; combining filters narrows results (AND)

# Show full details of a paper (shows associated files: PDF, MD, etc.)
paper lit show <kb-id> <lit-id> [--json] [--jq <expression>]

# Update paper metadata
paper lit update <kb-id> <lit-id> [options]
#   -t, --title <title>
#   --title-translation <translation>
#   -a, --author <author>
#   --abstract <abstract>
#   --summary <summary>
#   --url <url>
#   --doi <doi>
#   --keywords <comma-separated-keywords>

# Remove a paper (deletes DB record, source file; vectors remain in store)
paper lit remove <kb-id> <lit-id>

paper lit note — Literature Notes

Notes are key-value string pairs attached to a literature entry — useful for personal annotations.

paper lit note list <lit-id> [--json] [--jq <expression>]  # List all notes
paper lit note set <lit-id> <key> <value>       # Set a note
paper lit note remove <lit-id> <key>            # Remove a note

Note: the note commands take <lit-id> directly (not <kb-id> <lit-id>).

paper dep — Dependency Management

# Check all external dependencies
paper dep check

# Check a specific dependency
paper dep check opendataloader   # Checks Java runtime + @opendataloader/pdf package

opendataloader-pdf is used for high-quality PDF-to-Markdown conversion. It requires Java 11+ and the @opendataloader/pdf npm package (optional dependency). If unavailable, lit add silently skips the conversion; lit convert will report what's missing.

paper util — Utilities

# Convert a DOI to BibTeX citation (accepts DOI identifier or full URL)
paper util doi2bib <doi>

# Extract metadata from a PDF file (title, author, subject, keywords, DOI, dates)
paper util pdf-meta <file> [--json] [--jq <expression>]

Common Workflows

Start a new research project

  1. paper kb create "my-project" -d "Papers about X" — create a project-scoped KB
  2. paper lit add <kb-id> ./paper.pdf -t "Paper Title" — add papers (also works with .txt, .md, .tex)
  3. paper kb query <kb-id> "your research question" — search

Add a paper by DOI (Open Access)

  1. paper config set email "you@example.com" — set email for Unpaywall API (one-time)
  2. paper lit add <kb-id> --doi 10.1038/nature12373 — automatically downloads OA PDF and ingests it

Manage paper metadata

  1. paper lit list <kb-id> — find the literature ID
  2. paper lit update <kb-id> <lit-id> -a "Author Name" --keywords "ML,NLP"
  3. paper lit note set <lit-id> takeaway "Key insight from this paper"

Find the stored file for a literature

Source files are stored at <scope-dir>/files/<lit-id>.<ext> (e.g., .paper-manager/files/f47ac10b-58cc-4372-a567-0e02b2c3d479.pdf). To locate a literature's file, use paper lit list <kb-id> to get the literature ID, then look in the files/ directory under the appropriate scope directory (.paper-manager/ for project scope, ~/.paper-manager/ for user scope).

JSON Output (--json) and jq Filtering (--jq)

Most read commands support --json for machine-readable output and --jq <expression> for inline jq filtering.

--jq implies --json — you don't need to pass both. If both are provided, --json is ignored. The jq implementation is built-in (pure TypeScript via @eurfelux/jq-js), so no external jq binary is needed.

Examples:

# Get just the names of all knowledge bases
paper kb list --jq '.[].name'

# Get titles of papers by a specific author
paper lit search <kb-id> -a "Smith" --jq '[.[] | .title]'

# Extract DOI from a PDF
paper util pdf-meta paper.pdf --jq '.doi'

kb list --json

[
  {
    "id": "uuid",
    "name": "...",
    "description": "...",
    "embeddingModelId": "...",
    "scope": "project|user",
    "createdAt": "ISO",
    "updatedAt": "ISO"
  }
]

kb query --json

[
  {
    "pageContent": "chunk text...",
    "metadata": {
      "literatureId": "uuid",
      "source": "path",
      "loc": { "pageNumber": 1 }
    }
  }
]

lit list --json / lit search --json

[{ "id": "uuid", "title": "...", "author": "...", "keywords": [], "doi": "...", "knowledgeBaseId": "uuid", "createdAt": "ISO", "updatedAt": "ISO", ... }]

lit show --json

Single literature object (same shape as array elements above).

lit note list --json

{ "key1": "value1", "key2": "value2" }

Integration Guides

  • Unpaywall API — add Open Access papers by DOI (paper lit add --doi)
  • opendataloader-pdf — high-quality PDF-to-Markdown conversion with image extraction

Important Notes

  • All IDs (knowledge base, literature) are UUIDs — always use list commands to look them up
  • Adding a paper (lit add) extracts the file content (PDF or text), splits it into chunks, and creates vector embeddings — this calls the embedding API and may take some time
  • Text files have a 10 MB size limit
  • The kb remove command is destructive: it deletes the knowledge base, all its literatures, source files, and vector stores
  • The --user flag on config/kb create controls scope; omitting it uses project scope
  • Config values passed to config set are parsed as JSON first; if JSON parsing fails, the raw string is stored

Skill Maintenance

Skill version: v0.12.1. To update to the latest version, run npx/pnpx/bunx skills add paper-manager.