Cross-platform command-line documentation puller and converter.
It pulls docs from:
- GitHub repos (prefer raw
.md/.mdx) - HTML documentation sites (crawl + convert)
- PDFs (local or URL)
- EPUBs (local or URL)
And exports a clean folder of Markdown files plus an index.md.
These steps assume you have Python 3.10+ installed and available as python.
- Clone the repo and enter it:
git clone https://github.com/trebory6/DocHarvester.git
cd DocHarvester- Install with
pipx:
pipx install .- Verify the command is available:
docpull --helpTo upgrade later:
git pull
pipx reinstall .To uninstall:
pipx uninstall docharvester- Clone the repo and enter it:
git clone https://github.com/trebory6/DocHarvester.git
cd DocHarvester- Create a virtual environment:
python -m venv .venv- Activate it and install:
Windows PowerShell:
.\.venv\Scripts\Activate.ps1
pip install -U pip
pip install .
docpull --helpLinux/macOS:
source .venv/bin/activate
pip install -U pip
pip install .
docpull --help- Running later (venv):
- Windows PowerShell:
.\.venv\Scripts\Activate.ps1
docpull- Linux/macOS:
source .venv/bin/activate
docpullDocHarvester runs interactive by default:
docpullNon-interactive examples:
docpull pull --source "https://docs.example.com" --output "./Example Docs"
docpull pull --source "./book.epub" --output "./Book Markdown"
docpull pull --source "https://example.com/file.pdf" --output "./PDF Output" --yes
docpull pull --source "https://github.com/org/repo" --output "./Repo Docs" --githubYou can also use the project name command:
docharvester pull --source "https://docs.python.org/3/" --output "./Python Docs"docpull pull --source "https://docs.example.com" --output "./Docs" --max-pages 200 --max-depth 8 --delay 0.5
docpull pull --source "https://docs.example.com" --output "./Docs" --ignore-robotsIf you hit GitHub API limits, set GITHUB_TOKEN:
- Windows PowerShell:
$env:GITHUB_TOKEN="ghp_..."- Linux/macOS:
export GITHUB_TOKEN="ghp_..."Each run creates a subfolder inside your chosen output directory, and writes:
- One
.mdper page/chapter/file (depending on source) index.mdlinking to the pulled Markdown