Readwise Reader Archive
Date: 2026-06-15
Status: Approved
Tool: readwise-reader.py → new archive-html subcommand
Purpose
Archive HTML content of all Readwise Reader documents the user has actually
read (reading_progress > 0). Acts as a data sovereignty / exit strategy backup.
Complementary to go-org-readwise which handles highlight-rich content →
org-mode (the gap: read-but-not-highlighted documents don’t get archived).
Long-term, this tool may consolidate both workflows.
MVP — archive-html subcommand
Command
readwise-reader archive-html [--output DIR]
Pipeline:
- Fetch document list from Reader API (all locations, metadata only)
- Filter to documents with
reading_progress > 0 - Fetch HTML content for uncached documents (
withHtmlContent=true) - Save with human-readable filenames:
YYYY-MM-DD-slugified-title.html
File Layout
~/.local/share/readwise/
├── html/ # Archived HTML (persistent, incremental)
│ ├── 2026-03-15-nix-flake-patterns-by-tweag.html
│ ├── 2026-04-02-go-error-handling-rob-pike.html
│ └── 2026-05-20-kubernetes-gateway-api-overview.html
├── html-manifest.json # doc-id → filename mapping (for incremental)
├── reader-latest.json # Existing (unchanged)
└── reader-analyzed.json # Existing (unchanged)
HTML Content Fetching
- Reader API
withHtmlContent=truereturns Reader’s parsed/cleaned HTML (not raw original page). - Cached per document. A manifest (
html-manifest.json) maps doc IDs to filenames for incremental runs — only fetch docs not already archived. - Rate limited: 20 req/min for Reader API, with 3.1s delay between requests.
Filename Convention
YYYY-MM-DD-slugified-title.html where:
- Date is the
saved_atdate from Reader - Title is slugified (lowercase, non-alphanumeric → hyphens, truncated to ~80 chars)
- Collisions appended with
-2,-3, etc.
Browseable via nginx autoindex on; with no extra tooling.
NixOS Integration
Home-Manager Module
New home/common/services/readwise-archive.nix:
services.readwise-archive.enableservices.readwise-archive.interval(default:"daily")services.readwise-archive.outputDir(default:~/.local/share/readwise/html)- Systemd user service + timer (
Persistent = true) - Wraps
readwise-reader archive-htmlwith API key from passage
Deployment
- Host: okinawa (already runs readwise-sync and readwise-reader services)
- Nginx:
autoindex on;on the html directory, served on a port
Dependencies
- No new Python dependencies (no Jinja2, no Pagefind for MVP)
Future Phase — Static Site (nice-to-have)
All of the following can be built on top of the archived HTML cache:
readwise-reader sitesubcommand generating a proper static site- Jinja2 templates: index page (chronological list) + per-document pages with metadata headers
- CSS from
https://vincent.demeester.fr/style.css(light + dark mode) - Client-side search via Pagefind
- Tag/category filters (flux-filter style)
- Separate
reader.cssfor archive-specific styling
These become trivial to add since the HTML content is already cached and incrementally maintained by the MVP.
What This Is Not
- No AI/LLM summaries or analysis
- No server-side logic — purely static files
- Not public (homelab only, can change later)
- Not a replacement for
go-org-readwise(complementary, potential future consolidation)