Metadata-Version: 2.1
Name: abx-dl
Version: 1.13.154
Summary: All-in-one CLI tool to download and extract content from URLs
Keywords: scraping,crawling,downloading,internet archiving,web archiving,digipres,warc,preservation,backups,archiving,web,bookmarks,puppeteer,browser,download
Author: Nick Sweeting, ArchiveBox
License: MIT
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Environment :: Web Environment
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Education
Classifier: Intended Audience :: End Users/Desktop
Classifier: Intended Audience :: Information Technology
Classifier: Intended Audience :: Legal Industry
Classifier: Intended Audience :: System Administrators
Classifier: License :: OSI Approved :: MIT License
Classifier: Natural Language :: English
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Internet :: WWW/HTTP :: Indexing/Search
Classifier: Topic :: Internet :: WWW/HTTP :: WSGI :: Application
Classifier: Topic :: Sociology :: History
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: System :: Archiving
Classifier: Topic :: System :: Archiving :: Backup
Classifier: Topic :: System :: Recovery Tools
Classifier: Topic :: Utilities
Classifier: Typing :: Typed
Project-URL: Homepage, https://abx-dl.archivebox.io/
Project-URL: Source, https://github.com/ArchiveBox/abx-dl
Project-URL: Documentation, https://abx-dl.archivebox.io/
Project-URL: Bug Tracker, https://github.com/ArchiveBox/abx-dl/issues
Project-URL: Changelog, https://github.com/ArchiveBox/abx-dl/releases
Project-URL: Community, https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community
Project-URL: Donate, https://github.com/ArchiveBox/ArchiveBox/wiki/Donations
Requires-Python: <3.15,>=3.12
Requires-Dist: rich-click>=1.8.0
Requires-Dist: rich>=13.0.0
Requires-Dist: abxbus==2.5.76
Requires-Dist: abxpkg==1.13.42
Requires-Dist: abx-plugins==1.13.198
Requires-Dist: pydantic>=2.0.0
Requires-Dist: pydantic-settings>=2.0.0
Requires-Dist: platformdirs>=4.0.0
Requires-Dist: requests>=2.28.0
Requires-Dist: psutil!=7.2.2,>=7.2.1
Description-Content-Type: text/markdown

# ⬇️ `abx-dl`

> A simple all-in-one CLI tool to auto-detect and download *everything* available from a URL.

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
uvx abx-dl 'https://example.com'
```

<!--pytest-codeblocks:cont-->
<!--
```bash
test -s index.jsonl
test -s title/title.txt
test -s wget/example.com/index.html
```
-->

---

✨ *Ever wish you could `yt-dlp`, `gallery-dl`, `wget`, `curl`, `puppeteer`, etc. all in one command?*

`abx-dl` is an all-in-one CLI tool for downloading URLs "by any means necessary".

It's useful for scraping, downloading, OSINT, digital preservation, and more.
`abx-dl` provides a simpler one-shot CLI interface to the [ArchiveBox plugin ecosystem](https://plugins.archivebox.io/).

<img width="1000" height="1082" alt="Screenshot 2026-03-11 at 6 53 03 AM" src="https://github.com/user-attachments/assets/4e19d985-1a93-4f65-9970-2565be16b718" />


---

<br/>

#### 🍜 What does it save?

<!--
```bash
set -Eeuo pipefail
trap 'status=$?; printf "README crawl failed: %s (exit %s)\n" "$BASH_COMMAND" "$status" >failure.log; find . -maxdepth 3 -type f | sort >>failure.log; exit "$status"' ERR
scratch_dir="${ABX_DL_DOCS_OUTPUT_DIR:-$(mktemp -d)}"
mkdir -p "$scratch_dir"
cd "$scratch_dir"
exec >stdout.log 2>stderr.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
abx-dl --plugins=wget,title,screenshot,pdf,readability 'https://example.com'
```

<!--pytest-codeblocks:cont-->
<!--
```bash
test -s index.jsonl
test -s title/title.txt
test -s wget/example.com/index.html
test -s screenshot/screenshot.png
test -s pdf/output.pdf
test -s readability/content.html
grep -q 'Example Domain' title/title.txt
grep -q 'Example Domain' wget/example.com/index.html
# Readability extracts article body, which may omit the heading. The live
# example.com DOM now does; title extraction is checked separately above.
grep -Fq 'This domain is for use in documentation examples without needing permission.' readability/content.txt
grep -q '"plugin": "wget".*"status": "succeeded"' index.jsonl
grep -q '"plugin": "screenshot".*"status": "succeeded"' index.jsonl
grep -q '"plugin": "pdf".*"status": "succeeded"' index.jsonl
grep -q '"plugin": "title".*"status": "succeeded"' index.jsonl
grep -q '"plugin": "readability".*"status": "succeeded"' index.jsonl
```
-->

`abx-dl` runs all plugins by default (and auto installs dependencies). You can specify `--plugins=wget,favicon,title` or filters like `--output=html,pdf,ico,text/` to limit plugin selection.
- HTML, JS, CSS, images, etc. rendered with a headless browser
- title, favicon, headers, outlinks, and other metadata
- audio, video, subtitles, playlists, comments
- snapshot of the page as a PDF, screenshot, and [Singlefile](https://github.com/gildas-lormeau/single-file-cli) HTML
- article text, `git` source code
- [and much more](https://plugins.archivebox.io/)...

<br/>

#### 🧩 How does it work?

`abx-dl` uses the **[Plugin Library](https://plugins.archivebox.io/)** (shared with [ArchiveBox](https://github.com/ArchiveBox/ArchiveBox)) to run a collection of downloading and scraping tools.

Plugins are loaded from the installed [`abx-plugins`](https://pypi.org/project/abx-plugins/) package (or from `ABX_PLUGINS_DIR` if you override it) and execute in distinct phases:
1. **Install phase** runner reads plugins `config.json`: `required_binaries` and emits `BinaryRequestEvent`s for `abxpkg.binary_service.BinaryService`, which resolves or installs binaries using built-in providers such as env, pip, npm, brew, apt, cargo, and browser-specific providers. `abxpkg` owns the persistent binary cache; `abx-dl` only projects resolved paths into the current run's in-memory config.
2. **CrawlSetup hooks** (`on_CrawlSetup__*`) launch/configure expensive crawl-scoped processes like chrome, or trigger side effects. background hooks use their first stdout line as the readiness boundary and emit no stdout JSONL records.
3. **Snapshot hooks** (`on_Snapshot__*`) run per URL to extract content. background hooks use their first stdout line as the readiness boundary; JSONL records after that are `ArchiveResult`, `Snapshot`, and `Tag`.

Applications embedding the runtime use the same framework-free interfaces as
the CLI: `PluginCatalog` for inventory, `PluginConfigResolver` for config,
the service classes for explicit listener composition, typed events for phase
dispatch, `execute_hook()` for a single finite hook, and `OutputManifest` for
output metadata. These APIs accept plain mappings, filesystem paths,
environment variables, and CLI arguments; they do not depend on Django or an
application database.

Standalone `download()` attaches both `CrawlService` (plugin crawl hooks) and
`CrawlLifecycleService` (phase sequencing). Embedders can attach only the
listener suites whose behavior they want and dispatch the corresponding typed
events directly.

`parse_input(source_text, catalog, output_dir)` is the framework-free import
path for pasted text and bookmark/feed exports. It writes
`staticfile/stdin.txt`, runs only plugins declaring
`x-accepts-internal-input`, and returns metadata-preserving `Snapshot` facts at
depth zero. It does not create a crawl, database row, or synthetic URL.


<br/>

#### ⚙️ Configuration

Configuration is handled via environment variables plus a user config file under the platformdirs user config path (`<user-config>/abx/config.env`). Resolved binary paths are projected into the current run in memory; persistent binary state lives only in the abxpkg library cache:

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
abx-dl config                        # show all config (global + per-plugin)
abx-dl config --get WGET_TIMEOUT     # get a specific value
abx-dl config --set TIMEOUT=120      # set persistently (resolves aliases)
```

Output is grouped by section:
```text
# GLOBAL
TIMEOUT=60
USER_AGENT="Mozilla/5.0 ..."
...

# plugins/wget
WGET_BINARY="wget"
WGET_TIMEOUT=60
...

# plugins/chrome
CHROME_BINARY="chromium"
...
```

Common options:
- `TIMEOUT=60` - default timeout for hooks
- `USER_AGENT` - default user agent string
- `{PLUGIN}_BINARY` - path or name of the binary to use (e.g. `WGET_BINARY=wget` or `CHROME_BINARY=/usr/bin/chromium`)
- `{PLUGIN}_ENABLED=True/False` - enable/disable specific plugins
- `{PLUGIN}_TIMEOUT=120` - per-plugin timeout overrides

Aliases are automatically resolved (e.g. `--set USE_WGET=false` saves as `WGET_ENABLED=false`).

One-off config is easy via env vars or CLI args:

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
env \
  TIMEOUT=120 \
  WGET_TIMEOUT=120 \
  abx-dl \
    --dir=./config-example \
    --plugins=title,wget \
    --timeout=90 \
    'https://example.com'
```

<!--pytest-codeblocks:cont-->
<!--
```bash
test -s config-example/index.jsonl
test -s config-example/title/title.txt
test -s config-example/wget/example.com/index.html
grep -q 'Example Domain' config-example/title/title.txt
grep -q '"plugin": "wget".*"status": "succeeded"' config-example/index.jsonl
grep -q '"plugin": "title".*"status": "succeeded"' config-example/index.jsonl
```
-->

<br/>

---

<br/>

### 📦 Install

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
uv tool install abx-dl
abx-dl version
```

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
uvx abx-dl version
```

```bash
abx-dl install wget title
```

#### Docker

The image includes the downloader plugins and their dependencies. Mount an
output directory at `/out` to keep the downloaded files:

<!--pytest.mark.docker_required-->
```bash
set -Eeuo pipefail
output_dir="$(mktemp -d)"
image="${ABXDL_IMAGE:-archivebox/abx-dl:latest}"
trap 'rm -rf "$output_dir"' EXIT
docker run --rm \
  --volume "$output_dir:/out" \
  "$image" \
  --no-install --plugins=title,wget 'https://example.com'
test -s "$output_dir/index.jsonl"
test -s "$output_dir/title/title.txt"
test -s "$output_dir/wget/example.com/index.html"
grep -q 'Example Domain' "$output_dir/title/title.txt"
grep -q 'Example Domain' "$output_dir/wget/example.com/index.html"
```

To persist browser personas, also mount their directory at `/data/personas`.
The image does not create an anonymous persona volume.

<br/>

### 🔠 Usage

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
abx-dl --plugins=title,wget --dir=./downloads --timeout=120 'https://example.com'
```

<!--pytest-codeblocks:cont-->
<!--
```bash
test -s downloads/index.jsonl
test -s downloads/title/title.txt
test -s downloads/wget/example.com/index.html
```
-->

```text
# Default command - a bare URL archives with all enabled plugins:
abx-dl 'https://example.com'

# Select plugins by output type (mimetypes, categories, or file extensions):
abx-dl --output=html,pdf,video/ 'https://example.com'
abx-dl -o text -o image -o mp4 'https://example.com'

# Limit work to a subset of plugins by name:
abx-dl --plugins=wget,title,screenshot,pdf 'https://example.com'

# Skip auto-installing missing dependencies (emit warnings instead):
abx-dl --no-install 'https://example.com'

# Specify output directory (default is current working dir):
abx-dl --dir=./downloads 'https://example.com'

# Set timeout:
abx-dl --timeout=120 'https://example.com'
```

#### Commands

```text
abx-dl <url>                              # Download URL (default shorthand)
abx-dl plugins                            # Check + show info for all plugins
abx-dl plugins wget ytdlp git             # Check + show info for specific plugins
abx-dl install wget ytdlp git             # Pre-install plugin dependencies
abx-dl config                             # Show all config values
abx-dl config --get TIMEOUT               # Get a specific config value
abx-dl config --set TIMEOUT=120           # Set a config value persistently
```

#### Installing Dependencies

Many plugins require external binaries (e.g., `wget`, `chrome`, `yt-dlp`, `single-file`).

By default, `abx-dl` lazily installs missing dependencies as needed when you download a URL.
Use `--no-install` to skip plugins with missing dependencies instead. `install` runs only the pre-run dependency pipeline (`required_binaries` → `BinaryRequestEvent` → `BinaryEvent`) without starting crawl setup or snapshot extraction:

<!--
```bash
cd "$(mktemp -d)"
exec >stdout.log
```
-->
<!--pytest-codeblocks:cont-->
```bash
abx-dl install wget title
abx-dl plugins wget title
```

```text
abx-dl 'https://example.com'              # checks and installs missing deps before hooks run
abx-dl --no-install 'https://example.com' # skips plugins with missing deps and emits warnings
abx-dl install wget singlefile ytdlp      # installs dependencies for specific plugins only
abx-dl plugins                            # checks which dependencies are available/missing
```

Every preflight request is resolved through `abxpkg`. Compatible host binaries are selected first and projected into `ABXPKG_LIB_DIR/env/bin`; otherwise the configured managed provider installs and projects the dependency. Hook subprocesses then use the resolved Python or Node interpreter and projected runtime environment directly.

During capture, a dependency failure leaves that plugin's real hook to report its failure while unrelated plugins continue. Explicit `abx-dl install` commands still fail when a requested dependency cannot be installed. Crawl setup readiness failures remain fatal because later hooks may require those shared services.

The normal runtime flow after dependency preflight is:
- `CrawlEvent` (internal lifecycle root)
- `CrawlSetupEvent` → plugin `on_CrawlSetup__*` hooks
- `CrawlStartEvent` → `SnapshotEvent`
- `SnapshotEvent` → plugin `on_Snapshot__*` hooks
- `SnapshotCleanupEvent` / `CrawlCleanupEvent`

Hook output contract:
- `EXTRA_CONTEXT` is opaque correlation data reflected into output records only. Hooks must never inspect it. Snapshot hooks receive `--url`, `--snapshot-id`, and `--depth` as explicit arguments; archived titles/tags stay in filesystem data. Embedders provide `download(snapshot=Snapshot(...))` for an existing snapshot, not identity or depth hidden in configuration.
- binary preflight is driven by plugin `required_binaries` and handled by `abxpkg`, not by plugin hooks
- `on_CrawlSetup__*` background hooks emit a first stdout readiness line, but no stdout JSONL records
- `on_Snapshot__*` background hooks emit a first stdout readiness line; hook JSONL records after that are only `ArchiveResult`, `Snapshot`, and `Tag`
- the TUI and services consume structured events derived from those hook records

Dependencies are installed to `<user-config>/abx/lib/{arch}/` using the appropriate package manager:
- **pip packages** → `<user-config>/abx/lib/{arch}/pip/venv/`
- **npm packages** → `<user-config>/abx/lib/{arch}/npm/`
- **brew/apt packages** → system locations

You can override the install location with `ABXPKG_LIB_DIR=/path/to/lib abx-dl install wget`.

<br/>

---

<br/>

### Output Structure

By default, `abx-dl` writes results into the current working directory. Each run creates an `index.jsonl` manifest plus one subdirectory per plugin that produced output. If you want to keep runs isolated, `cd` into a scratch directory first or pass `--dir=/path/to/run`.

```bash
mkdir -p /tmp/abx-run && cd /tmp/abx-run
uvx --from abx-dl abx-dl --plugins=title,wget 'https://example.com'
```

```
./
├── index.jsonl             # Snapshot metadata and results (JSONL format)
├── title/
│   └── title.txt
├── favicon/
│   └── favicon.ico
├── screenshot/
│   └── screenshot.png
├── pdf/
│   └── output.pdf
├── dom/
│   └── output.html
├── wget/
│   └── example.com/
│       └── index.html
├── singlefile/
│   └── output.html
└── ...
```

<br/>

### All Outputs

- `index.jsonl` - snapshot metadata and plugin results (JSONL format, ArchiveBox-compatible)
- `title/title.txt` - page title
- `favicon/favicon.ico` - site favicon
- `screenshot/screenshot.png` - full page screenshot (Chrome)
- `pdf/output.pdf` - page as PDF (Chrome)
- `dom/output.html` - rendered DOM (Chrome)
- `wget/example.com/...` - mirrored site files
- `singlefile/output.html` - single-file HTML snapshot
- ... and more via plugin library ...

---

### Available Plugins

See the [`abx-plugins` marketplace](https://github.com/ArchiveBox/abx-plugins).

#### Snapshot / Extraction Plugins

- `ytdlp` - downloads media plus sidecars: audio, video, images/thumbnails, subtitles (`.srt`, `.vtt`), JSON metadata, and text descriptions.
- `gallerydl` - downloads gallery/media sets as images, videos, JSON sidecars, text sidecars, and ZIP archives.
- `forumdl` - exports forum/thread archives as JSONL, WARC, and mailbox-style message archives.
- `git` - clones repository contents including text, binaries, images, audio, video, fonts, and other tracked files.
- `wget` - mirrors pages and requisites as HTML, WARC, images, CSS, JavaScript, fonts, audio, and video.
- `archivedotorg` - saves a Wayback Machine archive link as plain text.
- `favicon` - saves site favicons and touch icons as image files.
- `modalcloser` - setup helper only; no direct archive files.
- `consolelog` - saves browser console events as JSONL.
- `dns` - saves observed DNS activity as JSONL.
- `ssl` - saves TLS certificate/connection metadata as JSONL.
- `responses` - saves HTTP response metadata as JSONL and can record referenced text, images, audio, video, apps, and fonts.
- `redirects` - saves redirect chains as JSONL.
- `staticfile` - saves non-HTML direct file responses such as PDF, EPUB, images, audio, video, JSON, XML, CSV, ZIP, and generic binary files.
- `headers` - saves main-document HTTP headers as JSON.
- `chrome` - manages shared browser state and emits plain-text and JSON runtime metadata.
- `seo` - saves SEO metadata such as meta tags and Open Graph fields as JSON.
- `accessibility` - saves the browser accessibility tree as JSON.
- `infiniscroll` - page-expansion helper only; no direct archive files.
- `claudechrome` - saves Claude-computer-use interaction results as JSON plus PNG screenshots.
- `singlefile` - saves a full self-contained page snapshot as HTML.
- `screenshot` - saves rendered page screenshots as PNG.
- `pdf` - saves rendered pages as PDF.
- `dom` - saves fully rendered DOM output as HTML.
- `title` - saves the final page title as plain text.
- `readability` - extracts article HTML, plain text, and JSON metadata.
- `defuddle` - extracts cleaned article HTML, plain text, and JSON metadata.
- `mercury` - extracts article HTML, plain text, and JSON metadata.
- `claudecodeextract` - generates cleaned Markdown from other extractor outputs.
- `htmltotext` - converts archived HTML into plain text.
- `trafilatura` - extracts article content as plain text, Markdown, HTML, CSV, JSON, and XML/TEI.
- `papersdl` - downloads academic papers as PDF.
- `parse_html_urls` - emits discovered links from HTML as JSONL records.
- `parse_txt_urls` - emits discovered links from text files as JSONL records.
- `parse_rss_urls` - emits discovered feed entry URLs from RSS/Atom as JSONL records.
- `parse_netscape_urls` - emits discovered bookmark URLs from Netscape bookmark exports as JSONL records.
- `parse_jsonl_urls` - emits discovered bookmark URLs from JSONL exports as JSONL records.
- `parse_dom_outlinks` - emits crawlable rendered-DOM outlinks as JSONL records.
- `search_backend_sqlite` - writes a searchable SQLite FTS index database.
- `search_backend_sonic` - pushes content into Sonic search; no local archive files declared.
- `claudecodecleanup` - writes cleanup/deduplication results as plain text.
- `hashes` - writes file hash manifests as JSON.
- and more via the [`abx-plugins` marketplace](https://github.com/ArchiveBox/abx-plugins)...

---

### AI Skill

This repo includes an `abx-dl` skill for coding agents that need to run the standalone ArchiveBox extractor pipeline without a full ArchiveBox install.

- Skill source: [`skills/abx-dl/SKILL.md`](./skills/abx-dl/SKILL.md)
- skills.sh page: https://skills.sh/archivebox/abx-dl/abx-dl

---

### Architecture

`abx-dl` is built on these components:

- **`abx_dl/catalog.py`** - Plugin discovery, selection, and config resolution from `abx-plugins` or `ABX_PLUGINS_DIR`
- **`abx_dl/executor.py`** - Hook execution engine with config propagation
- **`abx_dl/config.py`** - Environment variable configuration
- **`abx_dl/cli.py`** - Rich CLI with live progress display

### Related Projects

- `abxbus` https://abxbus.archivebox.io https://github.com/archiveBox/abxbus
- `abxpkg` https://abxpkg.archivebox.io https://github.com/archiveBox/abxpkg
- `abx-plugins` https://plugins.archivebox.io https://github.com/ArchiveBox/abx-plugins
- `archivebox` https://archivebox.io https://github.com/ArchiveBox/ArchiveBox
- And lots more...
  - https://github.com/stars/pirate/lists/internet-archiving
  - https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community

---

For more advanced use with collections, parallel downloading, a Web UI + REST API, etc.
See: [`ArchiveBox/ArchiveBox`](https://github.com/ArchiveBox/ArchiveBox)

`abx-dl` runs one snapshot at a time. `--snapshot-max-size` stops starting more
hooks once reported output reaches the budget; cleanup still runs to preserve
recordings. Accounting lives only in memory for that run, so retrying an output
directory starts fresh and preserves partial artifacts until hooks overwrite them.
Crawl-wide URL, size, and time limits belong to ArchiveBox and its database.
