Markdownee npm CLI
Use fetch for one page, crawl for stored collection, export for files, and
purge for storage cleanup. For application calls, use the npm library.
Install
Use Node.js 22.22.2+ on 22.x, 24.15.0+ on 24.x, or 26+. In a new project directory:
npm init -y
npm install @markdownee/markdownee
The commands below use HTTP-only Cheerio, which needs no browser. Adaptive and Chromium crawling require Chromium; Firefox requires its corresponding browser:
npx playwright install chromium
npx playwright install firefox
Install only the browser you select. Omitting --crawler-type cheerio uses adaptive
browser rendering; a nonzero --rendering-type-detection enables HTTP/browser sampling.
Fetch a page
npx markdownee fetch https://en.wikipedia.org/wiki/Web_scraping \
--crawler-type cheerio --save markdown-file -o page.md
Open page.md for the extracted article. Omit --save and -o to print Markdown to
stdout; diagnostics use stderr. Fetch follows no links and creates no Dataset or KVS.
Repeat --save with txt, markdown, html, minified-html, or original and a
file or stdout destination. Only one format can use stdout. -o is a literal
filename for one format, a prefix for several, or a directory when it ends in / or
already exists. Multiple HTML outputs use .html, .min.html, and .original.html.
Image saving requires file-only routes and at least one derived Markdown or HTML format. For example:
npx markdownee fetch https://en.wikipedia.org/wiki/Web_scraping \
--crawler-type cheerio --image-handling save \
--save markdown-file --save html-file -o out/page
This writes out/page.md and out/page.html, plus out/page.assets/ when images
survive extraction and download. Derived references become relative paths; captured
original HTML is unchanged. Text/original-only and stdout routes reject image saving.
Crawl, export, and remove storage
npx markdownee crawl https://en.wikipedia.org/wiki/Web_scraping \
--crawler-type cheerio --max-requests-per-crawl 1 \
--storage ./guide-storage --save markdown-kvs
npx markdownee export --storage ./guide-storage --output-dir ./guide-export
The crawl stores records and content; export writes available formats and
guide-export/manifest.json. Add --selector to follow links, with --globs /
--exclude and page/depth limits to bound the run. Globs alone discover no links.
--start-urls-file accepts a file containing one URL per line.
crawl retains successes when pages fail after retries. Exit codes are 0 for
success, 1 for a hard failure, and 2 for partial page/format failure. Export
includes every recorded outcome in its manifest. Its default output is
./markdownee-output.
To remove the storage created above:
npx markdownee purge --storage ./guide-storage
This permanently deletes its datasets, key-value stores, and request queues, without
confirmation or undo. crawl --purge performs the same cleanup before extraction.
Command reference
Storage
| Option | Description |
|---|---|
--storage <path> | Run storage root; default: ./storage or the XDG data directory. |
--purge | Delete the selected dataset, KVS, and request-queue buckets before crawling. |
--start-urls-file <path> | Read start URLs (one per line) from a file |
--config-file, -c <path> | Path to JSON config file |
For export, use --output-dir <path> (default ./markdownee-output) and --storage <path>.
Crawl settings
| Option | Description |
|---|---|
--crawler-type <type> | adaptive (default), firefox, chromium, or HTTP-only cheerio. |
--rendering-type-detection <ratio> | HTTP/browser sampling ratio 0–1 (adaptive only; default: 0). |
--markdown-discovery <mode> | Origin-published Markdown source: off (default), alternate, negotiate, or probe. |
--max-requests-per-crawl <n> | Request limit; omitted means unrestricted. |
--max-crawl-depth <n> | Link-depth limit; omitted means unrestricted, 0 means start URLs only. |
--max-results <n> | Max results per crawl; omit for no limit |
--initial-concurrency <n> | Starting parallel requests; omitted uses Crawlee's default. |
--max-concurrency <n> | Max parallel requests (default: 3) |
--max-retries <n> | Max request retries (default: 3) |
Markdown discovery supplies an alternate content source for all formats. The modes are cumulative: follow advertised same-origin links, also negotiate through Accept, then also try a .md sibling. Up to three alternates, a negotiated refetch, and a sibling can be attempted; robots.txt lookup can add an origin-level request. Crawler-path capabilities and per-origin budgets restrict these attempts. Rejected representations fall back to HTML. A Markdown-sourced record carries markdownSource; links come from available page HTML, which is absent when the page response itself is Markdown.
Crawl filtering
| Option | Description |
|---|---|
--globs <pattern> | Glob pattern to include (repeatable) |
--exclude <pattern> | Glob pattern to exclude (repeatable) |
--selector <css> | CSS selector for links to follow |
--keep-url-fragment | Preserve URL fragments |
--use-sitemaps | Enqueue URLs from the start origins' sitemap.xml. |
--respect-robots-txt | Honor robots.txt |
--deduplication <level> | minimal (URL), standard (plus canonical URL, default), or aggressive (plus content hash). |
Browser
| Option | Description |
|---|---|
--headless / --no-headless | Browser headless mode (default: headless) |
--wait-until <event> | Page load event: load, domcontentloaded, networkidle, commit |
--navigation-timeout <secs> | Navigation limit in seconds (default: 60). |
--wait-for-dynamic-content <secs> | Wait until network idle or the seconds limit; 0 disables the wait. |
--wait-for-selector <css> | CSS selector to wait for before extracting (fails on timeout) |
--soft-wait-for-selector <css> | CSS selector to wait for before extracting (continues on timeout) |
--block-media / --no-block-media | Block images, stylesheets, fonts, PDFs, and ZIPs (Chromium only; default: on). |
--ignore-cors-and-csp | Disable CORS/CSP restrictions |
--ignore-https-errors | Skip SSL certificate verification |
--close-cookie-modals / --no-close-cookie-modals | Attempt consent handling (default: on). |
--max-scroll-height <px> | Scroll limit in pixels (default: 5000; 0 disables scrolling). |
--user-agent <ua> | Custom User-Agent string |
Proxy & sessions
| Option | Description |
|---|---|
--proxy <url> | Proxy URL (repeatable) |
--proxy-rotation <strategy> | Rotation: recommended, per-request, until-failure |
--session-pool-name <name> | Named session pool for cross-run session sharing |
--max-session-rotations <n> | Session rotations per request after blocking (default: 10). |
Cookies & headers
| Option | Description |
|---|---|
--cookies <json> | JSON array of cookie objects |
--headers <json> | JSON object of custom HTTP headers |
Output & extraction
| Option | Description |
|---|---|
--save <token> | Repeat {txt,markdown,html,minified-html,original}-{dataset,kvs} (default: markdown-kvs). Repeat a format for both destinations; prefer KVS for large HTML bodies. |
--mode <mode> | precision (less noise), balanced (default), recall (more content), or whole-document cleanup with keep. |
--language <lang> | Filter by declared language (e.g. en); this is not statistical detection |
--image-handling <mode> | exclude (default), alt-text, resolved-url, or save. Crawl saves images to KVS; fetch saves sibling asset files. |
--max-image-edge <px> | Saved-image long-edge cap (default: 1568); 0 is uncapped; no upscaling. Ignored outside save mode. |
--rasterize-svg / --no-rasterize-svg | Saved SVG as PNG (default), or cleaned SVG source. Ignored outside save mode. |
--link-handling <mode> | include (default) or exclude (retain anchor text). |
--table-handling <mode> | include (default) or exclude (remove the subtree and its text). |
--comment-handling <mode> | include (default) or exclude detected user-comment sections, not HTML comment markup. |
--output-layout <layout> | Output envelope: minimal (default), standard, or enhanced |
--store-skipped-urls | Push skipped URL records to the dataset after crawl |
--verbose, -v | Enable verbose logging |
Load JSON configuration
Save the following as config.json, then run the command below it:
{
"startUrls": [{ "url": "https://en.wikipedia.org/wiki/Web_scraping" }],
"crawlerType": "cheerio",
"maxRequestsPerCrawl": 1,
"mode": "recall",
"outputLayout": "standard",
"save": ["txt-dataset"]
}
npx markdownee crawl --config-file config.json --storage ./config-storage
Configuration uses the CLI's camelCase schema, not YAML. Defaults apply first, JSON
values next, and explicit CLI flags last; explicit --proxy flags replace configured
proxy URLs. JSON URL and glob entries are { "url": "..." } and { "glob": "..." }
objects, while CLI flags take strings.
Unknown or Actor-only fields (datasetName, keyValueStoreName, requestQueueName,
Apify Proxy controls) fail validation. Ordinary proxies can use
proxyConfiguration.proxyUrls. Keep storage/purge orchestration on the command line.
outputLayout is independent of formats: minimal gives body text/HTML fragments,
standard adds ordinary metadata/front matter and complete HTML, and enhanced
adds extended metadata and crawl context. Original remains the captured input;
Markdown discovery can make that capture a derivative of the served representation.
Resolve the storage path
The first available value wins: --storage, MARKDOWNEE_STORAGE_DIR,
CRAWLEE_STORAGE_DIR, ./storage when .actor/ or ./storage/ exists, then
${XDG_DATA_HOME:-~/.local/share}/markdownee/storage.
Fetch excludes crawl/storage controls such as --globs, --max-crawl-depth,
--storage, and persistent session settings. Use each subcommand's --help for
its accepted flags. Continue with the npm library,
Apify Actor, or Python library.
Updated: September 25, 2026