Skip to main content

Markdownee npm CLI

Use fetch for one page, crawl for stored collection, export for files, and purge for storage cleanup. For application calls, use the npm library.

Install

Use Node.js 22.22.2+ on 22.x, 24.15.0+ on 24.x, or 26+. In a new project directory:

npm init -y
npm install @markdownee/markdownee

The commands below use HTTP-only Cheerio, which needs no browser. Adaptive and Chromium crawling require Chromium; Firefox requires its corresponding browser:

npx playwright install chromium
npx playwright install firefox

Install only the browser you select. Omitting --crawler-type cheerio uses adaptive browser rendering; a nonzero --rendering-type-detection enables HTTP/browser sampling.

Fetch a page

npx markdownee fetch https://en.wikipedia.org/wiki/Web_scraping \
  --crawler-type cheerio --save markdown-file -o page.md

Open page.md for the extracted article. Omit --save and -o to print Markdown to stdout; diagnostics use stderr. Fetch follows no links and creates no Dataset or KVS.

Repeat --save with txt, markdown, html, minified-html, or original and a file or stdout destination. Only one format can use stdout. -o is a literal filename for one format, a prefix for several, or a directory when it ends in / or already exists. Multiple HTML outputs use .html, .min.html, and .original.html.

Image saving requires file-only routes and at least one derived Markdown or HTML format. For example:

npx markdownee fetch https://en.wikipedia.org/wiki/Web_scraping \
  --crawler-type cheerio --image-handling save \
  --save markdown-file --save html-file -o out/page

This writes out/page.md and out/page.html, plus out/page.assets/ when images survive extraction and download. Derived references become relative paths; captured original HTML is unchanged. Text/original-only and stdout routes reject image saving.

Crawl, export, and remove storage

npx markdownee crawl https://en.wikipedia.org/wiki/Web_scraping \
  --crawler-type cheerio --max-requests-per-crawl 1 \
  --storage ./guide-storage --save markdown-kvs
npx markdownee export --storage ./guide-storage --output-dir ./guide-export

The crawl stores records and content; export writes available formats and guide-export/manifest.json. Add --selector to follow links, with --globs / --exclude and page/depth limits to bound the run. Globs alone discover no links. --start-urls-file accepts a file containing one URL per line.

crawl retains successes when pages fail after retries. Exit codes are 0 for success, 1 for a hard failure, and 2 for partial page/format failure. Export includes every recorded outcome in its manifest. Its default output is ./markdownee-output.

To remove the storage created above:

npx markdownee purge --storage ./guide-storage

This permanently deletes its datasets, key-value stores, and request queues, without confirmation or undo. crawl --purge performs the same cleanup before extraction.

Command reference

Storage

OptionDescription
--storage <path>Run storage root; default: ./storage or the XDG data directory.
--purgeDelete the selected dataset, KVS, and request-queue buckets before crawling.
--start-urls-file <path>Read start URLs (one per line) from a file
--config-file, -c <path>Path to JSON config file

For export, use --output-dir <path> (default ./markdownee-output) and --storage <path>.

Crawl settings

OptionDescription
--crawler-type <type>adaptive (default), firefox, chromium, or HTTP-only cheerio.
--rendering-type-detection <ratio>HTTP/browser sampling ratio 0–1 (adaptive only; default: 0).
--markdown-discovery <mode>Origin-published Markdown source: off (default), alternate, negotiate, or probe.
--max-requests-per-crawl <n>Request limit; omitted means unrestricted.
--max-crawl-depth <n>Link-depth limit; omitted means unrestricted, 0 means start URLs only.
--max-results <n>Max results per crawl; omit for no limit
--initial-concurrency <n>Starting parallel requests; omitted uses Crawlee's default.
--max-concurrency <n>Max parallel requests (default: 3)
--max-retries <n>Max request retries (default: 3)

Markdown discovery supplies an alternate content source for all formats. The modes are cumulative: follow advertised same-origin links, also negotiate through Accept, then also try a .md sibling. Up to three alternates, a negotiated refetch, and a sibling can be attempted; robots.txt lookup can add an origin-level request. Crawler-path capabilities and per-origin budgets restrict these attempts. Rejected representations fall back to HTML. A Markdown-sourced record carries markdownSource; links come from available page HTML, which is absent when the page response itself is Markdown.

Crawl filtering

OptionDescription
--globs <pattern>Glob pattern to include (repeatable)
--exclude <pattern>Glob pattern to exclude (repeatable)
--selector <css>CSS selector for links to follow
--keep-url-fragmentPreserve URL fragments
--use-sitemapsEnqueue URLs from the start origins' sitemap.xml.
--respect-robots-txtHonor robots.txt
--deduplication <level>minimal (URL), standard (plus canonical URL, default), or aggressive (plus content hash).

Browser

OptionDescription
--headless / --no-headlessBrowser headless mode (default: headless)
--wait-until <event>Page load event: load, domcontentloaded, networkidle, commit
--navigation-timeout <secs>Navigation limit in seconds (default: 60).
--wait-for-dynamic-content <secs>Wait until network idle or the seconds limit; 0 disables the wait.
--wait-for-selector <css>CSS selector to wait for before extracting (fails on timeout)
--soft-wait-for-selector <css>CSS selector to wait for before extracting (continues on timeout)
--block-media / --no-block-mediaBlock images, stylesheets, fonts, PDFs, and ZIPs (Chromium only; default: on).
--ignore-cors-and-cspDisable CORS/CSP restrictions
--ignore-https-errorsSkip SSL certificate verification
--close-cookie-modals / --no-close-cookie-modalsAttempt consent handling (default: on).
--max-scroll-height <px>Scroll limit in pixels (default: 5000; 0 disables scrolling).
--user-agent <ua>Custom User-Agent string

Proxy & sessions

OptionDescription
--proxy <url>Proxy URL (repeatable)
--proxy-rotation <strategy>Rotation: recommended, per-request, until-failure
--session-pool-name <name>Named session pool for cross-run session sharing
--max-session-rotations <n>Session rotations per request after blocking (default: 10).

Cookies & headers

OptionDescription
--cookies <json>JSON array of cookie objects
--headers <json>JSON object of custom HTTP headers

Output & extraction

OptionDescription
--save <token>Repeat {txt,markdown,html,minified-html,original}-{dataset,kvs} (default: markdown-kvs). Repeat a format for both destinations; prefer KVS for large HTML bodies.
--mode <mode>precision (less noise), balanced (default), recall (more content), or whole-document cleanup with keep.
--language <lang>Filter by declared language (e.g. en); this is not statistical detection
--image-handling <mode>exclude (default), alt-text, resolved-url, or save. Crawl saves images to KVS; fetch saves sibling asset files.
--max-image-edge <px>Saved-image long-edge cap (default: 1568); 0 is uncapped; no upscaling. Ignored outside save mode.
--rasterize-svg / --no-rasterize-svgSaved SVG as PNG (default), or cleaned SVG source. Ignored outside save mode.
--link-handling <mode>include (default) or exclude (retain anchor text).
--table-handling <mode>include (default) or exclude (remove the subtree and its text).
--comment-handling <mode>include (default) or exclude detected user-comment sections, not HTML comment markup.
--output-layout <layout>Output envelope: minimal (default), standard, or enhanced
--store-skipped-urlsPush skipped URL records to the dataset after crawl
--verbose, -vEnable verbose logging

Load JSON configuration

Save the following as config.json, then run the command below it:

{
  "startUrls": [{ "url": "https://en.wikipedia.org/wiki/Web_scraping" }],
  "crawlerType": "cheerio",
  "maxRequestsPerCrawl": 1,
  "mode": "recall",
  "outputLayout": "standard",
  "save": ["txt-dataset"]
}
npx markdownee crawl --config-file config.json --storage ./config-storage

Configuration uses the CLI's camelCase schema, not YAML. Defaults apply first, JSON values next, and explicit CLI flags last; explicit --proxy flags replace configured proxy URLs. JSON URL and glob entries are { "url": "..." } and { "glob": "..." } objects, while CLI flags take strings.

Unknown or Actor-only fields (datasetName, keyValueStoreName, requestQueueName, Apify Proxy controls) fail validation. Ordinary proxies can use proxyConfiguration.proxyUrls. Keep storage/purge orchestration on the command line.

outputLayout is independent of formats: minimal gives body text/HTML fragments, standard adds ordinary metadata/front matter and complete HTML, and enhanced adds extended metadata and crawl context. Original remains the captured input; Markdown discovery can make that capture a derivative of the served representation.

Resolve the storage path

The first available value wins: --storage, MARKDOWNEE_STORAGE_DIR, CRAWLEE_STORAGE_DIR, ./storage when .actor/ or ./storage/ exists, then ${XDG_DATA_HOME:-~/.local/share}/markdownee/storage.

Fetch excludes crawl/storage controls such as --globs, --max-crawl-depth, --storage, and persistent session settings. Use each subcommand's --help for its accepted flags. Continue with the npm library, Apify Actor, or Python library.

Updated: September 25, 2026