Skip to main content

Markdownee npm library

Use fetch() for one URL or createCrawler() to collect pages. Both return data to your Node.js application; the CLI provides shell commands.

Install

Use Node.js 22.22.2+ on 22.x, 24.15.0+ on 24.x, or 26+. In a new project directory:

npm init -y
npm install @markdownee/markdownee

Save the JavaScript examples as .mjs files and run them with node filename.mjs. They use HTTP-only Cheerio and need no browser. For browser-rendered pages, install Chromium with npx playwright install chromium, or Firefox with npx playwright install firefox, and choose the corresponding crawler type.

Fetch a page

Save as extract.mjs:

import { fetch } from "@markdownee/markdownee";

const { markdown } = await fetch(
  "https://en.wikipedia.org/wiki/Web_scraping",
  { crawlerType: "cheerio" },
);
console.log(markdown);
node extract.mjs

The program prints extracted Markdown. Fetch follows no links; request failure throws, and unavailable requested formats are omitted.

  • formats selects txt, markdown (default), html, minified-html, or original. The compact HTML result key is minifiedHtml; original is captured input.
  • crawlerType uses playwright-adaptive, playwright-firefox, playwright-chromium, or cheerio. Short browser aliases belong to the CLI.
  • imageHandling accepts exclude, alt-text, or resolved-url. Fetch rejects save; use CLI file output or a crawl for downloaded images.
  • Link, table, and user-comment handling accept include / exclude.
  • proxyConfiguration.proxyUrls accepts HTTP, HTTPS, SOCKS4, and SOCKS5 proxies. Apify Proxy configuration belongs to the Actor.
  • save, storageDir, includeHtml, crawl-frontier/concurrency controls, and persistent sessions are excluded from fetch. Request raw input through formats.

outputLayout applies to derived formats: minimal returns bodies/HTML fragments, standard adds ordinary metadata/front matter and complete HTML, and enhanced adds extended metadata and crawl information. It does not alter original HTML.

Collect a crawl

Save as crawl.mjs and run node crawl.mjs:

import { createCrawler } from "@markdownee/markdownee";

const crawler = createCrawler({
  crawlerType: "cheerio",
  save: ["markdown-dataset"],
  storageDir: "./guide-storage",
  selector: 'main a[href^="/wiki/"]',
  maxRequestsPerCrawl: 2,
  maxCrawlDepth: 1,
});
const { dataset, statistics, failures } = await crawler.run([
  "https://en.wikipedia.org/wiki/Web_scraping",
]);
await dataset.forEach((record) => {
  console.log(record.url, record.crawl?.depth);
});
console.log(statistics, failures.map((failure) => failure.url));

This prints successful URLs/depths and request statistics, and persists records in guide-storage. No selector means no link following. Request/result/depth limits are unrestricted when omitted; depth zero limits work to start URLs.

ResultDataset.export() returns the collected successes; forEach() visits them after collection, not as a live stream. failures holds exhausted requests and statistics counts request outcomes, not extracted-record or skipped-URL counts. Partial page failures preserve successes; invalid options and run-level errors throw.

Construction validates and snapshots camelCase options. Each run() owns a fresh queue, deduplication state, and logger. Without storageDir, records stay in memory; image save still persists image bytes in the CLI-resolved default storage. Calls sharing a directory share its stored data. includeHtml defaults to false and logLevel to warning. Bound maxResultsPerCrawl for large in-memory collections.

Optional markdownDiscovery uses off, alternate, negotiate, or probe. An accepted origin-published representation supplies all formats and adds markdownSource; failed discovery falls back to HTML. Link following uses available page HTML, so a Markdown-only page response has no HTML links to enqueue.

Export stored results

After the crawl above, save this as export.mjs and run node export.mjs:

import { runExportAction } from "@markdownee/markdownee/storage";

const result = await runExportAction({
  storageDir: "./guide-storage",
  outputDir: "./guide-export",
});
console.log(result.filesWritten, result.manifestPath);

Export writes available content files and guide-export/manifest.json. Its result also includes outputDir and recordsTotal. Filenames derive from title, URL, then page. The manifest records successful, failed, and skipped outcomes.

runPurgeAction({ storageDir }) permanently removes the selected datasets, key-value stores, and request queues without confirmation. It returns the resolved storageDir; it does not call process.exit().

Entry points and storage helpers

  • @markdownee/markdownee exports extraction APIs, types, and finite-value aliases such as SaveFormat, Save, and CrawlerType; aliases serialize as raw strings.
  • @markdownee/markdownee/schema exports MarkdowneeLibraryInput and MarkdowneeFetchInput validators and generated schema/presentation helpers.
  • @markdownee/markdownee/cli exports the import-safe Commander buildProgram().
  • @markdownee/markdownee/storage exports runExportAction, runPurgeAction, configureStorage, resolveStorageDir, Dataset, KeyValueStore, and Configuration.

Configure the storage directory before opening Crawlee stores. dataset.exportToJSON and exportToCSV require an explicitly supplied store; persistence does not remove returned results from memory. Library APIs have no public streaming or abort-signal contract. Navigation/request timeouts remain available through their options.

The CLI reference documents shared controls and storage resolution. Python Help covers that library's separate API and cancellation behavior.

Updated: September 25, 2026