Markdownee npm library
Use fetch() for one URL or createCrawler() to collect pages. Both return data
to your Node.js application; the CLI provides shell commands.
Install
Use Node.js 22.22.2+ on 22.x, 24.15.0+ on 24.x, or 26+. In a new project directory:
npm init -y
npm install @markdownee/markdownee
Save the JavaScript examples as .mjs files and run them with node filename.mjs.
They use HTTP-only Cheerio and need no browser. For browser-rendered pages, install
Chromium with npx playwright install chromium, or Firefox with
npx playwright install firefox, and choose the corresponding crawler type.
Fetch a page
Save as extract.mjs:
import { fetch } from "@markdownee/markdownee";
const { markdown } = await fetch(
"https://en.wikipedia.org/wiki/Web_scraping",
{ crawlerType: "cheerio" },
);
console.log(markdown);
node extract.mjs
The program prints extracted Markdown. Fetch follows no links; request failure throws, and unavailable requested formats are omitted.
formatsselectstxt,markdown(default),html,minified-html, ororiginal. The compact HTML result key isminifiedHtml; original is captured input.crawlerTypeusesplaywright-adaptive,playwright-firefox,playwright-chromium, orcheerio. Short browser aliases belong to the CLI.imageHandlingacceptsexclude,alt-text, orresolved-url. Fetch rejectssave; use CLI file output or a crawl for downloaded images.- Link, table, and user-comment handling accept
include/exclude. proxyConfiguration.proxyUrlsaccepts HTTP, HTTPS, SOCKS4, and SOCKS5 proxies. Apify Proxy configuration belongs to the Actor.save,storageDir,includeHtml, crawl-frontier/concurrency controls, and persistent sessions are excluded from fetch. Request raw input throughformats.
outputLayout applies to derived formats: minimal returns bodies/HTML fragments,
standard adds ordinary metadata/front matter and complete HTML, and enhanced
adds extended metadata and crawl information. It does not alter original HTML.
Collect a crawl
Save as crawl.mjs and run node crawl.mjs:
import { createCrawler } from "@markdownee/markdownee";
const crawler = createCrawler({
crawlerType: "cheerio",
save: ["markdown-dataset"],
storageDir: "./guide-storage",
selector: 'main a[href^="/wiki/"]',
maxRequestsPerCrawl: 2,
maxCrawlDepth: 1,
});
const { dataset, statistics, failures } = await crawler.run([
"https://en.wikipedia.org/wiki/Web_scraping",
]);
await dataset.forEach((record) => {
console.log(record.url, record.crawl?.depth);
});
console.log(statistics, failures.map((failure) => failure.url));
This prints successful URLs/depths and request statistics, and persists records in
guide-storage. No selector means no link following. Request/result/depth limits
are unrestricted when omitted; depth zero limits work to start URLs.
ResultDataset.export() returns the collected successes; forEach() visits them
after collection, not as a live stream. failures holds exhausted requests and
statistics counts request outcomes, not extracted-record or skipped-URL counts.
Partial page failures preserve successes; invalid options and run-level errors throw.
Construction validates and snapshots camelCase options. Each run() owns a fresh
queue, deduplication state, and logger. Without storageDir, records stay in memory;
image save still persists image bytes in the CLI-resolved default storage. Calls
sharing a directory share its stored data. includeHtml defaults to false and
logLevel to warning. Bound maxResultsPerCrawl for large in-memory collections.
Optional markdownDiscovery uses off, alternate, negotiate, or probe.
An accepted origin-published representation supplies all formats and adds
markdownSource; failed discovery falls back to HTML. Link following uses available
page HTML, so a Markdown-only page response has no HTML links to enqueue.
Export stored results
After the crawl above, save this as export.mjs and run node export.mjs:
import { runExportAction } from "@markdownee/markdownee/storage";
const result = await runExportAction({
storageDir: "./guide-storage",
outputDir: "./guide-export",
});
console.log(result.filesWritten, result.manifestPath);
Export writes available content files and guide-export/manifest.json. Its result
also includes outputDir and recordsTotal. Filenames derive from title, URL, then
page. The manifest records successful, failed, and skipped outcomes.
runPurgeAction({ storageDir }) permanently removes the selected datasets,
key-value stores, and request queues without confirmation. It returns the resolved
storageDir; it does not call process.exit().
Entry points and storage helpers
@markdownee/markdowneeexports extraction APIs, types, and finite-value aliases such asSaveFormat,Save, andCrawlerType; aliases serialize as raw strings.@markdownee/markdownee/schemaexportsMarkdowneeLibraryInputandMarkdowneeFetchInputvalidators and generated schema/presentation helpers.@markdownee/markdownee/cliexports the import-safe CommanderbuildProgram().@markdownee/markdownee/storageexportsrunExportAction,runPurgeAction,configureStorage,resolveStorageDir,Dataset,KeyValueStore, andConfiguration.
Configure the storage directory before opening Crawlee stores. dataset.exportToJSON
and exportToCSV require an explicitly supplied store; persistence does not remove
returned results from memory. Library APIs have no public streaming or abort-signal
contract. Navigation/request timeouts remain available through their options.
The CLI reference documents shared controls and storage resolution. Python Help covers that library's separate API and cancellation behavior.
Updated: September 25, 2026