Skip to Content
ResourcesIntegrationsDeveloper ToolsFirecrawl

Firecrawl

Service domainWEB SCRAPING
Firecrawl icon
Arcade OptimizedBYOCPro

Arcade.dev LLM tools for reading the web via Firecrawl

Author:Arcade
Version:4.0.1
Auth:No authentication required
12tools
12require secrets

Firecrawl toolkit lets Arcade agents read and extract content from the web via the Firecrawl API — covering single pages, full-site crawls, structured extraction, and targeted search across the open web, developer docs, GitHub issues, and research paper corpora.

Capabilities

  • Single-page and multi-page reading: Scrape a known URL or crawl an entire site, retrieving page content with job tracking for long-running crawls.
  • Site mapping: List all URLs on a site without fetching page content.
  • Structured data extraction: Describe desired fields in plain language (e.g. products, prices, contacts) and receive structured records; supports async job tracking.
  • Web search with optional content fetching: Search the open web and optionally retrieve full content of each result in one call.
  • Specialized search indexes: Query indexed developer documentation/repositories, GitHub issues only, or a scientific paper corpus (PubMed, bioRxiv, medRxiv, arXiv) with source-appropriate ranking and quoting.
  • Crawl lifecycle management: Check crawl progress, read pages collected so far, or cancel an active crawl by job ID.

Secrets

FIRECRAWL_API_KEY — Your Firecrawl API key, used to authenticate every request to the Firecrawl API. Obtain it by signing in to the Firecrawl dashboard and creating or copying an API key. Free-tier keys are available but have rate and credit limits; higher-volume usage requires a paid plan. Set this value as a secret in Arcade before invoking any tool in this toolkit.

See Arcade secrets docs for how to store secrets, or manage them directly at https://api.arcade.dev/dashboard/auth/secrets.

Available tools(12)

12 of 12 tools
Operations
Behavior
Tool nameDescriptionSecrets
Stop a crawl that is still running. Pages the crawl already collected stay readable by its job id.
1
Read many pages of one website and return each page's content. A crawl that outruns wait_seconds keeps running: the result carries its job_id for checking progress, reading the pages, or cancelling it later.
1
Collect structured records from the web from a plain-language description. Use this when the answer is fields such as products, prices, or contacts, rather than the page text itself. An extraction that outruns wait_seconds keeps running, and the result carries its job_id for collecting the results later.
1
Read the pages a crawl has collected so far.
1
Check how far a crawl has progressed, without fetching its pages.
1
Collect the results of an extraction that is already running.
1
List the URLs on a website, without reading the pages.
1
Read one web page whose URL is already known and return its content. It reads only the page at the given URL. It does not search the web, crawl other pages of the site, or pull structured fields out of the page text.
1
Search the web and optionally read each result page in the same call. Use this when no URL is known yet. It searches the open web rather than a corpus of scientific papers or indexed code documentation.
1
Search developer documentation and repositories, quoting the text that matched. Use this to answer a question about how a library or API works. It searches indexed repositories, not the open web.
1
Search GitHub issues across indexed repositories. Only issues are searched; documentation, READMEs, and pull requests are not.
1
Search published scientific papers and return their abstracts. This reads a corpus of paper records drawn from PubMed, bioRxiv, medRxiv, and arXiv. To find ordinary web pages that happen to sit on academic sites, use Search narrowed to research instead.
1
Last updated on