AutomationMart
Home/Browse/Scrape, search and browse the web with a Firecrawl AI agent webhook
n8n

Scrape, search and browse the web with a Firecrawl AI agent webhook

n8nn8n14 modulesv1.0
agentcodefirecrawlToollmChatOpenRouteroutputParserStructuredrespondToWebhookwebhook

Turn any prompt into structured web data. Send a POST request with a natural language prompt and an optional JSON schema, and get back clean, structured results scraped from the web by an AI agent powered by Firecrawl. Use Cases - Data Enrichment: Feed company names or URLs from your CRM and get back structured firmographic data (industry, funding, team size, tech stack). - Lead Generation: Ask the agent to find pricing, contact pages, or product details for a list of competitors. -

At a glance

Scrape, search and browse the web with a Firecrawl AI agent webhook is a ready-made n8n workflow you import as a workflow JSON file — no build required. It connects agent, code, firecrawlTool, lmChatOpenRouter. It's free to download. Follow the 5-step import below to go live in minutes.

Platform
n8n
Connects
agent, code, firecrawlTool, lmChatOpenRouter, outputParserStructured, respondToWebhook
Modules
14
Price
Free
Version
v1.0
Scrape, search and browse the web with a Firecrawl AI agent webhook workflow diagram

About this workflow

Turn any prompt into structured web data. Send a POST request with a natural language prompt and an optional JSON schema, and get back clean, structured results scraped from the web by an AI agent powered by Firecrawl. Use Cases - Data Enrichment: Feed company names or URLs from your CRM and get back structured firmographic data (industry, funding, team size, tech stack). - Lead Generation: Ask the agent to find pricing, contact pages, or product details for a list of competitors. - Market Research: Extract structured pricing plans, feature comparisons, or product catalogs from any website. - Content Aggregation: Pull structured news, events, or job postings from across the web on a schedule. - Sales Intelligence: Enrich prospect lists with company info, recent news, or tech stack details before outreach. How It Works 1. Receive Scrape Request receives a POST request with prompt and an optional outputschema. 2. Validate Output Schema checks the schema. If none is provided, it falls back to a permissive default. If the schema is malformed, it returns a clear error via Return Schema Error. 3. Research & Extract Web Data takes the prompt and uses the full Firecrawl toolkit to research the web: - Search (/search): Finds relevant pages and sources across the web. - Scrape (/scrape): Extracts clean, structured content from any URL. - Interact (interactContext, interact, interactStop): Lets the agent interact with scraped pages in a live session. After scraping a page, the agent can click buttons, fill forms, navigate dynamic content, and extract data that static scraping cannot reach, all without managing sessions manually. This combination gives the AI agent complete web navigation capabilities. It can discover sources, read pages, and interact with dynamic content autonomously. 4. Format Response to Schema (Structured Output Parser) formats the agent's response to match the provided (or default) schema. 5. Return Structured Results sends the structured JSON back to the caller. Setup Requirements - Firecrawl API Key: Sign up at firecrawl.dev and grab your API key. Connect it in the Firecrawl credential nodes. - LLM Provider: Configure your Primary Chat Model and Fallback Chat Model nodes (e.g., OpenRouter, OpenAI, Anthropic). The template uses two model nodes for reliability, plus a separate Parser Chat Model for the output parser. - n8n Instance: Self-hosted or cloud. Make sure the webhook node is set to accept POST requests. API Reference Endpoint Request Body | Field | Type | Required | Description | |-------|------|----------|-------------| | prompt | string | Yes | Natural language instruction for the agent | | outputschema | object | No | JSON Schema defining the desired output structure | Response Returns a JSON object matching the provided schema, or a flexible object if no schema was given. --- Testing Examples 1. Basic Request (No Schema) The agent decides the output structure on its own. Expected output: A JSON object with whatever structure the agent finds most appropriate for the data. Since no schema was provided, the internal default ({ "type": "object", "additionalProperties": true }) is used. 2. Request With a Custom Schema You define exactly the shape of data you want back. Expected output: 3. Invalid Schema (String Instead of Object) Expected output: 4. Invalid Schema (Array Instead of Object) Expected output: Same error response as above. 5. Invalid Schema (Missing type Property) Expected output: Same error response as above. 6. Invalid Schema (Invalid type Value) Expected output: Same error response as above. --- Workflow Architecture Schema Validation Logic The Validate Output Schema node runs this validation before passing data to the agent: - If outputschema is missing or null, the default permissive schema is used: { "type": "object", "additionalProperties": true }. - If outputschema is present, it must be a JSON object (not a string, array, or primitive). - It must have a type property with a valid value: object, array, string, number, or boolean. - If validation fails, the workflow returns an error response with a helpful message and example schema. Notes - The Format Response to Schema node (Structured Output Parser) requires the schema to be passed as a JSON string. The expression {{ JSON.stringify($('Validate Output Schema').item.json.outputschema) }} handles this conversion. - The agent has access to Firecrawl's full toolkit: search, scrape, and interact. With all three connected, the agent has complete web navigation powers. It can discover sources via search, extract content via scrape, and interact with dynamic JavaScript-heavy pages via interact. The interact tools let the agent scrape a page first and then continue working with it in a live session, clicking buttons, filling forms, and navigating deeper, all without manual session management. The agent autonomously decides which tools to use based on the prompt. - Response times vary depending on the complexity of the prompt and how many pages the agent needs to visit. Simple lookups take a few seconds; deep research can take longer.

n8n

How to import this n8n workflow

  1. 1

    Download the workflow JSON file after purchase.

  2. 2

    Open n8n → click the menu → Import from File.

  3. 3

    Select the downloaded JSON and import.

  4. 4

    Set up credentials for each node that requires them.

  5. 5

    Click Execute Workflow to test, then activate.

Setup guide

Setup guide included

Purchase to unlock the full step-by-step guide

Related N8n workflows

Extract and organize Colombian invoices with Gmail, GPT-4o & Google Workspace

Personal Invoice Processor This N8N workflow automates the extraction and organization of personal invoices in Colombia received via Gmail. It includes the following key steps: 🔁 Flow Summary 1. Email Trigger - Polls Gmail every 30 minutes for emails with .zip attachments (assumed to contain invoices). - Expects ZIP file following DIAN standards. 2. ZIP File Handling - Extracts all files. - Filters only PDF and XML files for processing. 3. Data Extraction & Processing - Uses LangChain Agent

Free

5 ways to process images & PDFs with Gemini AI in n8n

How it works Many users have asked in the support forum about different methods to analyze images and PDF documents with Google Gemini AI in n8n. This workflow answers that question by demonstrating five different approaches: - Single image with auto binary passthrough - The simplest approach using AI Agent's automatic binary handling - Multiple images with predefined prompts - For customized analysis with different instructions per image - Native n8n item-by-item processing - For handling

Free

Automate Morning Brew–style Reddit Digests and Publish to DEV using AI

This workflow contains community nodes that are only compatible with the self-hosted version of n8n. Who’s it for Community managers, content marketers, and builders who want a daily, skimmable update from a subreddit—automatically summarized, formatted, and cross-posted to DEV Community. Here is a Link to video hackathon detailing this build. What it does Collects fresh posts from a subreddit (seeded via RSS). Uses the Bright Data node to batch-scrape each post for richer fields (upvotes, co

Free

Monitor marketing job boards with Bright Data & GPT-4o for growing companies

This workflow automatically monitors marketing job boards to identify growing companies and potential business opportunities. It saves you time by eliminating the need to manually check job listings and provides insights into which companies are actively hiring and expanding their marketing teams. Overview This workflow automatically scrapes marketing job listings from Indeed and other job boards to extract company information, job details, and growth indicators. It uses Bright Data to access jo

Free

E-commerce product fine-tuning with Bright Data and OpenAI

This workflow contains community nodes that are only compatible with the self-hosted version of n8n. This workflow automates the process of scraping product data from e-commerce websites and using it to fine-tune a custom OpenAI GPT model for generating high-quality marketing copy and product descriptions. Main Use Cases Fine-tune OpenAI models with real product data from hundreds of supported e-commerce websites for marketing content generation. Create custom AI models specialized in writing

Free

Scrape detailed GitHub profiles to Google Sheets using BrowserAct

Scrape Detailed GitHub Profiles to Google Sheets Using BrowserAct This template is a sophisticated data enrichment and reporting tool that scrapes detailed GitHub user profiles and organizes the information into dedicated, structured reports within a Google Sheet. This workflow is essential for technical recruiters, talent acquisition teams, and business intelligence analysts who need to dive deep into a pre-qualified list of developers to understand their recent activity, repositories, and tech

Free

Track domain expiry dates with Google Sheets and WHOIS API

Automatically track domain expiry dates from Google Sheets, fetch real-time DNS expiry data via WHOIS API, and update expiry details back to your sheet with zero manual effort. --- Automated Domain Expiry Date Tracker with Google Sheets & WHOIS API Automate the entire process of monitoring domain expiry dates for all your websites directly from Google Sheets. This workflow reads domain names, fetches DNS SOA expiry information using the WHOIS API, converts timestamps into readable dates, and upd

Free

Generate text, image, and video-to-video clips with WAN 2.6 via KIE.AI

This n8n template provides a comprehensive suite of WAN 2.6 video generation capabilities through the KIE.AI API. The workflow includes three independent video generation workflows: text-to-video, image-to-video, and video-to-video. Each workflow can be used independently to create videos from different input types, making it perfect for content creators, marketers, and video production teams. Use cases are many: Create videos from text descriptions without any input media, transform static imag

Free

Reviews

No reviews yet

Be the first to buy and share your experience.

Leave a review

Sign in to share your experience with this workflow.

Log in to review
Free
No ratings yet

Create a free account to purchase workflows.

  • JSON blueprint — instant download
  • Setup guide PDF included
  • 5 downloads · valid 30 days
  • Works with n8n

Need help setting this up?

Book a 3-hour live setup session with an Agility consultant.

₹2,499/ session
3 hrs · video call
  • Configure live on Google Meet / Zoom
  • Free follow-up if workflow has defects
  • Platform expert assigned to you
Book installation session
Free