AutomationMart
Home/Browse/Build a company website RAG chatbot using Apify, Pinecone and Gemini
n8n

Build a company website RAG chatbot using Apify, Pinecone and Gemini

n8nn8n14 modulesv1.0
Gemini

AI chatbots are only as good as the data they learn from. Most large language models (LLM) rely only on their training datasets. If you want the chatbots to know more about your business, the best is to implement a retrieval-augmented generation (RAG) pipeline to train Gemini with your website data. This is what this workflow will help you to do. This workflow uses a scheduler to scrape a website on a regular basis using Apify; web pages are then indexed or updated in a Pinecone vector database.

At a glance

Build a company website RAG chatbot using Apify, Pinecone and Gemini is a ready-made n8n workflow you import as a workflow JSON file — no build required. It connects Gemini. It's free to download. Follow the 5-step import below to go live in minutes.

Platform
n8n
Connects
Gemini
Modules
14
Price
Free
Version
v1.0
Build a company website RAG chatbot using Apify, Pinecone and Gemini workflow diagram

About this workflow

AI chatbots are only as good as the data they learn from. Most large language models (LLM) rely only on their training datasets. If you want the chatbots to know more about your business, the best is to implement a retrieval-augmented generation (RAG) pipeline to train Gemini with your website data. This is what this workflow will help you to do. This workflow uses a scheduler to scrape a website on a regular basis using Apify; web pages are then indexed or updated in a Pinecone vector database. This allows the chatbot to provide accurate and up-to-date information. The workflow uses Google's Gemini AI for both embeddings and response generation. How does it work? This workflow is split into 2 sub-logics highlighted with green sticky notes: RAG Training logic Chatbot logic RAG training logic 1. Use the Apify Website Content Crawler to retrieve all content from your website 2. The Pinecone Vector Store node indexes the text chunk in a Pinecone index. 3. The Embeddings Google Gemini node generates embeddings for each text chunk Chatbot logic 1. The Chat Trigger node receives user questions through a chat interface. An AI Agent node handles those requests. 1. The AI Agent node uses a Vector Store Tool node, linked to a Pinecone Vector Store node in query mode, to retrieve relevant text chunks from Pinecone based on the user's question. 2. The AI Agent sends the retrieved information and the user's question to the Google Gemini Chat Model (gemini-pro). How to set up this template? All nodes with an orange sticky note require setup. Get your tools set up: 1 Google Cloud Project and Vertex AI API: Create a Google Cloud project. Enable the Vertex AI API for your project. Obtain a Google AI API key from Google AI Studio 2 Get an Apify account Create an Apify account 3 Pinecone Account: Create a free account on the Pinecone website. Obtain your API key from your Pinecone dashboard. Create an index named company-website in your Pinecone project. Configure credentials in your n8n environment for: Google Gemini(PaLM) Api (using your Google AI API key) Pinecone API (using your Pinecone API key) Setup trigger frequency: Edit the Schedule Trigger to match the frequency at which you wish to update your RAG If you want to train your chatbot only once, you can replace it with a click trigger. Set up the Apify node Authenticate (via OAuth or API) Set up your website URL in the JSON input FAQ What is RAG? RAG stands for retrieval-augmented generation. It is a technique that provides an AI model (such as a large language model) with additional data. That allows the LLM to give more up-to-date and topic-specific information. What is the difference between RAG and LLM? RAG is a way to complement an LLM by giving it more up-to-date information. You can think of the LLM as the CPU processing your question, and RAG as the hard drive providing information. Do I have to use my website as training data? No. Website Content Crawler can scrape any website. So you can, in theory, use this template to build a RAG for someone else. You can even combine data from multiple websites. Can I use another model other than Gemini? In theory, yes. You could replace the Gemini node with another LLM model. If you are looking for inspiration about RAG implementation with the Ollama model, check out this template.

n8n

How to import this n8n workflow

  1. 1

    Download the workflow JSON file after purchase.

  2. 2

    Open n8n → click the menu → Import from File.

  3. 3

    Select the downloaded JSON and import.

  4. 4

    Set up credentials for each node that requires them.

  5. 5

    Click Execute Workflow to test, then activate.

Setup guide

Setup guide included

Purchase to unlock the full step-by-step guide

Related N8n workflows

Extract structured invoice data from JotForm PDFs with GPT-4.1-mini & Sheets

Who this is for This workflow is designed for Finance teams, accounting professionals, and automation engineers. Use Case: Automates processing of invoice submissions received via JotForm. Core Function: Extracts structured data such as: Invoice number Client information Totals and tax amounts Line items or services Key Benefit: Eliminates manual data entry, saving time and reducing human error. Automation Goal: Streamline document handling with AI-powered PDF parsing and structured o

Free

Financial news digest with Google Gemini AI to Outlook email

Key Features - Looping source scraping: Collects content from news sites you have selected (it might not work for all of them however) - HTML extraction & cleaning: Parses, cleans, and filters messy website data to isolate only the most relevant content. - AI-powered synthesis: Uses Google Gemini (via LangChain agent) to summarize and structure financial news into a clear, bullet-pointed format. - Email-ready output: Generates styled HTML summaries with coral-colored headings, ideal for daily em

Free

Create dynamic crypto market monitors with Gemini AI and Telegram Bot

Template Description This description details the template's purpose, how it works, and its key features. You can copy and use it directly. Overview This is a powerful n8n "meta-workflow" that acts as a Supervisor. Through a simple Telegram bot, you can dynamically create, manage, and delete countless independent, AI-driven market monitoring agents (Watchdogs). This template is a perfect implementation of the "Workflowception" (workflow managing workflows) concept in n8n, showcasing how to achie

Free

Collect LinkedIn details and generate CV feedback with Gemini and Google Workspace

This workflow helps HR teams, career coaches, and training programs collect candidate data and automatically generate CV improvement recommendations and a cover letter draft. Candidates submit their LinkedIn profile URL, contact details, and an optional CV PDF using an n8n Form. The workflow logs submissions, processes the uploaded CV, and generates structured outputs in Google Docs using Gemini, then sends the candidate an email with the results. How it works Collect candidate data An

Free

Create structured Notion workspaces from notes & voice using Gemini & GPT

AI Assistant Workflow: Create Notion Workspaces from Notes & Voice Records 👤 Who is this for? This workflow is designed for anyone who loves Notion—from project managers, freelancers, to students—who want to turn scattered ideas, handwritten notes, or quick thoughts into fully structured Notion databases without the hassle of manual setup. 😩 The Problem You have a brilliant idea jotted down during a meeting or on a piece of paper. But turning that into a structured Notion workspace (for project

Free

Extract, Transform LinkedIn Data with Bright Data MCP Server & Google Gemini

Disclaimer This template is only available on n8n self-hosted as it's making use of the community node for MCP Client. Who this is for? The Extract, Transform LinkedIn Data with Bright Data MCP Server & Google Gemini workflow is an automated solution that scrapes LinkedIn content via Bright Data MCP Server then transforms the response using a Gemini LLM. The final output is sent via webhook notification and also persisted on disk. This workflow is tailored for:​ 1. Data Analysts : Who require st

Free

Daily crypto yield monitor: Track top Binance Earn APY rates with earnings calculator

Get top Binance Earn yields sent to Email This workflow automates the tracking of passive income opportunities on Binance by fetching real-time Flexible Earn APY rates, calculating potential returns, and delivering a daily summary to your inbox. Manually checking crypto savings rates is tedious. This template handles the complex authentication required by Binance (HMAC-SHA256 signing), filters through hundreds of assets to find the highest yields, and calculates exactly how much you would ea

Free

Generate SEO seed keywords using AI

What this workflow does: This flow uses an AI node to generate Seed Keywords to focus SEO efforts on based on your ideal customer profile. You can use these keywords to form part of your SEO strategy. Outputs: - List of 20 Seed Keywords Setup 1. Fill the Set Ideal Customer Profile (ICP) 2. Connect with your credentials 3. Replace the Connect to your own database with your own database Pre-requisites / Dependencies - You know your ideal customer profile (ICP) - An AI API account (either OpenAI or

Free

Reviews

No reviews yet

Be the first to buy and share your experience.

Leave a review

Sign in to share your experience with this workflow.

Log in to review
Free
No ratings yet

Create a free account to purchase workflows.

  • JSON blueprint — instant download
  • Setup guide PDF included
  • 5 downloads · valid 30 days
  • Works with n8n

Need help setting this up?

Book a 3-hour live setup session with an Agility consultant.

₹2,499/ session
3 hrs · video call
  • Configure live on Google Meet / Zoom
  • Free follow-up if workflow has defects
  • Platform expert assigned to you
Book installation session
Free