Compare GPT-4, Claude & Gemini Responses with Contextual AI's LMUnit Evaluation
PROBLEM Evaluating and comparing responses from multiple LLMs (OpenAI, Claude, Gemini) can be challenging when done manually. - Each model produces outputs that differ in clarity, tone, and reasoning structure. - Traditional evaluation metrics like ROUGE or BLEU fail to capture nuanced quality differences. - Human evaluations are inconsistent, slow, and difficult to scale. This workflow automates LLM response quality evaluation using Contextual AI’s LMUnit, a natural la
At a glance
Compare GPT-4, Claude & Gemini Responses with Contextual AI's LMUnit Evaluation is a ready-made n8n workflow you import as a workflow JSON file — no build required. It connects OpenAI, Gemini. It's free to download. Follow the 5-step import below to go live in minutes.
- Platform
- n8n
- Connects
- OpenAI, Gemini
- Modules
- 16
- Price
- Free
- Version
- v1.0

About this workflow
PROBLEM Evaluating and comparing responses from multiple LLMs (OpenAI, Claude, Gemini) can be challenging when done manually. - Each model produces outputs that differ in clarity, tone, and reasoning structure. - Traditional evaluation metrics like ROUGE or BLEU fail to capture nuanced quality differences. - Human evaluations are inconsistent, slow, and difficult to scale. This workflow automates LLM response quality evaluation using Contextual AI’s LMUnit, a natural language unit testing framework that provides systematic, fine-grained feedback on response clarity and conciseness. > Note: LMUnit offers natural language-based evaluation with a 1–5 scoring scale, enabling consistent and interpretable results across different model outputs. How it works - A chat trigger node collects responses from multiple LLMs such as OpenAI GPT-4.1, Claude 4.5 Sonnet, and Gemini 2.5 Flash. - Each model receives the same input prompt to ensure fair comparison, which is then aggregated and associated with each test cases - We use Contextual AI's LMUnit node to evaluate each response using predefined quality criteria: - “Is the response clear and easy to understand?” - Clarity - “Is the response concise and free from redundancy?” - Conciseness - LMUnit then produces evaluation scores (1–5) for each test - Results are aggregated and formatted into a structured summary showing model-wise performance and overall averages. How to set up - Create a free Contextual AI account and obtain your CONTEXTUALAIAPIKEY. - In your n8n instance, add this key as a credential under “Contextual AI.” - Obtain and add credentials for each model provider you wish to test: - OpenAI API Key: platform.openai.com/account/api-keys - Anthropic API Key: console.anthropic.com/settings/keys - Gemini API Key: ai.google.dev/gemini-api/docs/api-key - Start sending prompts using chat interface to automatically generate model outputs and evaluations. How to customize the workflow - Add more evaluation criteria (e.g., factual accuracy, tone, completeness) in the LMUnit test configuration. - Include additional LLM providers by duplicating the response generation nodes. - Adjust thresholds and aggregation logic to suit your evaluation goals. - Enhance the final summary formatting for dashboards, tables, or JSON exports. - For detailed API parameters, refer to the LMUnit API reference. - If you have feedback or need support, please email feedback@contextual.ai.
How to import this n8n workflow
- 1
Download the workflow JSON file after purchase.
- 2
Open n8n → click the menu → Import from File.
- 3
Select the downloaded JSON and import.
- 4
Set up credentials for each node that requires them.
- 5
Click Execute Workflow to test, then activate.
Setup guide
Setup guide included
Purchase to unlock the full step-by-step guide
Related N8n workflows
Flight deal analyzer with weather data using GPT, Google Flights & WordPress
Generate horror faceless shorts with OpenAI TTS, Replicate Video, and YouTube upload
Auto create & publish X-threads (Twitter) with GPT via Telegram & approval loop
Send automated replies on Telegram using ChatGPT assistant messages
Clone LinkedIn writing styles with AI analysis & prompt generation to Airtable
Create a Notion AI assistant with Google Gemini for managing tasks & content
Auto-generate blog & AI image from YouTube videos with Dumpling AI & GPT-4o
Convert YouTube videos into SEO blog posts with GPT-4o, Dumpling AI, and Flux
Reviews
No reviews yet
Be the first to buy and share your experience.
Leave a review
Sign in to share your experience with this workflow.
Create a free account to purchase workflows.
- JSON blueprint — instant download
- Setup guide PDF included
- 5 downloads · valid 30 days
- Works with n8n
Need help setting this up?
Book a 3-hour live setup session with an Agility consultant.
- Configure live on Google Meet / Zoom
- Free follow-up if workflow has defects
- Platform expert assigned to you