MCPLabby Inspectr

πŸ§ͺ Test and evaluate MCP Servers with LLMs

Test how well LLM agents use your MCP tools, compare different models, and track quality over time with automated testing and detailed reports.

localhost:5173
MCPLab dashboard showing evaluation results and test scenarios

LLM agent evaluations

Evaluate how models use your MCP tools

Give your models the same tasks. MCPLab runs the agents through their APIs, records their MCP tool calls, and checks the results against your evaluation criteria.

  • Claude

  • ChatGPT

  • Azure Foundry

One evaluation suite, consistent checks across models.

  • Google Gemini
  • DeepSeek
  • Ollama
  • Groq
  • Mistral
  • OpenAI-compatible

Measure agent behavior

Run repeatable scenarios and review pass rates, latency, and tool usage to understand how models handle tasks with your MCP tools.

Check tool use

Check required tools, call order, input arguments, and final answers. Add judge assertions when answer quality needs a closer look.

Inspect the evidence

Inspect captured tool calls and responses, review saved reports, and rerun scenarios as you improve your MCP server.

See It in Action

Rich visual reports, detailed traces, and interactive dashboards.

Dashboard
Dashboard
Run Evaluation
Run Evaluation
Evaluation Results
Evaluation Results
Run Detail
Run Detail
MCP Analysis
MCP Analysis
Tool Analysis
Tool Analysis
Reference Reports
Reference Reports
AI Assistant
AI Assistant

Track pass rates, latency trends and recent runs at a glance.

Core Capabilities

  • β€’HTTP SSE Transport for MCP servers
  • β€’Multi-LLM support (OpenAI, Claude, Azure)
  • β€’Rich assertions & variance testing
  • β€’Detailed JSONL trace logs

Analysis & Reporting

  • β€’Trend analysis & LLM comparison
  • β€’HTML, JSON, Markdown outputs
  • β€’Custom metrics & KPI tracking
  • β€’Markdown reports for each run

Developer Experience

  • β€’CI-friendly CLI for scheduled runs
  • β€’Reusable agent and server libraries
  • β€’Interactive HTML reports
  • β€’Multi-agent testing via CLI
Rover, the MCPLab browser companion

MCPLab Rover

Run your browser evaluations on repeat

Repeating the same prompts in Claude or ChatGPT? Queue your MCPLab scenarios and let Rover handle submission and response capture. Review the checked answers and saved results in MCPLab.

Automate the repeated work

Queue scenarios, submit prompts, and capture completed answers in Claude, ChatGPT, or a compatible learned provider.

Reuse your evaluation scenarios

Apply your MCPLab response checks and judge assertions to browser answers, with results saved in the same place.

Control conversation context

Start fresh conversations between scenarios or continue the same chat. Follow progress and stop active work from Rover.

Quick Start

Up and running in under a minute.

1. Install

$ npx @inspectr/mcplab --help

2. Create eval config

servers:
  my-server:
    transport: "http"
    url: "http://localhost:3000/mcp"

agents:
  claude:
    provider: "anthropic"
    model: "claude-haiku-4-5-20251001"
    temperature: 0

scenarios:
  - id: "basic-test"
    agent: "claude"
    servers: ["my-server"]
    prompt: "Use the tools to complete this task..."
    eval:
      tool_constraints:
        required_tools: ["my_tool"]
      response_assertions:
        - type: "regex"
          pattern: "success|completed"

3. Run evaluation

$ npx @inspectr/mcplab run -c eval.yaml

AI-Powered Tools

Built-in AI assistants to supercharge your workflow.

Scenario Assistant

AI chat to help design and refine evaluation scenarios. Describe what you want to test and get ready-to-use YAML configurations.

Result Assistant

AI chat to analyze and explain completed run results. Understand failures, spot patterns, and get actionable improvement suggestions.

MCP Tool Analysis

Automated review of your MCP tool definitions for quality, safety, and LLM-friendliness. Get recommendations before testing.

Agent Workflows

Use MCPLab with LLM agents

Install the `mcplab-assistant` skill and reuse the same prompts across Claude, OpenAI Codex, and similar coding agents.

Install the skill once

Use the Skills CLI installation flow documented in the MCPLab docs.

npx skills add https://github.com/inspectr-hq/mcplab --skill mcplab-assistant
Open installation guide

Prompt examples for any LLM agent

These prompts are agent-neutral and work as reusable starting points.

Config authoring

Generate a minimal, valid starter config before scaling up scenarios and agents.

Use the mcplab-assistant skill to draft a minimal MCPLab eval config with one scenario and OAuth client-credentials auth.
Run + compare

Run one config across multiple agents and summarize performance differences clearly.

Use mcplab-assistant to run this config and compare claude-haiku vs gpt-4o-mini with --agents, then summarize pass-rate differences.
Result analysis

Analyze run artifacts and return concrete fixes tied to failed scenarios.

Use mcplab-assistant to analyze this run directory and explain failing scenarios with concrete fixes and a rerun command.

Documentation

Everything you need to go from setup to deeper analysis

Start quickly, then dive deeper with guides for setup, scenario design, app workflows, debugging, and advanced evaluation analysis.

Overview
What MCPLab does and when to use it.
Open docs
Installation
Install MCPLab and configure your API keys.
Open docs
Quick Start
Write your first eval and see results in under 5 minutes.
Open docs
Setting Up Evaluations
Set up a robust evaluation workflow before running your first full test suite.
Open docs
Scenario Configuration
Detailed guide for writing MCPLab scenarios and assertions.
Open docs
Libraries & Refs
Reuse shared servers, agents, and scenarios across evaluation configs.
Open docs
Running Evaluations
The mcplab run command and all its options.
Open docs
Configuration
Write eval.yaml β€” servers, agents, scenarios, assertions, and auth.
Open docs
Reports & Output
What MCPLab writes after a run and how to work with it.
Open docs
Results Query
LLM-first querying of run artifacts with mcplab results.
Open docs
CI/CD
Run MCPLab in GitHub Actions and other CI pipelines.
Open docs
Command Reference
All mcplab commands and their flags in one place.
Open docs
Starting the App
Launch the MCPLab web UI and find your way around.
Open docs
Configurations
Browse, filter, and run eval configs from the Configurations page.
Open docs
Running Evaluations
Launch and monitor evaluations from the web UI.
Open docs
Rover Browser Agents
Run repeatable browser evaluations in Claude, ChatGPT, and compatible browser agents.
Open docs
Set Up Rover
Start MCPLab, configure a browser agent, install Rover, and connect it.
Open docs
Run Evaluations
Run managed evaluations and use Rover's local queue in the active chat.
Open docs
Learn Browser Providers
Teach Rover how to work with another compatible browser agent.
Open docs
Review Rover Results
Monitor Rover work and review captured browser evaluation results.
Open docs
LangSmith Integration
Export evaluation traces to LangSmith for rich model and tool-call debugging.
Open docs
Analysing Results
Understand run output, compare agents, and browse markdown reports.
Open docs
AI Assistants
Use the Scenario and Result AI assistants to work faster.
Open docs
MCPLab Assistant Skill
Install and use the mcplab-assistant skill from skills.sh in Codex/Claude-style agent workflows.
Open docs
OAuth Debugger
Debug OAuth 2.0 authorization flows for MCP servers step by step.
Open docs
Scenario Setup in the App
Create and manage evaluation scenarios directly in the MCPLab app UI.
Open docs
MCP Tool Analysis
Review MCP tool definitions for quality and LLM-readiness.
Open docs
Library
Manage reusable agents and servers shared across eval configs.
Open docs
Configuration Schema
Complete field reference for eval.yaml.
Open docs
Tool and Response Assertions
Assertion guide with examples for tool checks, response checks and agent judge.
Open docs
Environment Variables
All environment variables read by MCPLab.
Open docs