Test how well LLM agents use your MCP tools, compare different models, and track quality over time with automated testing and detailed reports.

LLM agent evaluations
Give your models the same tasks. MCPLab runs the agents through their APIs, records their MCP tool calls, and checks the results against your evaluation criteria.
Claude
ChatGPT
Azure Foundry
One evaluation suite, consistent checks across models.
Run repeatable scenarios and review pass rates, latency, and tool usage to understand how models handle tasks with your MCP tools.
Check required tools, call order, input arguments, and final answers. Add judge assertions when answer quality needs a closer look.
Inspect captured tool calls and responses, review saved reports, and rerun scenarios as you improve your MCP server.
Rich visual reports, detailed traces, and interactive dashboards.








Track pass rates, latency trends and recent runs at a glance.

MCPLab Rover
Repeating the same prompts in Claude or ChatGPT? Queue your MCPLab scenarios and let Rover handle submission and response capture. Review the checked answers and saved results in MCPLab.
Queue scenarios, submit prompts, and capture completed answers in Claude, ChatGPT, or a compatible learned provider.
Apply your MCPLab response checks and judge assertions to browser answers, with results saved in the same place.
Start fresh conversations between scenarios or continue the same chat. Follow progress and stop active work from Rover.
Up and running in under a minute.
1. Install
2. Create eval config
servers:
my-server:
transport: "http"
url: "http://localhost:3000/mcp"
agents:
claude:
provider: "anthropic"
model: "claude-haiku-4-5-20251001"
temperature: 0
scenarios:
- id: "basic-test"
agent: "claude"
servers: ["my-server"]
prompt: "Use the tools to complete this task..."
eval:
tool_constraints:
required_tools: ["my_tool"]
response_assertions:
- type: "regex"
pattern: "success|completed"3. Run evaluation
Built-in AI assistants to supercharge your workflow.
AI chat to help design and refine evaluation scenarios. Describe what you want to test and get ready-to-use YAML configurations.
AI chat to analyze and explain completed run results. Understand failures, spot patterns, and get actionable improvement suggestions.
Automated review of your MCP tool definitions for quality, safety, and LLM-friendliness. Get recommendations before testing.
Agent Workflows
Install the `mcplab-assistant` skill and reuse the same prompts across Claude, OpenAI Codex, and similar coding agents.
Use the Skills CLI installation flow documented in the MCPLab docs.
These prompts are agent-neutral and work as reusable starting points.
Generate a minimal, valid starter config before scaling up scenarios and agents.
Run one config across multiple agents and summarize performance differences clearly.
Analyze run artifacts and return concrete fixes tied to failed scenarios.
Documentation
Start quickly, then dive deeper with guides for setup, scenario design, app workflows, debugging, and advanced evaluation analysis.