Skip to main content
The Model Context Protocol (MCP) lets AI coding tools read and write your Braintrust data directly. Query production logs, author prompts and scorers, configure monitoring, and run evals from Claude Code, Cursor, Codex, VS Code, and any other MCP-compatible client.

Set up the MCP server

The examples below use https://api.braintrust.dev/mcp, the US data plane. If your organization is on the EU data plane, use https://api-eu.braintrust.dev/mcp instead. If you self-host, use the read-only MCP URL in your Data plane settings.
The Braintrust plugin for Claude Code wraps the MCP server and adds tracing capabilities. If you’ve installed that plugin, you don’t need to configure MCP separately.
1

Install Claude Code

If you haven’t already, install Claude Code.
2

Set your API key

Set the BRAINTRUST_API_KEY environment variable with your API key:
3

Add the Braintrust MCP server

Add Braintrust MCP server with API key authentication:
Alternatively, you can use OAuth authentication instead of an API key:
1

Install Claude Desktop

If you haven’t already, download and install Claude Desktop.
2

Add the Braintrust MCP server

Follow the Claude Desktop documentation to create a custom connector with the following details:
  • Name: Braintrust
  • URL: https://api.braintrust.dev/mcp
Claude Desktop uses OAuth 2.0 for authentication. You don’t need to provide an API key in the connector configuration - you’ll authenticate when you first use the server.
1

Install Codex

If you haven’t already, install Codex.
2

Set your API key

Set the BRAINTRUST_API_KEY environment variable with your API key:
3

Add the Braintrust MCP server

Edit ~/.codex/config.toml and add the Braintrust MCP server configuration:
This configures Codex to read your Braintrust API key from the BRAINTRUST_API_KEY environment variable.
4

Verify the setup

Launch Codex with the environment variable set:
Run the /mcp command to verify Braintrust is installed and accessible.
The Braintrust extension for Cursor automatically configures the MCP server for you. If you’ve installed that extension, you don’t need to configure MCP separately.
1

Install Cursor

If you haven’t already, download and install Cursor.
2

Add the Braintrust MCP server

Click to automatically add the Braintrust MCP server: Add to CursorOr manually add to .cursor/mcp.json:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.Cursor also supports OAuth authentication. If you omit the headers field, Cursor will prompt you to authenticate via OAuth when you first use the server.
1

Install VS Code

If you haven’t already, download and install Visual Studio Code.
2

Install an AI assistant extension

VS Code requires an AI assistant extension that supports the Model Context Protocol (MCP). Popular options include:Install one of these extensions from the VS Code marketplace.
3

Add the Braintrust MCP server

Add the Braintrust MCP server to your VS Code settings, either in workspace settings or user settings:
  • Workspace settings - Create or edit .vscode/mcp.json in your project:
  • User settings - Add to your VS Code user settings (Cmd+, / Ctrl+, → Search for “mcp”):
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.VSCode also supports OAuth authentication. If you omit the headers field, VSCode will prompt you to authenticate via OAuth when you first use the server.
4

Restart VS Code

Reload the VS Code window (Cmd+R / Ctrl+R) or restart VS Code to apply the configuration.
1

Install Devin Desktop

If you haven’t already, install Devin Desktop.
2

Add the Braintrust MCP server

Edit ~/.codeium/windsurf/mcp_config.json and add the Braintrust server:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.
3

Restart Devin Desktop

Close and reopen Devin Desktop to load the new MCP server configuration.
1

Install Gemini CLI

If you haven’t already, install Gemini CLI.
2

Set your API key

Set the BRAINTRUST_API_KEY environment variable with your API key:
3

Add the Braintrust MCP server

Edit ~/.gemini/settings.json and add the Braintrust MCP server configuration:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.
4

Verify the setup

Launch Gemini CLI and run the /mcp command to confirm the Braintrust server is connected.
1

Install Antigravity

If you haven’t already, install Antigravity.
2

Open the MCP configuration

Open Antigravity settings, go to the Customizations tab, and select Open MCP config to edit mcp_config.json (located at ~/.gemini/config/mcp_config.json).
3

Add the Braintrust MCP server

Add the Braintrust server to mcp_config.json:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.
4

Refresh the server list

Save the file, then refresh the Installed MCP servers section to load the new configuration.
1

Install Zed

If you haven’t already, install Zed.
2

Add the Braintrust MCP server

Open your Zed settings (Cmd+, on macOS / Ctrl+, on Windows/Linux) and add the Braintrust server under context_servers:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.If you omit the headers field, Zed prompts you to authenticate via OAuth when you first use the server.
3

Verify the setup

Open the Agent Panel settings and confirm the Braintrust server appears in the context servers list with a green indicator.
1

Install Amp

If you haven’t already, install Amp.
2

Add the Braintrust MCP server

Edit ~/.config/amp/settings.json and add the Braintrust server under amp.mcpServers:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.
3

Verify the setup

Restart Amp, then run amp mcp list to confirm the Braintrust server is connected.
For automatic tracing of OpenCode sessions, consider the Braintrust plugin for OpenCode.
1

Install OpenCode

If you haven’t already, install OpenCode.
2

Add the Braintrust MCP server

Edit your OpenCode configuration file and add the Braintrust MCP server:
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.
3

Restart OpenCode

Restart OpenCode to apply the configuration changes.
1

Install Warp

If you haven’t already, download and install Warp.
2

Add the Braintrust MCP server

Open Warp and navigate to Settings > AI > MCP Servers. Add a new server with the following details:
  • Name: Braintrust
  • URL: https://api.braintrust.dev/mcp
  • Header: Authorization: Bearer YOUR_BRAINTRUST_API_KEY
Replace YOUR_BRAINTRUST_API_KEY with your actual API key.
3

Verify the setup

Once added, the Braintrust MCP server will be available in Warp’s AI agent. You can verify the connection from the MCP Servers settings page.
Any MCP-compatible client can connect to the Braintrust MCP server. The server is available at:
Streamable HTTP transportPass your API key as a bearer token in the Authorization header:
Most clients that support remote MCP servers accept a URL and optional headers. Refer to your client’s documentation for where to configure these. SSE-only MCP clients will not work with the Braintrust server.OAuth authenticationThe Braintrust MCP server also supports OAuth 2.0. If your client supports OAuth-based MCP authentication, you can connect without an API key and authenticate interactively.

Install the Braintrust SDK

Ask your AI assistant to set up Braintrust in your project:
Your assistant reads the docs://sdk-install resource, detects your programming language and frameworks, installs the appropriate SDK, and configures auto-instrumentation. Once complete, it runs your app, verifies traces are being logged, and provides a permalink to view them in Braintrust.

Supported tools

Braintrust MCP provides read and write tools for the data and objects in your Braintrust organization. Your AI assistant can call several of them in one task. For example, it can query your logs, create a scorer from what it finds, and then run an eval to test it. The tools below are grouped by what you would use them for.
Write tools act on your Braintrust organization using the permissions of your authenticated account. Configure your MCP client to require confirmation before it runs a write tool.

Explore your data

  • sql_query - Query experiments, datasets, and logs using SQL. Supports SELECT, FROM, WHERE, GROUP BY, ORDER BY, and LIMIT.
  • infer_schema - Discover the available fields, data types, and most common values in experiments, datasets, or logs.
  • summarize_experiment - Get aggregated performance metrics for an experiment, optionally compared to a baseline.
Example prompts:
  • “Show me the last 10 logged requests with errors”
  • “Compare accuracy scores between my GPT-4 and Claude experiments”
  • “What were the costs for my recent chatbot experiments?”
  • “What fields are available in my experiment data?”
  • “Show me the schema for production logs”
  • “What metadata fields exist in this dataset?”
  • “Summarize the results of my latest A/B test”
  • “Compare my experiment against the baseline”
When a result exceeds 1 MB, sql_query uploads it to object storage and returns an overflow envelope instead of inline rows. The envelope includes an overflow_url (a signed URL to the JSON result), a byte_length, a row_count (when available), and an instructions field describing how to retrieve the full result. Set return_url: true to request a URL even when the result is below the threshold, which is useful when you want to download or save results without putting them in model context. Field values in the result are truncated to preview_length characters (1024 by default). Set preview_length: -1 to include untruncated field values. See SQL for query syntax, and View logs for the equivalent in the UI.

Find and share objects

  • list_recent_objects - List recently created projects, experiments, datasets, prompts, or functions you have access to.
  • resolve_object - Convert names to IDs or vice versa, and parse Braintrust URLs. Useful for looking up IDs before querying.
  • generate_permalink - Generate a direct web link to a Braintrust object for sharing or bookmarking.
Example prompts:
  • “Show me my recent experiments in the ‘chatbot’ project”
  • “List datasets in the recommendation engine project”
  • “What projects do I have access to?”
  • “Find the ID for my ‘sentiment-analysis’ experiment”
  • “What’s the name of experiment abc123?”
  • “Parse this Braintrust URL and tell me what object it points to”
  • “Create a link to share my experiment results”
  • “Generate a permalink to the customer-reviews dataset”

Configure Topics

These tools configure the Topics pipeline, which preprocesses traces into text, extracts facets from that text, and clusters the results. Each one expects your assistant to load the braintrust/topics-workflow skill first.
  • create_preprocessor - Create a versioned preprocessor from inline JavaScript that converts raw trace data into text.
  • test_preprocessor_on_trace - Run a saved, global, or inline preprocessor on up to 50 span, trace, or group references without writing to the source trace.
  • create_facet - Create a versioned facet that extracts a short summary from spans or traces. Facet extraction always uses Braintrust’s built-in facet model.
  • test_facet_on_trace - Run an inline facet definition on up to ten span, trace, or group references without writing the result to the source trace.
  • enable_topics_automation - Enable Topics for a project. This seeds processing for new traffic and doesn’t rewind historical data.
  • set_topics_automation - Update an existing Topics automation’s facets, scope, filters, sampling, or timing.
  • rewind_topics_automation - Rewind an existing Topics automation to process historical data from a start time or a recent window.
Example prompts:
  • “Set up Topics for my project”
  • “My traces don’t store conversation text on LLM spans. Write a preprocessor that works with my trace shape”
Rewinding a Topics automation processes historical traces and draws from your monthly model credits.

Build monitor views

  • generate_monitor_chart - Preview a monitor chart for project logs without modifying a saved view.
  • list_monitoring_views - List a project’s saved monitor views and chart IDs.
  • get_monitoring_view - Inspect a saved monitor view, including its options and ordered chart definitions.
  • create_monitoring_view - Create a project-scoped monitor view, optionally containing charts you already previewed.
  • update_monitoring_view - Insert, update, or remove charts in an existing monitor view, one edit at a time or several in bulk.
Example prompts:
  • “Create a dashboard for daily cost analysis”
  • “Add a p95 latency chart to my error monitoring view”
See Dashboards for the equivalent in the UI.

Manage automations and alerts

  • list_automations - List a project’s automations, including online scoring rules, alerts, exports, retention policies, and Topics automations. Filter by automation_id, name, or kind. Returns complete configurations, so you can inspect an automation before updating it.
  • set_automation_status - Pause or activate an alert, scheduled job, or online scoring rule.
  • create_log_alert - Create an alert for individual matching project logs. Use config.interval_seconds to throttle repeated notifications.
  • create_environment_update_alert - Create an alert for environment updates. Use config.environment_filter to limit notifications to specific environment slugs.
  • create_threshold_alert - Create an alert for an aggregate over a recent window of project data, evaluated on a schedule. Use it for averages, counts, rates, percentages, percentiles, and distributions.
  • create_scheduled_loop_job - Create a Loop job that runs on an interval or cron schedule over a recent window of project data.
Example prompts:
  • “What automations are configured in this project?”
  • “Alert me when the error rate goes above 2% over the last hour”
  • “Pause the online scoring rule you just created”
See Alerts for delivery channels and tuning.

Author prompts and evaluators

  • create_prompt - Create a versioned prompt from a completion-style prompt or chat messages. Set if_exists to replace to save a new version, or ignore to leave an existing prompt unchanged.
  • create_evaluator - Create a versioned LLM or inline code evaluator. Set output_type to score for numeric scores or classification for categorical labels.
  • test_evaluator - Run a saved, global, or inline evaluator against span, trace, or group references without writing results to the source trace.
  • update_online_scoring_rule - Save or rewind an online scoring rule that runs saved evaluator functions. New rules default to paused.
Example prompts:
  • “Write a scorer that detects the errors in these logs, then test it on a few traces”
  • “Create an LLM-as-a-judge scorer for helpfulness”
  • “Set up online scoring with the scorer you just created”
See Write prompts and Write scorers for details.

Run evals and edit datasets

  • run_eval - Run an experiment with a hosted dataset, inline rows, or a prior experiment as input data, any saved or inline task, and zero or more saved or inline scorers. When a prior experiment supplies the data, its outputs become expected values unless an expected value was already recorded.
  • edit_dataset_rows - Insert, update, or delete up to 100 dataset rows. Target a dataset by ID or name, and set create_if_missing to create a new named dataset.
Example prompts:
  • “Run an eval comparing these two prompts on my regression dataset”
  • “Add these traces to my regression dataset and set the expected output”
edit_dataset_rows can permanently delete dataset rows. Review the operations your assistant proposes before approving them.
run_eval creates an experiment and can execute your code or call AI providers, so it incurs compute and model usage.
See Run evaluations and Datasets for details.

Manage project settings

  • get_project_settings - Return a project’s typed settings, including the effective default preprocessor. An unset default resolves to the built-in thread preprocessor.
  • set_project_default_preprocessor - Set or clear a project’s default preprocessor. Pass null to restore the built-in default. This changes the default used by facets and other project functions that don’t select a preprocessor explicitly. Expects the braintrust/topics-workflow skill to be loaded first.
Example prompts:
  • “What preprocessor is my project using by default?”
  • “Make the preprocessor you just created the project default”
See Projects for the equivalent in the UI.

Search docs and load skills

  • search_docs - Search Braintrust documentation to find relevant guides, API references, and code examples.
  • load_braintrust_skill - Load a Braintrust workflow guide before using the tools it covers. Available skills are braintrust/automations-workflow, braintrust/evaluator-workflow, and braintrust/topics-workflow.
See Available skills for what each skill covers and which tools expect one. Example prompts:
  • “How do I create a custom scorer?”
  • “Show me examples of SQL queries”
  • “What’s the difference between experiments and project logs?”

Available skills

Skills are workflow guides your assistant loads with load_braintrust_skill and then follows. Where a tool reference tells your assistant what a tool does, a skill tells it the order to do things in, what to validate at each stage, and when to ask you for input.
  • braintrust/topics-workflow - Configure, evaluate, and improve the Topics pipeline, covering preprocessors, facets, scope, and Topics automations.
  • braintrust/evaluator-workflow - Create, test, refine, deploy, and rewind evaluators, and apply them to production logs with an online scoring rule.
  • braintrust/automations-workflow - Set up, validate, and manage alerts and scheduled Loop jobs, including threshold-triggered work, Slack and webhook delivery, and refining existing automations.
The Topics tools expect braintrust/topics-workflow to be loaded first, so your assistant validates the preprocessor and facets against real traces before it saves anything or enables an automation. Loading a skill is read-only and costs one tool call.

Available resources

MCP resources provide contextual documentation that AI assistants can read to perform tasks more effectively.
  • docs://sdk-install - Step-by-step guidance for installing the Braintrust SDK into a project, setting up tracing, configuring auto-instrumentation, and running your first eval.
  • docs://sql - Documentation for the sql_query tool, including syntax, available fields, and examples.
  • docs://url-formats - Reference for Braintrust URL patterns, used by the resolve_object tool.
  • docs://experiments - Background on Braintrust experiments and how to create them.
docs://sdk-install has companion resources for Python, TypeScript, Go, Java, Ruby, and C#. Your assistant reads the one matching your project automatically.

Troubleshooting

Invalid client errors: Verify the URL is exactly https://api.braintrust.dev/mcp (no trailing slash). Connection timeouts: Check internet connection. Corporate networks may need to allowlist api.braintrust.dev and *.braintrust.dev. MCP server not appearing: Restart your AI tool and verify JSON configuration syntax. Server URL errors on a self-hosted deployment: The MCP server derives its own address from the forwarding headers your ingress sets. If it reports that it could not determine the server URL, set the MCP_SERVER_URL environment variable on your data plane to your API URL.

Next steps