> For the complete documentation index, see [llms.txt](https://mcp-test-kitchen-docs.cakewalk.security/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://mcp-test-kitchen-docs.cakewalk.security/scenarios/scenarios.md).

# How Scenarios Work

One scenario is active per user at a time. It decides what the server does when your client calls run\_configured\_test\_scenario.

A scenario is a saved failure. You pick one in the console, save its parameters and connect your client. From then on the server behaves the way that scenario says, on the call that scenario names.

There is one endpoint and, for most scenarios, one tool.

***

## One Tool Carries Every Scenario

Every MCP session exposes `run_configured_test_scenario`. What it does depends entirely on the scenario saved against your account, not on the arguments you pass.

```json
{
  "name": "run_configured_test_scenario",
  "arguments": { "message": "optional label echoed back for debugging" }
}
```

The `message` argument is optional and never changes behavior. It comes back inside the result, which is useful when you are correlating calls by hand.

One scenario adds to this pattern. `compat.sdk_v2` keeps `run_configured_test_scenario`, where it returns an index of the matrix, and registers three dedicated tools alongside it: `simulate_ticket_close_mrtr`, `show_negotiated_mcp_protocol` and `verify_order_region_param_header`.

Alongside the tool, every session advertises the same four catalog items regardless of which scenario is active: `valid_resource`, `invalid_resource`, `valid_prompt` and `invalid_prompt`. They let you exercise `resources/read` and `prompts/get` against valid and schema invalid responses without touching your scenario selection.

***

## Six Scenarios Ship Today

Each one breaks a different layer, so the layer is the thing you choose between.

| ID                                                                            | Layer it breaks           | What it does                                                             |
| ----------------------------------------------------------------------------- | ------------------------- | ------------------------------------------------------------------------ |
| [`baseline`](/scenarios/scenarios/baseline.md)                                | None                      | Answers normally, with an optional delay. The control case               |
| [`errors.http_status_sequence`](/scenarios/scenarios/http-status-sequence.md) | HTTP transport            | Returns a configured sequence of status codes across successive requests |
| [`errors.jsonrpc_error`](/scenarios/scenarios/jsonrpc-error.md)               | JSON-RPC envelope         | Returns HTTP 200 carrying an `error` object                              |
| [`errors.tool_error`](/scenarios/scenarios/tool-error.md)                     | Tool result               | Returns a valid result carrying `isError: true`                          |
| [`elicitation.approval`](/scenarios/scenarios/elicitation-approval.md)        | Server to client requests | Calls `elicitation/create` before completing the tool call               |
| [`compat.sdk_v2`](/scenarios/scenarios/sdk-v2-compat.md)                      | Protocol revision         | Exercises the 2026-07-28 multi round trip path and the older fallback    |

### Where each failure is injected

Three of the six break a different layer of the same request. The layer decides what your client has to read to notice.

```mermaid
%%{init: {"flowchart":{"nodeSpacing":80,"rankSpacing":60,"curve":"linear"}}}%%
flowchart TD
    A(["Your client sends a POST"]) --> B{"HTTP status"}
    B -->|"4xx or 5xx"| E1(["errors.http_status_sequence<br/>No body to parse"])
    B -->|"200"| C{"JSON-RPC envelope"}
    C -->|"error object"| E2(["errors.jsonrpc_error<br/>Tool never runs"])
    C -->|"result object"| D{"isError on the result"}
    D -->|"true"| E3(["errors.tool_error<br/>Tool ran and failed"])
    D -->|"not set"| OK(["Success"])
```

Read it downward. Every layer you pass returns a more successful looking response than the one above it, and `errors.tool_error` sits at the bottom because a failed tool result is delivered inside a success.

`elicitation.approval` needs a client that advertises elicitation support and handles `elicitation/create`. Support splits by client surface rather than by vendor, so check your specific surface before selecting it.

***

## A Selection Applies to Your Next Session

Saving a scenario does not change a connection that is already open. The server reads your saved selection when your client sends `initialize` without an `Mcp-Session-Id`, and builds a new session only if that selection differs from the one you are already holding.

In practice, reconnecting your client is enough.

{% hint style="warning" %}
Reconnecting with the **same** scenario and the **same** parameters reuses your existing session, including its invocation counter. If you are testing `onInvocation` and the failure does not land where you expect, that counter is why. Use **Terminate** in the console's Scenarios pane to force a fresh one. The button appears only while the console knows about a live session, and it learns that when the page loads, so reload the console if your client connected afterwards.
{% endhint %}

You hold one active scenario session at a time. Saving a different scenario, or the same scenario with different parameters, replaces it on your next connection.

***

## Invocation Counters Decide When a Failure Lands

Several scenarios take an `onInvocation` parameter. The server counts requests within your session and applies the failure when the count matches exactly, so the third call of a scenario set to `onInvocation: 2` succeeds.

This is the parameter that separates a testbed from a fixed mock. A server that fails every time gets caught immediately. The failure that costs you arrives after things worked.

Counters start at 1, live on the session and reset only when the session is replaced.

### Three counters, not one

Which requests a scenario counts depends on the layer it breaks. This catches people out, so it is worth knowing before you set `onInvocation`.

| Counter        | Counts                                             | Used by                                                                  |
| -------------- | -------------------------------------------------- | ------------------------------------------------------------------------ |
| `http_request` | Every POST to the endpoint, including `initialize` | `errors.http_status_sequence`                                            |
| `tools_call`   | Every `tools/call`, counted before the tool runs   | `errors.jsonrpc_error`                                                   |
| Per tool name  | Calls that reach a specific tool's body            | `errors.tool_error`, `baseline`, `elicitation.approval`, `compat.sdk_v2` |

The first row is the one that surprises people. `initialize` is a POST, so a status sequence starting with an error hits your handshake rather than your first tool call.

The second row has a consequence too. `errors.jsonrpc_error` answers at the wire layer, so the tool body never runs and the per tool counter does not advance for the call it intercepts.

***

## Parameters Are a JSON Object

Every scenario defines its own parameter shape. The console validates what you save and rejects malformed JSON, non object payloads and null collections before the selection is stored.

Omit a field and its default applies. Save `{}` and every default applies. Each scenario page lists its fields, types and defaults, and the console offers **Load example** for a working starting point.

***

## What a Scenario Will Not Do

Worth knowing before you design a test around one.

* **No pass or fail.** Nothing here checks your client against the specification. The scenario produces a failure and the record shows what happened. The reading is yours.
* **It does not test your server.** The direction is fixed: this server misbehaves, your client responds. Point [MCP Inspector](https://github.com/modelcontextprotocol/inspector) at a server instead.
* **The transport is stateless.** The server cannot bridge state across a live connection, so stateful paths stay out of scope. This is why `compat.sdk_v2` reports which cells of the matrix a stateless host can cover.
