PaPoo
cover

Stop Testing MCP Servers Through the Model

What surprised me here is not the advice itself, but how bluntly the piece says the quiet part out loud: if your MCP server only “works” when an AI client happens to use it correctly, you don’t really know anything yet. That feels right. I’ve seen too many agent-tool integrations where people blame the model for a bug that was really a bad schema, a sloppy error shape, or a server that was already broken before the prompt even entered the picture.

The part I’d actually try first is the author’s insistence on treating MCP like a normal protocol, not some mystical agent layer. Start by sending initialize by hand. Then hit tools/list. Then call the tool directly. That’s the sort of boring discipline that saves hours. If the server can’t survive a raw request, the Inspector is just a nicer place to watch it fail.

I also like the push to test the descriptions as if they were code. That’s easy to dismiss until you’ve watched a model choose the wrong tool because two names overlap, or because a description is vague enough to invite guesswork. The article is strongest when it treats tool metadata as part of the contract, not documentation garnish. That seems especially important for Claude-style clients, where the model is making decisions from what you exposed, not from what you meant.

Where I’m a little less convinced

The piece is very confident that MCP surfaces are small enough to test exhaustively. Maybe for some servers, yes. But I wouldn’t overgeneralize that. Once a tool wraps a messy backend, the protocol layer may be small while the behavior space is not. You can still get a false sense of completeness if the tests are only checking schema validation and one happy-path fake backend.

I’m also a bit wary of the “just use a fake backend” advice when teams get lazy with it. Fakes are great for deterministic tests, but they can drift from the real service in exactly the places that matter: pagination quirks, auth edge cases, rate limits, weird serialization. So yes, use the fake. Just don’t pretend that means you’re done.

What I do buy completely is the warning about unreadable tool results. If a tool returns a giant blob, then of course the agent starts acting lost. That’s not the model being dumb; that’s you making the answer hard to consume. The article’s suggestion to keep the returned content minimal and expose the rest as a resource sounds like the kind of design cleanup that improves both humans and agents.

The broader message is pretty simple: if you want Claude or any other model to use your MCP server reliably, make the protocol surface deterministic before you let the model anywhere near it. That’s less glamorous than “AI debugging,” but it’s the part that actually scales.


Reference: How to Test and Debug an MCP Server: From the Inspector to Automated Tool Tests

同じ著者の記事