Building on MCP
Most first MCP servers are complete, correct, and unusable
An MCP server exposes your system to someone else's AI, not to a developer. Mapping one tool per API endpoint feels thorough and makes the model do five hops to answer one question. Here is how to spot the pattern and design task-shaped tools instead.
You have an API. It has endpoints, docs and tests, and it works.
So when someone asks for an MCP server, the obvious move is to wrap it. One endpoint, one tool. Done by lunch.
That is the first server almost everyone builds. It is complete, it is correct, and nobody can use it.
What does “MCP server” actually mean?
An MCP server is a program that exposes your system’s actions and data to an AI application through the Model Context Protocol, the open standard Anthropic published in November 2024 (opens in a new tab). The official docs (opens in a new tab) call it a USB-C port for AI: one standard plug, so Claude, ChatGPT or any other host can connect to your product without a custom integration each time.
That is the textbook meaning. Here is the one that matters when you build one:
An MCP server is not an API with a new hat. It is your system, exposed to someone else’s AI.
The caller is a model. It is not a developer. And it reads your server very differently.
The server that passes every test
Take a billing API, mapped one tool per endpoint:
list_customersget_customerupdate_customerlist_subscriptionsget_subscription
Textbook. Also useless.
Ask it “is Sarah in arrears?” and the model has to find Sarah, fetch her record, list her subscriptions, pick the right one, then work out what “arrears” means from fields it has never seen before.
Five calls for one question. Every hop is a chance to pick the wrong tool, pass the wrong ID, or give up and guess.
It passes every test you would write. That is the problem. The tests are written for a developer, and a developer is never going to call it.
You have this somewhere, I would bet on it. A server, a plugin, an internal agent tool, where the tool list reads like the table of contents of your API reference.
Why one tool per endpoint fails
A developer reading your API has advantages the model does not. The model:
- sees your tool list fresh, every session
- has no screen to check what happened
- cannot ask a colleague which endpoint is the right one
- gets one shot at the right call, from the name and description alone
Endpoints are designed for someone who can hold the whole system in their head. Tools are chosen by something that only knows what you told it in one line.
Anthropic’s engineering team names the same trap in its guide to writing tools for agents (opens in a new tab): a common mistake is tools that merely wrap existing API endpoints. Their example is calendar software. Instead of list_users, list_events and create_event, build one schedule_event tool that finds availability and books the slot.
That is the whole idea. Wrap the job, not the endpoint.
The menu gets read before anyone orders
There is a second cost, and we walked into it this week.
Every tool you expose arrives with a name, a description and a schema. The host loads those into the model’s context before your user has typed a word. Fifty thin tools means fifty descriptions sitting in the model’s head, most of them for jobs nobody asked about.
On 27/09/2026 we found that our unattended Claude Code jobs had been loading every MCP connector we had installed. Together they came to roughly 586,000 tokens of setup. The context limit was 200,000.
So every one of those jobs failed before the model read the task. It never got to the question. It was still reading the menu.
Anthropic makes the same point in its note on code execution with MCP (opens in a new tab): tool definitions loaded up front can push an agent through hundreds of thousands of tokens before it reads a request. Our fix was to load only the connectors a job needs. The lesson for your server is the same one from the billing API: fewer, better tools.
How do you design task-shaped tools?
Here is the method we use now. It fits on an index card.
- List the questions, not the endpoints. Write down the ten things a real person asks about your product. “Is this customer in arrears?” “Which events sell out fastest?” Those are your tools.
- One tool per question.
get_customer_status(email)returns the customer, their plan, their last invoice and whether they owe money. One call. - Return what the next decision needs. Not the raw database row. Plain fields with plain names, so the model does not have to interpret
status_code: 4. - Run the “I want to” test. Read each tool name aloud after “I want to”. “I want to check a customer’s status” sounds like a task. “I want to list subscriptions” sounds like a database query. Merge or cut the second kind.
- Count your tools. If the list is longer than the questions you wrote in step 1, you have wrapped the API again.
Side by side:
| One tool per endpoint | Task-shaped tools | |
|---|---|---|
| Built from | The API reference | The questions users ask |
| Calls per question | Often 3 to 5 | Usually 1 |
| Where it goes wrong | Wrong tool, wrong ID, wrong order | Missing a task you did not list |
| Context used before the first question | Grows with every endpoint | Grows with every real job |
| Who it is designed for | A developer with the docs open | A model with one line of description |
The test worth keeping
If you only take one thing, take the “I want to” test. Open your server, or the tool list of the one you are planning, and read every name aloud.
If it sounds like something a person would ask for, keep it.
If it sounds like something a database would do, the model is about to do five hops to answer one question. And it will do them confidently.