Quattr leads AEO, SEO, and content rankings on G2 Spring 2026. View our G2 badges →
Request demo
Request demo

Working with AI

Evaluating MCPs before you connect them

Whose data, what governance, what metering, the eight questions that separate a teammate from a firehose.

Key takeaways

  • Every connection you add becomes a voice your assistant trusts. Interview it before it gets one.
  • Start with the data: where it came from, how stale it is, and what sits outside the dataset.
  • The question demos never show is what the tool does when it cannot answer. A refusal you can trigger on demand is the whole test.
  • Metering and identity are not small print. Per-query pricing teaches your team to stop checking, and an unlabeled answer in a multi-account world is somebody else's right number.
  • Adoption is decided at the IT desk, so write the eight answers into the approval ticket before anyone asks.

An analyst on your team asks a new tool something it has no data for. It answers anyway, fluently, in the same confident register it uses for everything else. Nobody catches it, because a wrong answer and a right one look identical when both are written well.

Every connector you add becomes a voice your assistant trusts. An MCP is the standard that lets an assistant read a data source directly, so you ask questions in the assistant instead of exporting spreadsheets. That voice is the price of the convenience.

So the question is whether this one is a teammate, an accountable source with known data and known limits, or a firehose answering everything with equal confidence. Eight questions separate them. They take about an hour.

Ask them before you connect. Afterwards the answers blend into everything else and the audit gets much harder.

Whose data is it, and how fresh?

Questions one through three, all about the data.

Provenance first. Does the tool answer from your data, an index of the public web, a panel, or a model's memory? None of those is wrong. Each answers a different question, and a tool that will not say which has already failed.

Then freshness. What is the lag, and does the tool tell you per answer or leave you guessing? An as-of date, the date the data was last complete, stops a reporting lag from reading as a loss. A number without an age is a guess wearing confidence.

Then coverage. What is inside the dataset and what is not, and does the tool say so at the edges or quietly extrapolate?

Provenance also predicts how your tools will disagree later. When an index tool and a first-party connection return different numbers for one question, the difference is almost always provenance or freshness. A roster of documented answers turns a credibility crisis into a lookup.

See also: why every answer carries its scope and its as-of date →

What happens when it does not know?

Question four, and no demo shows it.

Ask it something its data cannot support and watch. A teammate says it cannot answer, and why. A firehose answers anyway, and that answer eventually reaches a deck with your name on the meeting invite.

The buyers who adopt this fastest ask it as one pointed worry. Will it invent numbers. Enterprise search teams already using assistants daily bring that question to every new connection, and the tools that survive answer it with a demonstrable no, backed by refusals you can trigger on demand.

Question five is scope control. Can you pin dates, segments and filters, and does the answer echo the scope it used? Unscoped answers are how two people quote one tool against each other in a meeting.

Eight questions, one hour, and the difference between a teammate and a firehose.

Ask it yourself

What data sources does this connection use, how fresh is each, and what happens when I ask something it can't support?

Access, cost, and whose account answers

Questions six through eight, the ones your security review will care about.

Access. What can it read, what can it change, how does it authorize? A read-only scope through a browser sign-in is a different risk class from a tool holding write scopes or long-lived keys. The piece on managing MCP permissions carries the full version.

Metering. What does a question cost, and does the pricing shape behavior? Per-query pricing teaches teams to hesitate, and that hesitation never appears on the invoice.

Identity. Whose account is answering, and can you see it? Multi-tenant tools should label every result with the account it belongs to, because the wrongest number available is another company's right number. That sounds cosmetic until a consultant with five client accounts quotes the wrong tenant into a board deck.

See also: the security review answers, question by question →

Adoption is won at the IT desk

Your practitioners are enthusiastic in the first meeting. Somebody else owns the timeline.

Connector approval policies, desktop-install rules and the security questionnaire decide when your team gets to use the thing. Teams already fluent in ChatGPT and Claude still wait on tickets. The tools that win are the ones whose eight answers are documented and pasteable into the approval request.

Budget anxiety shows up in the same conversations. Teams worry about burning limited assistant tokens pasting exports into chat, and a native connection answers that too: the data arrives as results rather than a pasted spreadsheet.

The checklist matters because operational questions decide adoption more than feature lists do. Teams that pre-write the approval ticket from the eight answers report the shortest queues, and the ticket doubles as onboarding for whoever requests next.

Give every candidate the same test

One question, asked of every tool you are weighing, chosen so answering it requires knowing your own limits.

A traffic spike that needs a real-or-noise call works well. It reveals whether the tool has any concept of statistical significance, a check on whether a change is bigger than normal week-to-week wobble, or whether it just narrates the delta.

The workflow card below shows what a passing answer looks like, a worked example of one question answered end to end: a claim arriving with a test, a scope and an as-of date attached. The update-week and spike-verdict scenarios in the gallery make good test questions, because their correct answers include a refusal.

The workflow that does this: Conversion spike: real? →

The interview scales to the roster

Most teams end up with several connections. An index tool for market context, a first-party one for ground truth, usually more after that.

The eight questions are how that roster stays coherent. Every member interviewed the same way. Every disagreement traceable to a documented difference in provenance or freshness, rather than sitting there as a mystery nobody settles before the meeting ends.

Our piece on working across several connectors covers running the roster day to day, and the comparison pages hold completed interviews of the major options. Steal the format.

These Working with AI articles are that interview applied to one tool in public. An hour spent interviewing a tool is the cheapest diligence in your stack, and the only hour that gets more expensive the longer you wait.

See also: running several connectors without them arguing →

Frequently asked

Can we run the interview after connecting?
You can, and it is harder. Once the tool is answering, its output blends into everything else your assistant says, and separating a provenance problem from a prompt problem takes longer than the original hour would have.
What if a tool will not answer some of the eight?
That is an answer. A vendor who cannot say where the data comes from, or what happens at the edges of coverage, has told you which of the two categories it is in.
Do the eight questions apply to non-search connectors?
Yes. Provenance, freshness, coverage, refusal, scope, access, metering and identity are properties of any data source, and a docs connector that invents an answer costs you the same credibility a search tool does.

Request a demo Take a test drive Steal the prompts