What jumps out to me here is that this is less about hallucinations than about accountability. That sounds subtle, but in an MCP setup it’s the whole ballgame. If an agent says something correct but points to the wrong tool or record, a source-blind checker can happily green-light it. That feels like a real failure mode, not an academic nitpick.
I think the paper’s strongest idea is also the most boring one: keep the source IDs attached all the way through the pipeline. Boring in the good sense. A lot of agent eval stories quietly flatten everything into one blob of “context,” and then act surprised when attribution breaks. If your answer says “according to the account record” but the fact actually came from a policy doc, that is not just a citation bug. In a medical or customer-support setting, it changes what a human reviewer thinks they can trust.
What I like is that they don’t pretend one generic faithfulness score solves this. They’re explicitly separating “is this fact supported somewhere?” from “is it supported by the source the answer names?” That distinction matters more as agents get more tool-heavy. The more an agent mixes search results, records, databases, and metadata, the less meaningful pooled evidence becomes.
The part I’m less impressed by is the comfort of the evaluation numbers. Catching 138 of 139 blocked claims sounds strong, but the source-selection story is still shaky in the harder setting: 50.3% exact-source accuracy among similar sources is basically a coin flip with a slight edge. That doesn’t invalidate the system, but it does tell you where the hard problem really is. Blocking wrong answers is one thing. Correctly resolving which of several plausible sources a claim came from is the real knife fight.
That also makes me wonder how usable this is outside controlled traces. The setup seems to depend on the agent preserving decent tool provenance in the first place. If the trace is messy, or the answer is synthesized across several tools, I suspect the nice clean source-ID story gets harder fast. Maybe that’s exactly why they keep the verifier conservative and allow repair/fallback rather than forcing a confident answer. That feels honest.
I also appreciate the practical edge of the paper: they’re not just measuring support, they’re asking what a system does after it blocks. That’s where a lot of safety work gets fake. A blocker that just says no is easy to admire and hard to ship. A fallback answer is less glamorous, but probably what a production team actually needs if they care about not inventing provenance.
The broader implication, at least to me, is that MCP agents need a better contract than “grounded.” Grounded in what, exactly? If you’re building tools on top of Claude or any other model that routes across multiple sources, this distinction will matter sooner than people think. The eval gate should probably ask whether the answer is supported by any retrieved text, or by the specific source it claims. Those are not the same release criteria.
Reference: Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents