Trust in AI goes well beyond whether the facts are right.
That gap is the subject of our first paper as Musa Guide, which we will present at Fantastic Futures 2026, the AI4LAM conference in Washington this September. This year's theme is "Trust in the Loop", and our contribution is called Visitor in the Loop: Trust and Mutual Learning in AI-Mediated Museum Interpretation, written with Hendrik Schäfer and, from Museo Miraflores, Maria Gadsden Escobar and Hari Castillo.
When museums ask us about trust, the first question is almost always about hallucination. It is a fair question, and it still gets most of the attention. It is also much less of an issue today than it was two years ago.
With a verified knowledge base and good retrieval behind it, a model can be kept to the material it was given, and made to say plainly when the answer is not in there. Those are solvable problems with known solutions, and the solutions keep getting better without anyone in the sector having to do anything.
So if accuracy were the whole of trust, this would already be a settled topic.
The part that does not get solved by better models is selection. What goes in, what stays out, what order it is told in, and whether the whole thing holds someone's attention for more than a minute.
Those are editorial decisions, and they are the same decisions a curator makes when writing a wall text or scripting an audio guide. A visitor standing in front of an object has no way to audit them. They are trusting that somebody made those calls well.
That is a much harder thing to measure than factual accuracy. You can test whether a guide said something false. Testing whether it said the right thing, in the right order, to this particular person, is a question about judgement.
The interesting result for us was how much the answer to that question changes visitor behaviour, with everything else held constant.
Running the same system with and without carefully curated narratives, explorative interactions went from under 5% to about a third. Visitors who had a story to follow were far more likely to step off it and ask something of their own. Retention roughly doubled.
That is the opposite of what people usually assume about conversational interfaces. The intuition is that structure constrains people and open-ended chat frees them. In practice a blank prompt in front of an unfamiliar object mostly produces silence, or "tell me about this". A curated path gives a visitor something to react to, and reacting is where the conversation actually starts. We wrote about that tension in more detail in Infinite Paths for Everyone.
There is a second half to this, which is what comes back the other way.
When visitors talk to a guide, a museum hears, often for the first time, what its audience is actually curious about at the object, in the visitors' own words and across every visit rather than in a survey afterwards. That is where interest sits, and where the interpretation has gaps.
A house can then send its content where the questions are, and over time cover a growing share of what visitors genuinely ask. The loop runs in both directions: the curator shapes the visit, and the visits reshape what the curator writes next.
The conclusion we land on is a simple one. The trustworthy part of an AI guide is not the model. It is the person who decided what the guide should say, and who keeps deciding as the questions come back.
That is why we build for curators rather than around them, and why "AI does the interpretation" was never the product we wanted to sell.
We will be at Fantastic Futures in Washington this September. If you are going, we would like to hear how your institution is thinking about this, particularly if you disagree with us.
You can also read the Museo Miraflores case study for how this played out in a live deployment, or our State of AI in Museums 2026 report for the wider picture.