Generative AI has spent much of the past few years being evaluated through the language of capability: what models can generate, how quickly they can reason, how convincingly they can communicate, how much work they can automate and, increasingly, how independently they can act.
But another dimension of AI adoption can't be captured by benchmarks alone.
What happens when the system gets a person wrong?
Not merely when it hallucinates a date, invents a citation or produces an awkward sentence, but when it misunderstands someone’s identity, culture, gender, age, disability or circumstances. What happens when that output is no longer confined to a private conversation with a chatbot, but becomes part of how a company communicates with customers, evaluates people, provides services, or makes decisions?
At that point, bias stops being an abstract property of a model.
It becomes an organisational problem.
That distinction is central to the work of Catharine Montgomery, founder of Washington, D.C.-based Better Together Agency, whose recent work sits at the intersection of communications, reputation, generative AI and bias. Her perspective is unusual precisely because she does not approach responsible AI primarily from model architecture or computer science. She approaches it from the point where technology meets people — and where an output can become a customer experience, a communications decision, a reputational event or, in more consequential settings, part of a decision affecting someone’s life.
Montgomery has now spent several years researching public perceptions of bias in generative AI. Better Together Agency’s 2025 Building Consumer Trust: Addressing Bias in Generative AI report found that 33% of respondents would consider stopping their use of a generative-AI tool they perceived as biased. More strikingly, 92% said companies must address generative-AI bias.
The distinction between those numbers matters.
The 92% figure is not a measure of consumers threatening to abandon a brand. It measures expectation: an unusually broad signal that people believe responsibility does not end with the developer of the underlying model. Montgomery was explicit about that distinction when she sent me her responses, correcting an earlier formulation of my first question and asking that I not use an unsupported 84% brand-disengagement figure.
That precision is significant in a conversation about AI trust. Responsible AI depends not only on identifying erroneous information produced by machines, but also on ensuring that humans do not amplify inaccurate claims about those systems.
The broader findings of the 2025 study are equally revealing. Fairness and lack of bias ranked behind only accuracy and reliability among the characteristics respondents valued in generative-AI products. The report also found strong preferences for human oversight and transparency around AI involvement.
Montgomery’s argument therefore goes beyond the familiar proposition that bias is an ethical concern. Bias can affect whether technology works effectively for the people expected to use it — and consequently whether those people trust the organisation putting the technology in front of them.
That concern has also shaped TogetherAI, Better Together Agency’s communications and marketing platform. Rather than asking organisations to rely entirely on the assumptions embedded in a generic model, TogetherAI uses an organisation’s own material, language and standards as context. Montgomery describes this broader philosophy as building “born-inclusive technology”: incorporating lessons about representation and bias during product development rather than attempting to repair them only after deployment.
Yet bringing AI closer to an organisation’s own knowledge introduces another problem.
Organisations themselves are not neutral datasets.
A company’s policies, archives, brand guidelines and institutional language may contain historical assumptions, omissions and blind spots. Retrieval from proprietary material can make an AI system more organisationally specific, but specificity is not the same thing as fairness. A system can faithfully reproduce an organisation’s voice while also faithfully reproducing its exclusions.
That makes governance much more difficult than simply putting a human somewhere “in the loop”.
Who is that human? What evidence can they see? Are they empowered to challenge the system? Can they stop an automated process? What happens when an AI-generated conclusion passes through enough people that it begins to acquire institutional authority simply because everyone assumes someone else has already checked it?
These questions become increasingly important as AI moves from experimentation into infrastructure.
Montgomery’s current research is broadening the conversation further. Better Together Agency is conducting its third annual Biases in Generative AI study and is deliberately seeking responses not only from regular AI users but also from people who do not use generative AI. On LinkedIn, Montgomery has described the research as an attempt to understand bias, accuracy, representation, misinformation, trust and accountability from a broader section of the US public.
That matters because much of the public conversation about artificial intelligence is still produced by a remarkably self-selecting group: technologists, founders, investors, researchers, policymakers, journalists and highly engaged users.
Those communities are important.
They are not everyone.
A technology can affect someone without that person choosing to become an expert in it. An applicant can be evaluated by an AI-assisted process without using ChatGPT. A customer can encounter synthetic content without knowing how it was produced. A citizen can be affected by automated analysis without understanding the architecture behind it.
The legitimacy of AI systems therefore cannot depend exclusively on the opinions of those most comfortable using them.
I spoke with Catharine Montgomery about the relationship between bias and trust, when technological failures become organisational accountability problems, the limits of grounding AI in company material, what meaningful human oversight actually requires, and why responsible-AI research needs to listen to people beyond the AI-literate communities currently shaping so much of the debate.
Your research found that 33% of consumers would consider stopping use of a biased generative AI tool, while 92% believe companies must address generative AI bias. What do organisations still misunderstand about the connection between AI bias and trust?
I think organizations underestimate how personal a biased answer can feel. If a tool gets your identity or circumstances wrong, you do not separate the software from the organization that put it in front of you. AI is already part of how organizations communicate and serve people, so this is relevant now. For me, governance becomes real when someone can say, “This is wrong,” reach a person who will listen and see the organization correct it.
You come at AI bias not only as a technology issue, but through crisis communications and reputation. At what point does an AI failure stop being a model problem and become an organisational accountability problem?
CNN recently reported (https://keyt.com/politics/cnn-us-politics/2026/09/18/exclusive-us-military-had-close-call-after-using-ai-for-false-intelligence-report-sources-say/) a military near miss involving an AI-assisted intelligence report. A chatbot misidentified cargo on a Chinese ship. An analyst put the finding into a formal report, and personnel prepared to intercept the vessel. The error was caught before the planned operation.
That is the point I worry about: a wrong output can gain authority as people pass it along. The organization has to examine who checked that specific answer, who could challenge it and why it got as far as it did. AI is an extraordinary tool, but people remain accountable for the decisions made with it.
TogetherAI is built around an organization’s own materials, voice and standards rather than relying solely on a generic model’s defaults. How much can that approach reduce bias—and what risks remain if an organisation’s own source material contains blind spots?
Using an organization’s own material gives TogetherAI something real to work from: its language, standards and history. But those files can carry blind spots too. A brand guide may describe some audiences well and barely acknowledge others. A policy may repeat an assumption nobody has questioned in years.
Better Together Agency can review those materials with a client before they are uploaded. Who is represented? Who is missing? What language needs another look? Then we need to test the answers the tool produces. An output can sound exactly like an organization and still get something important wrong.
Our first two years of Biases in Generative AI survey research informed how we developed TogetherAI’s approach to bias. That is what I mean by born-inclusive technology. We brought what we were learning about people’s experiences into the development process, and we keep questioning both the source material and the output.
You’ve argued that AI should support human judgment rather than replace it. In practice, where should organisations draw that line, and how do they prevent “human in the loop” from becoming little more than a rubber stamp?
The person has to be able to question the answer. If I use AI to work on a LinkedIn post, I do not publish the first draft. I talk with the tool, reject angles that miss the point, add what I know about the audience and decide what sounds like me. It helps me explore ideas, while my experience tells me which ideas are useful. When an output could affect someone’s job, safety or access to a service, a reviewer also needs the evidence, time and authority to say no. A person clicking “approve” is not enough.
Your latest Biases in Generative AI research is deliberately reaching beyond the people who build and regularly use these systems. What are we missing by listening disproportionately to the AI-literate—and what would you most like the next wave of research to reveal?
I have to watch my own assumptions. I spend time in AI Slack groups, WhatsApp chats and conferences where people know these tools. It is easy to leave those rooms thinking everyone feels the same way.
People I speak with elsewhere in the country sometimes tell me they feel overwhelmed or do not know where to begin. Those conversations are not survey findings, but they are a reason to ask better questions. This year’s survey includes people who do not use generative AI. I want to understand their choices and what they expect from organizations using AI, because those decisions can affect them too. We are still collecting responses, so I do not want to claim what the results will show.
From Model Risk to Institutional Responsibility
An important conceptual shift runs through Montgomery’s answers.
The first era of public generative-AI criticism concentrated heavily on the model: hallucination, bias, unsafe outputs, incomplete training data, alignment and the limitations of probabilistic systems.
All of those questions remain important.
But organisations are now placing these systems between themselves and the world.
Once that happens, the relevant unit of analysis changes.
A model may generate the answer, but an organisation decides where to deploy it. Someone decides what information it can access. Someone defines the review process. Someone decides when a machine-generated conclusion is sufficiently trustworthy to enter a workflow. Someone determines whether a customer can appeal. Someone decides whether employees have the authority to challenge an automated result.
The model therefore becomes only one component in a larger socio-technical system.
This is why Montgomery’s description of governance is so useful: governance becomes real when a person can say “This is wrong” and something happens.
A policy document alone is not governance.
A human approval button is not necessarily governance either.
Meaningful oversight requires the ability to inspect evidence, question an output, understand the consequences of accepting it and, crucially, reject it.
As AI agents gain greater autonomy, that distinction will matter even more. The central question will increasingly move from: Can the AI do this? to Under what conditions should an organisation allow the AI to do this without intervention?
And then to the more uncomfortable question: Who remains accountable when it goes wrong?
Montgomery’s observation about institutional authority matters here. AI outputs do not remain static. They travel through organisations.
A generated statement becomes an analyst’s note.
The note enters a presentation.
The presentation informs a meeting.
The meeting produces a decision.
At each stage, the origin of the information becomes less visible while its perceived legitimacy can increase.
By the end of that chain, people may no longer be evaluating an uncertain machine output. They may believe they are reading an organisational fact.
That is why human judgment cannot simply exist at the final checkpoint. Critical scrutiny has to survive the entire information chain.
There is also a deeper challenge around inclusion.
Much of the AI industry understandably focuses on improving models, benchmarks and safeguards. Yet people ultimately experience artificial intelligence through interfaces, products, workplaces and institutions.
Someone who has never heard the term “large language model” can still discover that an AI system does not understand their accent.
Someone who has never experimented with image generators can still be misrepresented by synthetic media.
Someone with no interest in AI can still have their application, insurance claim, customer-service request, or employment experience influenced by an AI-assisted system.
Non-users are therefore not outside the AI economy.
Increasingly, there may be no such thing.
That is what makes Montgomery’s attempt to bring non-users into the next phase of her research worth watching. Her latest survey was still collecting responses when she answered my questions, and she is careful not to pre-empt its findings.
That restraint matters too.
At a moment when AI discourse frequently races ahead of evidence, waiting to hear what people actually say is itself a useful principle.
The next phase of responsible AI will not be solved entirely inside laboratories or engineering teams. It will also be shaped by whether the organisations deploying these technologies create meaningful mechanisms for challenge, correction and accountability.
Because as AI becomes embedded in everyday institutional life, people will increasingly judge it through something much older than technology.
Whether they can trust the people behind it.
More Interviews
Building Creative Machines has published dozens of exclusive conversations with founders, researchers, executives, artists and other leading voices shaping artificial intelligence and its impact on society.
Editorial Disclosure
This interview was conducted independently by Building Creative Machines. No payment, sponsorship, or other editorial consideration was received in connection with its publication. The views expressed are those of the interviewee.



