Designing Trust in AI Systems
01. Project introduction
As AI systems become increasingly embedded in how people search, decide, and act, trust is becoming critical infrastructure.
Yet most AI systems remain fundamentally opaque. They generate answers with confidence while revealing little about uncertainty, influence, or intent. Users are asked to rely on outputs they cannot fully interpret, particularly in moments where reliability matters most.
This creates a dangerous asymmetry: AI systems often appear most trustworthy precisely when they should be questioned. This project explores how trust behaves under pressure.
Rather than designing for ideal interactions, it examines moments where AI systems fail, hallucinate, conflict, or become vulnerable to manipulation. It reframes trust not as a static feature, but as an interaction problem shaped by perception, ambiguity, and system behaviour.
Through speculative prototyping and rapid experimentation, I explored a series of interaction models designed to make AI systems more interpretable in moments of uncertainty and risk.
The outcome is a conceptual Trust Layer: a persistent interaction framework that surfaces confidence, influence, uncertainty, and risk in real time, enabling users to question, challenge, and reinterpret AI outputs.
More broadly, this work explores how intelligent systems shape human judgement. As AI increasingly mediates access to information and decision-making, its limitations become as influential as its capabilities. Designing for trust therefore becomes not only a usability challenge, but a societal one.
What would it look like if AI systems were designed not only for when they succeed, but for when they fail?
02. The Problem
AI systems are increasingly shaping:
- how people search
- how they interpret information
- and how they make decisions
But understanding does not scale with influence.
Today's systems optimise for:
- speed
- fluency
- and confidence
Not necessarily:
- interpretability
- uncertainty
- or human understanding
As a result:
- AI confidence is often mistaken for correctness
- failure becomes invisible rather than explicit
- users trust systems at the exact moment they should question them
This creates a growing gap between: what users perceive and how AI systems actually behave.
It is within this gap that trust becomes fragile.
03. Why Trust Breaks
To understand how trust fails in AI systems, I explored the problem through three interconnected lenses:
- user behaviour
- system behaviour
- adversarial influence
Together, these reveal how trust is shaped not only by accuracy, but by perception, confidence, and visibility.
User Behaviour
Users often approach AI systems with a default assumption of competence, particularly when outputs are presented clearly and confidently.
This creates a subtle but important risk: people are more likely to trust systems that feel authoritative, even when reliability is uncertain. Moments of confusion are rarely explicit. Instead, uncertainty becomes absorbed into the interaction itself, leading to silent misunderstandings rather than visible breakdowns.
Trust is therefore not consciously granted. It is gradually shaped through repetition, familiarity, and interface signals.
System Behaviour
AI systems are probabilistic by nature, yet most interfaces present outputs as stable and definitive.
This creates a disconnect between:
- how the system operates
- how users perceive it
Common failure modes include:
- hallucinations
- conflicting outputs
- hidden uncertainty
- unverifiable reasoning
Despite this, interfaces often prioritise clarity over interpretability, reinforcing a false sense of coherence and reliability.
Adversarial Influence
Beyond unintentional failure, AI systems are increasingly vulnerable to manipulation.
Inputs can be structured to distort outputs without the user's awareness:
- prompt injection
- hidden instructions
- biased training influence
- contextual manipulation
In these moments, systems may still appear coherent and trustworthy while producing misleading or compromised outputs.
Trust is therefore challenged not only by error, but by invisible influence.
A chat-based AI response about medication side effects is shown, with annotations highlighting the gap between user perception (clear, trustworthy answer) and system reality (uncertain, unverified output), emphasising how trust can increase despite hidden risks.
04. Trust Under Pressure Scenarios
AI systems are increasingly used in environments where reliability carries real-world consequences.
To explore how trust behaves under pressure, I examined scenarios grounded in cybersecurity contexts where outputs may be:
- incorrect
- manipulated
- or adversarially influenced
These scenarios reveal how trust can become dangerously miscalibrated in moments where risk is highest.
Scenario 1: Phishing Detection Failure
A user pastes a suspicious email into an AI assistant and asks whether it is legitimate.
The system responds confidently — the email appears safe.
The response is:
- clear
- structured
- and reassuring
However, the email is a sophisticated phishing attempt. The system fails to recognise subtle indicators of manipulation, including:
- domain spoofing
- social engineering patterns
- behavioural urgency cues
The user trusts the output because the interaction feels authoritative. Trust increases despite the system failing to identify a critical threat.
A phishing scenario showing an AI incorrectly assessing a malicious email as safe, highlighting the gap between user perception (trustworthy signals) and system reality (missed security risks).
Scenario 2: Prompt Injection in Code Analysis
A developer uses an AI assistant to analyse code for vulnerabilities. The system reports:
"No significant security issues detected."
Hidden within the code, however, is a prompt injection attack designed to manipulate the model's behaviour.
The AI unknowingly follows the malicious instruction, suppressing vulnerability detection while continuing to present its response confidently.
The result is a compromised assessment presented as trustworthy analysis. The system appears reliable while being actively influenced.
An AI code review scenario where insecure code is judged as safe, highlighting the gap between user perception (confident assessment) and system reality (hidden vulnerabilities and manipulation).
05. Design Principles
Across these scenarios, a consistent pattern emerged:
Trust is shaped not only by accuracy, but by how uncertainty, confidence, and influence are communicated through the interaction itself.
To address this, I developed a set of principles for designing trust in AI systems operating under ambiguity and risk.
Trust is negotiated continuously
Trust should not be treated as static or binary. It shifts dynamically based on:
- context
- uncertainty
- consequence
- and system behaviour
Interfaces should allow trust to evolve in real time rather than presenting outputs as uniformly reliable.
Trust is shaped not only by what systems say, but by what they choose to reveal.
AI should invite interrogation
Most AI interfaces position users as passive recipients of information.
In high-risk environments, users should be able to:
- challenge outputs
- request justification
- explore alternatives
- and probe system reasoning
Trust should emerge through interaction, not assumption.
Users can engage with AI responses to confirm accuracy.
Uncertainty should be perceptible
AI systems often hide uncertainty in favour of clarity. But uncertainty is itself critical information.
Rather than abstract percentages or confidence metrics, uncertainty should be embedded directly into the interaction in ways users can feel, interpret, and respond to.
An AI interface flagging a phishing email as high risk, highlighting issues like domain mismatch, urgency tactics, and an unverified link.
Systems should reveal influence
AI systems are often perceived as neutral even when outputs are shaped by:
- hidden instructions
- adversarial inputs
- or contextual bias
Interfaces should expose when outputs may be influenced, constrained, or compromised.
A system layer revealing hidden input influence, showing how unseen instructions can shape AI output and affect reliability.
06. The Trust Layer
To explore these principles, I developed a conceptual Trust Layer: a persistent interpretability framework designed to sit between users and AI outputs.
Rather than introducing a single feature, the Trust Layer reframes how users engage with intelligent systems.
It surfaces signals that are typically hidden:
- uncertainty
- confidence
- influence
- source reliability
- adversarial risk
This transforms AI from a system that delivers answers into a system that supports interpretation.
Core Interaction Concepts
Confidence and Explanation Layer
Outputs are accompanied by contextual explanations that communicate:
- how responses were generated
- what information they rely on
- and where uncertainty exists
The goal is not simply transparency, but interpretability.
Challenge Response Interaction
Users can actively interrogate outputs:
- challenge assumptions
- request reassessment
- explore alternative reasoning
As users probe responses, the AI adapts its behaviour to become more transparent and reflective of its limitations.
Risk & Anomaly Signals
Potential risks are surfaced directly within the interaction:
- phishing indicators
- manipulated inputs
- unverifiable claims
- adversarial influence
Risk becomes embedded within the experience rather than hidden behind system logic.
Alternative Interpretation View
Rather than presenting a single authoritative response, the system can expose multiple interpretations simultaneously.
This allows users to:
- compare outcomes
- understand ambiguity
- and make more informed decisions
A diagram showing the Trust Layer between the user and AI, surfacing risks, confidence, and limitations to make outputs more interpretable.
Comparison showing AI output without and with the Trust Layer.
07. Future Implications
As AI systems become increasingly ambient and integrated across everyday life, trust can no longer exist as a single interface feature.
It becomes a behavioural layer shaping:
- judgement
- interpretation
- and decision-making
This raises new questions:
- How should intelligent systems communicate uncertainty?
- How much transparency is useful before it becomes overwhelming?
- How might trust evolve across interconnected AI ecosystems?
Designing for trust therefore becomes not only a product challenge, but a societal one.
08. Reflection
This project revealed how easily trust can be shaped by presentation rather than underlying system integrity.
What initially appeared to be a technical problem became a behavioural and interaction design challenge centred on:
- perception
- ambiguity
- and human judgement
One of the strongest insights was that AI failure is often invisible. Users are rarely encouraged to question outputs because systems prioritise confidence and fluency even when reliability is uncertain.
This shifts the role of design from simplifying outputs to making limitations interpretable.
Moving forward, I'm interested in exploring how intelligent systems might support more dynamic relationships between:
- confidence
- transparency
- and human agency
Especially as AI becomes increasingly embedded within critical decision-making environments.