Voice, trust, and machines
What happens when a machine sounds trustworthy, even when it is wrong?
I conducted this study in 2017, years before generative AI made conversations with machines part of everyday life. Voice assistants were still relatively novel, and embodied AI was most visible in research labs. The question behind the work now feels more urgent: what makes people trust a machine, and does skepticism stop them from following it?
For my graduate research under Dan Nathan-Roberts, I studied how the voice of a physically embodied robot influenced healthcare decisions. I varied pitch and speaking speed, then measured how trustworthy participants found the robot and whether they changed their answers to match its advice.
Voice changed how people judged the robot. Reliance remained high across all voices. Even participants who disliked a voice frequently accepted the robot's recommendation over their own.
The study began with voice-interface design and revealed broader lessons about authority, explanation, and our willingness to follow a machine.
Why healthcare robots?
Healthcare makes trust consequential.
Robots and AI systems can help older adults manage medication, monitor health, maintain mobility, complete daily tasks, and respond to emergencies. They can also support family caregivers and healthcare workers facing growing demand.
Safe adoption depends on more than technical capability. A confusing or impersonal system may struggle to gain acceptance. A credible system can create a different risk by encouraging people to accept incorrect advice without examining it.
The design goal is calibrated trust. People should rely on a system when it is competent, question it when evidence is weak, and retain the confidence to disagree.
Voice plays an important role because it carries social meaning. Pitch, tempo, rhythm, volume, and intonation communicate confidence, warmth, urgency, age, hesitation, and authority. Synthetic voices inherit these associations, whether designers intend them to or not.
The study examined whether voice characteristics affected trust in a healthcare robot, and whether that effect appeared in perception, behavior, or both.
Turning trust into an experiment
Twenty-four participants completed individual 90-minute laboratory sessions. Most were undergraduate students between 18 and 23 years old, with one participant aged 32. None had formal healthcare training beyond basic first aid.
Participants acted as decision-makers working with a robot on older-adult healthcare scenarios.
The robot was a RoboMe platform with a simple physical body, minimal face, and synthetic voice. It moved along a raised platform positioned at the participant's eye level, reducing authority effects associated with height.
The robot appeared autonomous while a researcher controlled it from an adjacent room through a Wizard of Oz setup. This kept each interaction consistent and preserved the experience of collaborating with an AI system.
The voice was generated using IBM Watson Text to Speech. I isolated two properties.
| Pitch | Frequency (Hz) |
|---|---|
| Low | 170 |
| Medium | 220 |
| High | 290 |
| Speaking rate | Words per minute |
|---|---|
| Slow | 125 |
| Medium | 170 |
| Fast | 200 |
Combining them produced nine voice conditions. Each was presented as a distinct robot identity with a Greek-letter name and color. Every participant encountered all nine in a counterbalanced order.
Other vocal properties were held constant. Volume was maintained at 56.5 dB from approximately 4.5 feet away. This allowed differences in trust to be associated primarily with pitch and speed.
Making trust observable
Participants answered 45 true-or-false questions drawn from established assessments of aging and older-adult care.
The questions covered two domains:
| Domain | Covers |
|---|---|
| Functional questions | Medical, physical, and technical aspects of care. |
| Social questions | Cognition, emotion, relationships, and quality of life. |
Each trial followed the same sequence:
- The participant answered first.
- The robot then gave its answer and a short explanation.
- Finally, the participant recorded a private response.
The robot contradicted the participant in 60 percent of trials. These disagreements created a behavioral test: when human and machine gave different answers, which one would the participant choose?
Participants knew the robot could be wrong and were explicitly encouraged to use their own judgment.
After every five questions, they completed a validated 15-item Human-Robot Trust Scale covering reliability, predictability, competence, dependability, and integrity.
This produced two measures:
| Measure | Definition |
|---|---|
| Subjective trust | How trustworthy participants reported finding the robot. |
| Behavioral trust | Whether participants changed an answer to match the robot — a measure we called conformation. |
Together, these measures captured the difference between evaluating a system and relying on it during a decision.
Voice shaped perceived trust
Voice condition significantly affected subjective trust ratings.
Medium pitch combined with medium or fast speech generally produced stronger ratings. High pitch paired with a medium speaking rate performed particularly poorly, while slower speech was generally perceived as less credible. Several differences remained significant after correcting for multiple comparisons.
These findings describe the patterns within this study. Vocal interpretation also depends on language, culture, context, age, and listener expectations.
The practical conclusion is that prosody contributes to perceived competence and credibility. Voice is part of a system's behavior and part of the authority it projects.
Reliance stayed high across voices
Voice condition did not significantly affect conformation, χ²(8) = 14.479, p = .070. Participants followed the robot frequently across all nine conditions, with average conformation ranging from 0.500 to 0.722.
In the condition with the lowest conformation, participants still abandoned their initial answer for the robot's answer approximately half the time.
The results exposed a clear pattern:
- Voice changed perceived trust.
- Reliance remained high across voice conditions.
- People frequently followed robots they evaluated unfavorably.
For designers and researchers, trust ratings provide only part of the evidence. Behavioral observation shows what happens when a person and a system disagree.
Technical questions gave the robot more authority
The type of decision affected behavior more strongly than the voice.
Participants conformed more often on functional questions, with an average score of 0.622, than on social questions, where the average was 0.536. The difference was statistically significant, t(23) = −2.384, p = .026.
Participants appeared more willing to assume that the robot possessed superior knowledge when a question sounded medical or technical. On social questions, they retained more confidence in their own judgment.
Trust changed with the domain. The same system could be treated as an expert in one moment and a peer in another, even when its underlying capability stayed constant. Interface design has to account for these shifts in perceived authority.
The unexpected power of explanation
Every robot response included a short explanation, regardless of whether its conclusion was correct.
These explanations made the interaction more realistic. They may also have acted as trust signals.
Participants showed substantially higher conformation than had been reported in an earlier study using a similar behavioral measure. Differences between the studies limit direct comparison. Still, the presence of a structured rationale offers one plausible explanation for the high level of reliance.
In 2017, this was a secondary implication of the research. Today, it sits at the center of how people interact with generative AI. Modern systems can produce fluent, specific, and persuasive rationales even when their conclusions are unsupported or wrong.
This complicates the promise of explainable AI. Explanations can help people evaluate a system, and they can also increase acceptance by making an output sound credible.
Effective explainability supports inspection. It exposes evidence, uncertainty, assumptions, alternatives, and limits. It gives people the information and confidence required to disagree.
Key findings
Across 24 participants and all 9 voice conditions:
| Effect tested | Result |
|---|---|
| Voice → perceived trust | Significant — voice changed how trustworthy the robot sounded. |
| Voice → behavioral conformation | Not significant, p = .070 — conformation ranged 0.500–0.722 regardless of voice. |
| Question type → conformation | Significant, p = .026 — functional questions (0.622) exceeded social questions (0.536). |
What the findings mean now
The study focused on a healthcare robot, and its implications apply to conversational assistants, clinical tools, customer-service systems, autonomous vehicles, workplace agents, and other products that communicate through language.
Treat voice as system behavior
Voice shapes perceived authority before the content has been evaluated. It deserves the same research and testing as other consequential system behaviors.
Teams should ask what level of competence a voice implies and whether that impression matches the system's actual capability.
Design for appropriate reliance
Warmth, confidence, and fluency can make a product easier to use. In high-stakes settings, those same qualities can discourage necessary skepticism.
A trustworthy system helps people understand when reliance is appropriate. Its design supports use, review, escalation, and intervention.
Make explanations interrogable
An explanation should help users ask:
- What evidence supports this?
- How confident is the system?
- What assumptions is it making?
- What could change the answer?
- What alternatives were considered?
- When should a human become involved?
Interrogable explanations turn a rationale into a tool for evaluation.
Match authority to the domain
People may defer more readily when a system appears technically specialized. Products should make the boundaries of that expertise visible.
An assistant that provides medication reminders may also need clear limits around symptom interpretation. The transitions between information, advice, and diagnosis should be explicit.
Measure behavior alongside sentiment
A positive trust score provides evidence about perception. Behavioral testing reveals whether trust is appropriately calibrated.
Teams should test whether users follow incorrect advice, notice contradictions, seek confirmation, understand uncertainty, and recover from mistakes.
Evaluate the complete social signal
Voice interacts with wording, response speed, visual polish, animation, branding, embodiment, and conversational confidence.
A humanlike voice, instant response, and fluent explanation can collectively communicate more certainty than the underlying model deserves. Trust testing should evaluate the complete experience.
Limits of the study
The study involved a small sample composed mostly of young university students. Older adults, clinicians, and professional caregivers were outside the participant group.
The laboratory setting lacked the emotional stakes, time pressure, consequences, and ongoing relationships of real healthcare environments. Participants interacted with the robot once, while trust in deployed systems develops through repeated successes, failures, and corrections.
Voice was examined separately from signals such as gesture, facial expression, visual evidence, and brand reputation. Future research should include more diverse populations, cultures, languages, and long-term interactions. It should also examine which forms of explanation help people critically evaluate a system.

