All research

RESEARCH / 2026

Effects of Evidence Grounding and Source Attribution in Generative AI Tutors: A 2×2 Factorial Study

A controlled study with 36 university participants investigating how grounding and attribution interact in perceived trust, user experience, and learning.

Authors: Bingxu Hou, Luocheng Xie, Lijie Zheng

My role
First author · Study lead
Status
Submitted to IHCI 2026 · Decision pending
Period
Mar – Sep 2026

Research and role statuses reflect the source record as of 29 August 2026. Submitted does not mean accepted.

2×2 factorial design36 participantsUEQ / S-TIAS

Turning a product decision into a question

Building AI-PPT Tutor raised a question that the interface alone could not answer: does a tutor change people’s experience because its answers are grounded in course evidence, because it displays sources, or because of the combination? I wanted to separate evidence grounding from visible source attribution, instead of treating a citation interface as proof that the underlying system was helpful.

Controlling the system before testing it

We used a 2×2 between-subjects factorial design. Thirty-six university students were randomly assigned to baseline, grounding-only, attribution-only, or combined conditions, with nine in each group. The experimental build used React/Next.js and qwen-plus with a fixed, manually verified evidence index. Keeping this index separate from the later OCR/vector-retrieval product pipeline reduced retrieval uncertainty during the manipulation.

The system generated structured claims and attached citations only when retrieved evidence directly supported the full claim, with page navigation and highlighted passages. This made evidence behaviour something we could manipulate and inspect, not simply a visual decoration.

Technical and methodological focus: control what differs between conditions, and make it inspectable.

My role from protocol to analysis

As first author and study lead, I personally hosted the experimental workflow and collected participant data and interviews. My work covered ethics materials, the learning quiz, UEQ user-experience measures, S-TIAS trust measures, manipulation checks, the interview protocol, and factorial analysis. I worked on the experimental system as a co-creator and developer alongside Luocheng Xie; the manuscript authors are Bingxu Hou, Luocheng Xie, and Lijie Zheng.

What the findings support

Grounding × attribution showed a significant interaction on perceived trust (β = 2.52, 95% CI [0.77, 4.27], p = .019). We did not find significant effects on UEQ or learning performance. The combined condition’s higher mean quiz score was descriptive, not evidence that the system improved learning.

The study is complete and the manuscript was submitted to IHCI 2026; the decision remains pending in the 29 August 2026 record. For me, the work connects implementation with disciplined evaluation: a working feature is a starting point for a question, not the answer to it.

Explore the related project

AI-PPT Tutor

NEXT / LET’S TALK

Let’s take the conversation further.

Interested in the technical work, a research question, or collaborating? I’d be glad to talk.

Let’s connectseverushou@outlook.com