About this trial
Patients increasingly consult artificial intelligence (AI) chatbots such as ChatGPT for health information before clinical visits, yet the impact of an actual orthopedic consultation on patient trust in AI-derived information remains unknown. This prospective longitudinal observational study quantifies how a single orthopedic outpatient consultation modifies patient trust in AI chatbots, the concordance between AI-derived and physician-delivered information, and patient anxiety, using a paired pre-post survey design supplemented by a matched physician-side assessment. Adult patients (18 years and older) presenting to two orthopedic outpatient clinics in Cyprus complete a brief pre-consultation questionnaire (T0) capturing demographics, AI use patterns, prior AI consultation regarding the current complaint, baseline trust, expectations, and anxiety. Immediately after their consultation they complete a second questionnaire (T1) assessing concordance with physician advice, trust change, consultation facilitation, post-consultation anxiety, and future intention. The consulting physician completes a brief 30-second post-visit form capturing whether AI was discussed, the medical accuracy of AI-derived information conveyed by the patient, and the effect of the AI discussion on consultation duration. The primary outcomes are the paired within-patient change in AI trust between T0 and T1 and physician-patient concordance on AI versus physician advice. Target enrollment is 180 to obtain 150 paired completed assessments.
Eligibility criteria
This trial does not accept healthy volunteersQualifiers
Age 18 years or older
Presenting to an orthopedic outpatient clinic for any consultation
Able to read and respond to a Turkish-language questionnaire
Provides informed consent
Disqualifiers
Inability to complete a self-report questionnaire (e.g., severe cognitive impairment, language barrier)
Re-presentation within the same recruitment window (each patient is enrolled only once)
Refusal of consent for either T0 or T1
Trial population
Consecutive adult patients (18 years and older) presenting to the participating orthopedic outpatient clinics during the recruitment window who consent to participate. No specific orthopedic diagnosis is required.
Trial design
Cohort
Prospective
Treatments tested in this trial
Not listed
Trial groups
Trial outcomes
Primary outcomes
Mean within-patient change in self-reported trust in artificial intelligence-derived health information, measured by a study-specific 5-point Likert item (T0.11) and a study-specific 3-level categorical change item (T1.4).
Trust in AI-derived health information is assessed pre-consultation by a study-specific single-item 5-point Likert scale (item T0.11: "How much do you trust the AI's answer?"; anchors 1 = not at all, 5 = completely), administered only to patients who reported pre-consultation AI use (item T0.9 = Yes). Post-consultation, trust change is reassessed by a study-specific 3-level categorical item (item T1.4: increased trust / unchanged / decreased trust). For paired analysis, the post-consultation score is derived by mapping T1.4 categories to integer shifts (+1 / 0 / -1, with floor 1 and ceiling 5) relative to T0.11. Unit of measure: Likert score points on a 1-5 scale (continuous derived score) and proportion of patients per 3-level category. Primary analysis: paired Wilcoxon signed-rank test on the derived continuous score; sensitivity analysis: McNemar test on the 3-level categorical change.
Patient-physician concordance on artificial intelligence-versus-physician medical advice agreement, measured by Cohen's kappa coefficient between a study-specific 4-category patient item (T1.2) and a study-specific 5-point physician-rated AI medical accu
Concordance is assessed by Cohen's kappa coefficient comparing patient-reported AI-physician concordance (item T1.2: fully concordant / partially concordant / discordant / physician did not address; dichotomized to concordant vs. non-concordant) and physician-reported AI medical accuracy (item H2: 5-point Likert anchored 1 = entirely incorrect to 5 = entirely correct; dichotomized at ≥ 3 as concordant). Unit of measure: kappa coefficient (range -1 to +1) with 95% confidence interval, and percentage of dyads classified as concordant on each instrument.
Secondary outcomes
Mean within-patient change in self-reported anxiety, measured by an 11-point 0-to-10 visual analogue scale anchored 0 = no anxiety and 10 = worst possible anxiety (items T0.14 baseline, T1.5 post-consultation).
Anxiety is measured pre-consultation (item T0.14) and post-consultation (item T1.5) using the same 0-to-10 visual analogue scale. Within-patient change is calculated as T1.5 minus T0.14. Unit of measure: scale points (range -10 to +10). Analysis: paired t-test with Wilcoxon signed-rank as sensitivity analysis; Cohen's d effect size reported.
Percentage of enrolled patients reporting pre-consultation artificial intelligence use for the current orthopaedic complaint, measured by a study-specific single-item yes/no question (T0.9).
Proportion of enrolled patients responding "Yes" to item T0.9 ("Before today's appointment, did you ask an AI chatbot a question about this health concern?"). Unit of measure: percentage of participants, reported with exact (Clopper-Pearson) 95% confidence interval.
Percentage of pre-consultation artificial-intelligence users whose physician independently confirmed that AI was raised during the consultation, measured by a study-specific yes/no physician item (H1).
Among patients responding "Yes" to T0.9, the proportion in whom the treating physician independently reported "Yes" to item H1 ("Did the patient raise AI during this consultation?"). Unit of measure: percentage of patients with exact 95% confidence interval.
Percentage of consultations in which the physician reported that the artificial-intelligence discussion shortened, did not change, or prolonged the encounter, measured by a study-specific 3-category physician item (H3).
Among consultations in which the patient raised AI (H1 = Yes), the physician's categorical rating of effect on consultation duration (H3: "shortened" / "no change" / "prolonged"). Unit of measure: percentage of consultations per category (descriptive).
Other outcomes
Exploratory association between categorical post-consultation trust change and demographic predictors, estimated by multinomial logistic regression with the study-specific 3-level trust change item (T1.4) as the outcome and age band, sex, education level
Multinomial logistic regression model: outcome = T1.4 (decreased / unchanged / increased trust, reference category = unchanged); predictors = age band (5-level), sex (3-level), education (5-level), employment status, weekly internet-use frequency. Unit of measure: adjusted odds ratios with 95% confidence intervals.
Internal consistency of a four-item artificial-intelligence trust subscale, measured by Cronbach's alpha across items T0.11 (baseline trust), T1.4 (post-consultation trust change, linearly recoded), T1.7 (future-use intention), and T1.8 (recommendation
Cronbach's alpha is estimated on the final analytic sample using the four trust-related Likert items listed. Unit of measure: alpha coefficient (range 0 to 1) with bootstrap 95% confidence interval.
Sponsors and contacts
Click on the lead sponsor to view all of their trials.
Utku Gürhan
Lead sponsor
University of Kyrenia
Sponsor institution