Abstract
Ethics education in medical training remains difficult to standardize and sustain. Many curricula still rely heavily on didactic teaching rather than immersive ethical reasoning. Although large language models (LLMs) can generate structured analyses of ethical dilemmas, they are not designed to facilitate embodied, conversational engagement that mirrors real-world ethics discussions. We developed CALEB (Conversational Agent Learning Ethics Bot), a domain-specific, extended reality (XR)-enabled conversational agent designed to simulate pragmatic, case-based moral deliberation. CALEB integrates a curated medical ethics knowledge base, structured persona design, and bounded generative architecture to promote focused, dialogical interaction. We conducted a two-phase feasibility evaluation. In phase 1, CALEB and GPT-4o accessed through the ChatGPT interface were compared using standardized transcript outputs generated from matched medical ethics prompts. In phase 2, medical ethicists interacted with CALEB in a live XR setting and provided formative post-session feedback. GPT-4o generated substantially longer and comprehensive responses. In contrast, CALEB produced significantly shorter but more principle-dense responses and was the only system to consistently demonstrate first-person and emotionally interpretive language. Computational analysis revealed higher empathy scores for CALEB in interpretive dimensions (p < 0.001). In the live XR phase, post-session feedback suggested that CALEB was perceived more favorably in an interactive, embodied setting than in transcript-only review. This study provides evidence that domain-specific agents like CALEB are feasible and pedagogically distinct from general-purpose LLMs. Foundational and specialized systems may serve complementary roles in advancing scalable, interactive medical ethics education.
Keywords
Introduction
Ethics is a core element of medical practice, shaping clinicians’ responsibilities, guiding patient care, and shaping the education of future physicians. 1 Ethics education requires physicians to not only learn ethical principles, but also to engage in applying them in practice. Thus, ethics is taught through multiple paradigms. One established model is the epistemological approach, where learners acquire knowledge of ethical theories and principles through didactic instruction. Another is the pragmatic-hermeneutical, dialogical approach, which emphasizes moral case deliberation with competing perspectives and values and fosters reflective judgment. 2 Despite its importance, ethics training in medical education remains inconsistent, with a scarcity of dedicated resources, and significant variation in curricular emphasis, standardization, and reinforcement.3–5
Advancements in generative artificial intelligence (AI), particularly through natural language processing (NLP) and large language models (LLM)s, raise questions on whether AI can close the gaps in medical ethics training. 6 LLMs are powerful tools for interpreting semantics, engaging complex prompts, and generating naturalistic dialogue, all of which may allow for practice in ethical reasoning. While early trials suggest LLMs may improve performance in clinical assessments such as Objective Structured Clinical Examinations (OSCEs), evidence for their use in ethics education remains limited. 7
Most commercially available LLMs, such as GPT-4o, are not designed with pedagogical needs in mind. They are trained on vast general-domain corpora (e.g., “Common Crawl,” Wikipedia, professional exams)8,9 but lack domain-specific fine-tuning for health care ethics, mechanisms for weighting clinically authoritative sources, or the capacity to manifest native empathy. 10 In response, we developed CALEB (Conversational Agent Learning Ethics Bot), an artificially intelligent conversational agent (AICA) designed to support the pragmatic-hermeneutical approach to ethics instruction. Leveraging prior work in extended reality (XR) AICAs and guidelines for the safe and ethical creation of AICAs, we designed CALEB as a holographic avatar to enhance presence, interactivity, and engagement in moral case deliberation.11–13
Here, we conducted a two-phase feasibility evaluation of CALEB for medical ethics education. First, we performed a transcript-based comparison between CALEB and GPT-4o accessed through the ChatGPT interface using matched medical ethics prompts. Second, we gathered formative expert feedback after live interaction with CALEB in an XR environment. This design allowed us to examine transcript-level differences between a domain-specific conversational agent and a general-purpose LLM, while also exploring how CALEB was perceived when used in the embodied, interactive format for which it was designed.
Methods
Ethics statement
This research involved neither human subjects in the sense of regulated research (45 CFR 46), nor collection of identifiable information from individuals, thus exempt from IRB review. No human subjects consent was necessary, as no individuals were studied, harmed, or had their data collected beyond voluntary participation of expert reviewers in evaluation activities that involved no personal data collection.
Study design
This study was designed as a two-phased mixed-methods feasibility evaluation. Phase 1 consisted of transcript-based comparative analyses of CALEB and GPT-4o accessed through the ChatGPT interface. In this phase, both systems received identical text prompts across 24 medical ethics topics, generating 48 total transcripts for quantitative transcript coding, computational empathy classification, and blinded expert transcript review. Phase 2 consisted of a formative live interaction assessment of CALEB alone within the XR application. In this phase, medical ethicists interacted with CALEB in real time and provided post-session qualitative feedback. Because GPT-4o was not evaluated under matched live XR conditions, Phase 2 was not treated as a direct comparative evaluation between systems.
Development of CALEB
CALEB was engineered as a domain-specific generative conversational agent tailored for medical ethics pedagogy. The system was constructed using Convai, a platform for avatar-based dialogue generation. The development team reviewed previous methodology by Sardesai et al. to leverage a hybrid architecture by combining foundational LLMs (Claude 3.7 Sonnet) with customizable knowledge bases and narrative scripting, enabling the creation of contextualized interactive agents. 14
CALEB was configured to simulate a medical ethics professor persona, grounded in a pragmatic-hermeneutical and dialectical framework. The agent’s dialogic behavior was designed to support recursive ethical reasoning, case-based discussion, and affective engagement. To achieve this, CALEB’s backend incorporated:
Custom Knowledge Embedding: A curated corpus of medical ethics literature, case studies, and pedagogical prompts was uploaded via a retrieval-augmented generation (RAG) system using curriculum outlines from Weill Cornell Medicine’s ethics course. This corpus was structured using semantic tagging and hierarchical topic modeling to facilitate dynamic retrieval and context-aware response generation with model access to a specific reference list (Supplementary Appendix SA1).
Narrative Scripting and Persona Modeling: CALEB’s persona was instantiated via a structured narrative scaffold, including specific pre-prompting professional background as an ethics professor, teaching philosophy, and rhetorical style. This scaffold was encoded using a no-code scripting interface, allowing for conditional branching, topic-sensitive dialogue transitions, and safety guardrails based on previous guidelines to prevent agent use of explicit language or engage in dialogue related to violent, sexual, self-harm, and hate/harassment content. 11
Multimodal Interaction Design: The agent supported both text and speech-based input/output modalities. Natural language understanding and generation pipelines were integrated with Convai’s speech-to-text and text-to-speech modules, enabling real-time conversational flow and embodied interaction through a 2D or 3D avatar interface.
To ensure epistemic reliability and mitigate hallucination risks inherent in LLMs, CALEB’s responses were constrained by a bounded generative framework. This involved:
Prompt Engineering and Response Filtering: Initial content generation was seeded using Claude 3.7 Sonnet with domain-specific prompts and reduced temperature settings (0.49/1) to allow hierarchical preference of the initial curricular corpus.
Fallback and Refusal Mechanisms: When queried outside its domain or narrative scope, CALEB was pre-prompted to employ refusal strategies, explicitly acknowledging limitations in its knowledge base—thus preserving pedagogical integrity and transparency.
The agent was deployed in an XR application created with Unity for the Meta Quest 2/3 headset. The avatar was designed with expressive facial animations, gaze tracking, and spatial audio cues to simulate human-like presence and responsiveness. The virtual environment was also passthrough-enabled, allowing users to see CALEB as a 3D hologram in their real-life environments or in a blacked-out virtual reality environment.
Ethics topics and AI responses
The authors (R.J. and J.R.) comprised ethical simulations in 24 topics based on curricular consensus and generated transcripts in response to specific queries on each topic with each of the AI models (Supplementary Appendix SA2).
Although Supplementary Appendix SA1 contains 25 subtopics specific to CALEB’s RAG, one subtopic (AI and algorithmic bias in clinical decision support) was prospectively excluded prior to simulation generation. Prompting AI systems to evaluate the ethical appropriateness of AI itself was judged to risk reflexive or self-referential responses that might compromise the validity of the larger comparison. Therefore, 24 topics were used for all transcript analyses.
Each transcript was then quantitatively evaluated for fixed content metrics, including word count, paragraph count, as well as coded for ethical principles referenced, citations referenced, and instances where first-person language, empathetic language, or non-verbal stage directions were used. Shapiro-Wilk tests were performed for normality across each variable measured (Supplementary Appendix SA3).
Computational empathy analysis was conducted using a publicly available machine learning model specifically developed to assess empathetic communication in text-based mental health support contexts based on the methodology detailed in the Sharma et al. 15 Although the classifier was originally developed for text-based peer-support interactions in mental health contexts, it was used here as an exploratory tool to characterize empathy-related linguistic patterns in AI-generated medical ethics transcripts. The classifier outputs were therefore interpreted as linguistic markers rather than direct measures of empathetic capacity or educational effectiveness.
The model was designed to quantify three core dimensions of empathy: emotional reaction (ER)—expressing emotions of compassion, concern, and warmth toward the seeker, interpretations (IP)—communicating an understanding of feelings and experiences of the seeker, and explorations (EX)—exploring feelings and experiences to improve understanding of the seeker. Each component was scored on a three-point ordinal scale—0 is no communication, 1 is weak communication, and 2 is strong communication. To account for model stochasticity, such as variability in model outputs due to random factors in training, each of the three empathy classifiers (ER, IP, and EX) was independently retained and applied three times per transcript, using the original Reddit-based dataset provided in the GitHub repository. This retraining approach helps to account for output variability by treating each run as an independent estimate.
Each transcript thus received three labels per empathy dimension, one from each independent model run. A majority voting procedure was used to determine a final consensus label for each dimension: the most frequently occurring label across the three runs was selected as the final empathy score for that transcript. For transcripts consisting of multiple interactions (several messages exchanged), the maximum of the consensus scores across all interactions was taken as the overall empathy rating for that transcript. This ensured that peak expressions of empathy were captured in the final analysis. Finally, these consensus empathy scores were joined with metadata indicating the AI model responsible for each transcript.
Expert analysis
A structured evaluation form was developed to assess the quality and pedagogical value of the AI-generated transcripts. The form included seven questions rated on a five-point Likert scale (Supplementary Appendix SA4). Three medical ethicists (authors P.G., D.M., and E.G.) who were co-authors but were not involved in the development of the CALEB model independently reviewed all 48 randomized transcripts while blinded to the AI model that generated them. Since the reviewers were from the same institution and served as co-authors, this component was treated as internal formative expert feedback rather than independent external validation. Following the transcript evaluations, ethicists participated in a live interaction session with CALEB. Following their interactions, they completed free-text response forms to further evaluate their experience.
Written feedback was analyzed using thematic analysis, following a widely used stepwise approach that began with conducting line-by-line inductive coding to capture salient ideas in participants’ written responses. Codes were recorded in a non-mutually exclusive manner to ensure that overlapping constructs were retained and then grouped into higher-order categories and themes to reflect recurrent patterns. After generating themes inductively, alignment was established with evaluation frameworks for conversational agents to increase comparability and rigor. Specifically, quality frameworks for intelligent conversational agents were previously described by Radziwill and Benton, 16 which emphasize functionality, humanity, affect/empathy, ethics, and accessibility. A secondary taxonomy for health care chatbot evaluation proposed by Abbasian et al., 17 was also utilized, which organizes metrics into four domains: accuracy, trustworthiness, empathy, and performance. Combining both frameworks, codes were re-examined deductively to ensure coverage of six hybrid domains: Performance, Humanity, Affect/Empathy, Accessibility, Accuracy, and Trustworthiness. Finally, a cross-comparison analysis was conducted between the qualitative themes from the live sessions and the previous quantitative ratings of the transcripts to identify areas of convergence and divergence. This multi-stage process allowed us to capture both emergent participant perspectives and theoretically grounded dimensions of conversational agent evaluation (Supplementary Appendix SA5).
Statistical analysis
Statistical analyses were conducted in RStudio. Normality was assessed using Shapiro-Wilk testing, supplemented by visual inspection of Q-Q plots and histograms (Supplementary Appendix SA3). Since CALEB and GPT-4o responses were generated from the same set of 24 standardized prompts, phase 1 transcript comparisons were treated as matched by prompt. Because several transcript variables were non-normally distributed and each CALEB response was matched to a GPT-4o response generated from the same prompt, paired Wilcoxon signed-rank tests were used for transcript and ratio metrics. Ratio analyses excluded matched prompt pairs where the denominator was zero or undefined. Computational empathy labels were treated as paired ordinal outcomes and analyzed descriptively with exploratory paired non-parametric comparisons where testable. Expert transcript ratings were ordinal Likert-scale data and were compared using paired Wilcoxon signed-rank tests within each ethicist and rating item. Given the small single-center panel of three ethicists, expert rating p-values were interpreted as exploratory and formative rather than confirmatory.
Results
The CALEB program was deployed as a dual virtual reality or augmented reality application using a 3D human avatar with expressive facial animations, gaze tracking, and spatial audio (Fig. 1A). Prior to launching the XR version of the application, initial outputs of CALEB were tested as a chatbot application similar to the ChatGPT interface (Fig. 1B). CALEB and GPT-4o transcript outputs were then evaluated using three complementary modalities: (1) quantitative transcript analysis of 48 standardized ethics discussions, (2) computational assessment of empathy using a validated NLP model, and (3) expert review by medical ethicists, including both blinded transcript evaluation and live interaction with CALEB in the XR application.

Phase 1: Transcript-based comparative analyses
Quantitative transcript features
In the phase 1 analysis, quantitative transcript coding was performed across 48 transcripts generated from 24 matched medical ethics prompts, with one CALEB and one GPT-4o transcript per prompt. This was independently performed by three authors (R.J., J.R., and V.B.). CALEB produced shorter responses than GPT-4o (median 172 words vs. 2558 words, respectively; p < 0.001), with fewer paragraphs, fewer ethical principles referenced and fewer citations or references (Table 1). When normalized to word count, however, CALEB referenced ethical principles more densely than GPT-4o (34.5 vs. 266.7 words per ethical principle; p < 0.001) and incorporated citations or references more densely (175.0 vs. 238.1 words per citation/reference; p = 0.0042). Only CALEB used first-person language (median 3 vs. 0; p < 0.001) and emotional or empathetic phrasing (median 4 vs. 0; p < 0.001).
Phase 1 Matched-Prompt Transcript Comparison of CALEB and GPT-4o
Values are reported as median (IQR) across matched prompt pairs. Each prompt generated one CALEB transcript and one GPT-4o transcript. P-values were generated using paired Wilcoxon signed-rank tests. Ratio metrics represent word count per citation/reference and word count per ethical principle, with lower values indicating greater density. The word count per citation/reference analysis included 12 matched pairs because pairs with zero or undefined citation/reference denominators were excluded.
*p < 0.05, **p < 0.01, ***p < 0.001.
Relative performance transcript analysis shows a low level of congruence between the structural and stylistic aspects of CALEB and GPT-4o responses (Fig. 1C). CALEB outputs focused on fewer core ethical principles and emphasized the emotional resonance of dilemmas. CALEB transcripts incorporated non-verbal stage directions (e.g., “sighs deeply”) and first-person phrasing. In contrast, GPT-4o generated longer responses with multiple hierarchical categories, producing structured outlines.
Computational empathy classification
Transcript outputs were assessed using a computational empathy classifier across three empathy-related domains: ER, IP, and EX. Figure 2 presents the distribution of classifier labels across the 24 transcripts generated by each model for each empathy domain. The largest contrast was observed in the IP domain, where CALEB received “Strong Communication” labels in 18 of 24 transcripts compared with 1 of 24 GPT-4o transcripts; this difference was significant in paired ordinal analysis (p < 0.001). By contrast, ER scores were low across both models and did not differ significantly (p = 1.000), while EX scores showed no variation because all transcripts from both models were classified as “no communication”. These findings suggest that CALEB more frequently used language classified as interpretive understanding. Since the classifier was developed for text-based peer-support communication and applied here to AI-generated medical ethics dialogue, these findings were interpreted as exploratory linguistic signals rather than direct evidence of empathetic capacity.

The bar chart shows the distribution of classifier labels across 24 transcripts per model for each empathy related domain: emotional reactions (ER), explorations (EX), and interpretations (IP). Each transcript received a consensus label for each domain using a three-point ordinal scale: 0 = no communication, 1 = weak communication, and 2 = strong communication. Consensus labels were derived through a majority voting across three independent model runs.
Blinded ethicist transcript ratings
Three medical ethicists completed blinded ratings of the 48 randomized transcripts. These ratings reflected transcript-only evaluation and did not include the later live XR interaction. Two of three ethicists generally rated GPT-4o transcripts more favorably than CALEB transcripts for comprehensiveness, nuance, clarity, appropriate length, usefulness for ethics tutoring and overall ethical response quality (Table 2, Supplementary Appendix SA5 for full analysis). Ethicists’ comments described GPT-4o-generated responses as “a helpful outline of the issues in a structure that one can then build on to create a talk or a paper,” whereas CALEB’s responses were “way too short.” Regarding CALEB emotionally framing its responses, the ethicists commented that cues could be seen as awkward or repetitive: “Can we drop the stage-directions? They are a bit cringy.” Given the small number of expert reviewers, these ratings are best interpreted as descriptive formative feedback rather than confirmatory evidence of model superiority.
Phase 1 Blinded Transcript Ratings by Three Medical Ethicists
Values are reported as median (IQR) across 24 matched prompt pairs per system. Ratings were completed on a six-point Likert scale, with 0 indicating strongly disagree and 5 indicating strongly agree. Each ethicist independently rated all 48 randomized transcripts while blinded to model identity. P-values were calculated using paired Wilcoxon signed-rank tests within each ethicist and rating item. These ratings reflect transcript-only evaluation and do not include the later live XR interaction with CALEB. Given the small single-center expert panel, inferential comparisons should be interpreted as exploratory and formative rather than confirmatory validation evidence.
*p < 0.05, **p < 0.01, ***p < 0.001.
Phase 2: Formative Post-Session feedback after live XR interaction with CALEB
In phase 2, post-session feedback after live XR interaction with CALEB suggested that the agent was perceived more favorably in its intended interactive format than in static transcript review. This phase did not include matched live interaction with GPT-40 and was therefore not interpreted as a direct comparative assessment between systems. Thematic coding of ethicist feedback identified five recurrent domains across the 48 total coded excerpts: humanity (29%) and usefulness (25%) were most frequently endorsed, followed by accuracy (21%), flexibility (15%), and comprehensiveness (10%) (Supplementary Appendix SA5). Ethicists described CALEB as providing ethically supportable and nuanced responses, maintaining conversational flow and handling iterative follow-up questions appropriately. They also emphasized its human-like tone, empathy and value for practicing ethical dialogue. One ethicist remarked that interacting with CALEB “very much resembled a discussion we might have at a medical ethics committee meeting.”
Live interaction also revealed CALEB’s shortcomings. One ethicist noted that in discussing the allocation of deceased donor kidneys, CALEB proposed potential policy changes already implemented in the United States system. Another gap was observed in research ethics, where CALEB acknowledged the risks of exploitation but not the equally important risk of excluding vulnerable populations from research. While no clear “hallucinations” were observed, these lapses suggested incomplete domain coverage in certain scenarios.
Discussion
This study presents a mixed-methods comparison of a domain-specific conversational agent (CALEB) to a foundational LLM (GPT-4o). Our findings demonstrate that both systems can accurately convey ethical principles, yet they differ fundamentally in style, emphasis, and pedagogical value. Consistent with prior observations, 18 GPT-4o provides exhaustive coverage and structured frameworks; in contrast, CALEB delivers emotionally resonant, conversational dialogue that mirrors the lived experience of ethical deliberation. Additionally, CALEB’s responses were 95% shorter but contained ethical principles at 7.7x higher density, suggesting that brevity need not sacrifice principle coverage when systems are designed with domain expertise. Overall, CALEB’s outputs reflect its custom knowledge embedding, narrative scripting, and persona design.
These fundamental differences in style and substance suggest that both modalities may be useful for medical ethics teaching, though perhaps in different settings. CALEB’s live format may be better suited for problem-based learning formats that encourage critical thinking and Socratic-style learning. In contrast, GPT-4o’s strengths may be in providing comprehensive education of ethical frameworks and detailed implementation guidance. From a clinical perspective, CALEB appears well-suited to simulation training and debriefing, fostering more interpretation-oriented empathetic output (i.e., reassurance and acknowledgement).
These different systems may play complementary roles in global medical ethics education given its current limitations. Current gaps include lack of curricular standardization, limited opportunities for deliberate practice in clinical settings, insufficient faculty support, and over-reliance on didactic lectures that fail to cultivate critical thinking or contextual sensitivity.3,5,19 Even when formal ethics didactic instruction exists, many schools provide little funding, rely on non-specialist faculty, or omit key domains such as end-of-life care or emerging technologies.3,20,21 These shortcomings persist into residency education as well22–27 and are seen in curricula worldwide, especially outside North America and Europe.22,27 The need for improved ethical foundations and competency also exists at the faculty level.21,28,29 Conversational agents like CALEB may address a critical need at all stages of training through scalable, interactive, dialectical pedagogy, and case deliberation. 2
Our analysis here also displays a fundamental problem in critically evaluating AICAs. Static transcript review and live XR interaction assessed different conditions and were not directly comparable tests of model superiority. In phase 1, blinded transcript review generally favored GPT-4o because its responses were longer, more comprehensive and more conventionally structured. In Phase 2, post-session feedback suggested that CALEB was perceived more favorably when used in the modality for which it was designed for: real-time, embodied, dialogical interaction. This suggests that some pedagogical value may be interaction-dependent. Live interaction via turn-taking dynamics, embodied presence (avatar gaze and expression), immediate responsiveness to clarifying questions could be factors for future development and evaluation. Extended interaction also allowed CALEB to provide more structured perspectives, respond to follow-up questions and sustain more natural ethics dialogue. These findings, overall, suggest that transcript-only evaluations may not fully capture the perceived values of AICAs whose intended use depends on real-time interaction.
Additionally, our study was limited, as it was conducted at a single institution using a small number of ethicist reviewers. As a result, the expert review component should be interpreted as formative feasibility input rather than independent validation, as findings may be sensitive to individual rater variation and local expectations about ethics teaching. Additionally, although transcript ratings were blinded to model identity, the practical effectiveness of blinding may have been limited by distinctive stylistic features of the two systems, including response length, first-person phrasing, stage directions and outline-like formatting. These features may have allowed reviewers to infer which system generated some transcripts. Finally, the methodology for CALEB highlights the benefits and flaws of deliberately curated domain expertise. Despite extensive training using curriculum-based ethics references, shortcomings were identified by the ethicists, such as incomplete coverage of allocation policies or research ethics. Unlike GPT-4o, which evolves unpredictably through continuous training on a broad, uncontrolled corpus, CALEB can be systematically modified in response to expert feedback to improve future outputs. Thus, broader validation across cultural, legal, and professional settings is essential for future studies, as is analysis of dynamic in-person experiences.
In conclusion, AI systems may be effective partners in medical ethics education but have distinct and complementary strengths. Our analysis demonstrates that domain-specific agents like CALEB are a feasible, accurate source of information and that their unique qualities may be well-suited for ethics engagement in certain circumstances. Future research should explore scalability, cultural adaptation, and integration into existing training pathways, advancing a new paradigm of AI-assisted ethics education.
Supplemental Material
sj-docx-1-mxr-10.1177_29941520261469342 — Supplemental material for A Conversational Ethics Bot (CALEB) Versus GPT-4o for Medical Ethics Education: Transcript Analysis and Extended Reality Feasibility Study
Supplemental material, sj-docx-1-mxr-10.1177_29941520261469342 for A Conversational Ethics Bot (CALEB) Versus GPT-4o for Medical Ethics Education: Transcript Analysis and Extended Reality Feasibility Study by Rohan Jotwani, Vansh D. Barot, Peter A. Goldstein, Alexandros Sigaras, Sandhya Sriram, Robert S. White, Debjani Mukherjee, Ezra Gabbay, June M. Chan, Julia Steigerwald Schnall, and John E. Rubin
Footnotes
Author Disclosure Statement
P.A.G. is a co-inventor on patents related to the development of novel alkylphenols for the treatment of neuropathic pain and serves on the Scientific Advisory Board for Akelos, Inc. (New York, NY), a research-based biotechnology company that has secured a licensing agreement for the use of those patents. R.J. serves as a consultant to Abbott Laboratories and MaryAnn Liebert publishers, serves voluntarily on the scientific advisory board for 3DOrganon, and has received research support from the Nannette Laitman Foundation, Meta/Bodyswaps, and the Accreditation Council for Graduate Medical Education. None of these relationships present any direct conflicts of interest to the content of this article. D.M. is a paid consultant on a grant-funded research project at the Hastings Center for Bioethics, an independent, nonpartisan, nonprofit interdisciplinary research institute. This relationship does not present any direct conflicts of interest to the content of this article.
Funding Information
The Nanette Laitman Education Scholar Award in Entrepreneurship and the Josiah Macy Jr. Foundation through the Macy Faculty Scholars Program.
Abbreviations
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
