Background. Equipping learners with AI-ethical literacy is a priority for higher education, yet item-level self-reports leave two observables unrecorded: the structure of lexical co-use in learners’ written responses and the polarity of their free text. Neither is a measure of learning. A further inferential hazard is that shared-term statistics in participant–morpheme networks are sensitive to response length, so an apparent rise in shared vocabulary can reflect how much students write rather than lexical alignment; this confound is the target of the present study. We present Multimethod Ethics-Learning Evaluation (MELE), a multimethod evaluation workflow integrating paired post-reading/post-discussion statistical analysis, length-controlled inference on bipartite participant–morpheme networks, and BERT-based sentiment classification, demonstrated here as a single-group proof-of-concept on a facilitated case discussion in AI ethics. Methods. Twenty graduate students from diverse disciplines at a Japanese university deliberated an authentic AI job-screening case grounded in the Japanese shลซkatsu (job-hunting) context. Because T1 was administered after individual case reading, the paired contrast reflects change associated with the facilitated discussion rather than with case exposure as such. Three complementary analyses were applied: Holm/Bonferroni-corrected Wilcoxon tests, an exact conditional within-participant token-permutation reference test for discourse convergence with Monte Carlo-approximated -values, and BERT sentiment classification (exploratory). Findings. The shared-term share of the participant–morpheme networks rose descriptively between administrations, while post-discussion responses were substantially shorter. Under the length-controlled within-participant token-permutation test, neither observed increase was statistically significant relative to its length-conditioned null distribution (two-sided Monte Carlo-approximated -values; Q6 ; Q7 ), whereas a naive whole-bag permutation test would have reported both as significant; smaller effects cannot be excluded. Perceived fairness (Q2) rose significantly and survived multiple-comparison correction ( , , Holm/Bonferroni ; , large); this single-item effect is, however, sensitive to floor-responder exclusion: it attenuates to non-significance ( , ) when the five participants who began at the scale floor are removed, so we report it as preliminary rather than established. The other four items showed non-significant directional gains; participant-level model-estimated polarity shifted positively but non-significantly. Implications. Methodologically, MELE contributes a length-controlled discourse-inference procedure for participant–morpheme graphs: we show that the conventional degree-preserving null is degenerate for the star configurations underlying the shared-term share and that naive convergence permutation tests are confounded by response length, and we provide an exact conditional reference test with Monte Carlo-approximated -values and a bounded simulation study. Under within-participant exchangeability the test is finite-sample valid by construction for any fixed configuration of response lengths (length imbalance can reduce power but cannot legitimately inflate the false-positive rate), with the simulation study serving as an implementation check. In one of the ten null cells the estimated false-positive rate ( ) lay above the nominal , so the simulation does not by itself demonstrate calibration at that level, and performance beyond the evaluated cells remains to be characterized. Substantively, the fairness item showed a floor-sensitive, case-bound shift in this brief, single-case discussion. In this proof-of-concept, then, the workflow’s principal yield is the length-controlled discourse-inference procedure; the fairness shift is a single preliminary observation, and the discourse and sentiment components are hypothesis-generating rather than confirmatory, and we make no claim of cross-component corroboration. As a single-group proof-of-concept ( ), all findings require controlled, longitudinal, cross-cultural replication; key limitations are discussed.
- Inicie sesión para enviar comentarios