Accepted Papers
5th International AIxIA Workshop on Artificial Intelligence for Healthcare
Perugia, Italy, 06-09 October 2026
Review phase completed Program TBA. 📆⏰
5th International AIxIA Workshop on Artificial Intelligence for Healthcare
Perugia, Italy, 06-09 October 2026
Please find next the list of accepted papers along with authors, abstracts and some additional pieces of information.
CephAlignNet: A Coarse-to-Fine Deep Learning Framework for Cephalometric Landmark Localization
Authors: Simone Carlassare, Alice Bizzarri, Riccardo Zese
Student paper: YES
Keywords: Cephalometry; Artificial Intelligence; Medical Imaging; Landmark Detection; Deep Learning; Orthodontics; Heatmap Regression
Abstract: Cephalometric analysis is a key diagnostic procedure in orthodontics that relies on the accurate identification of anatomical landmarks on lateral cephalometric radiographs. Although recent AI-based methods have significantly improved automatic landmark localization, many existing solutions are trained on limited datasets, support only a restricted number of landmarks, or show reduced generalization across imaging devices. This paper presents CephAlignNet, a deep learning framework for automatic localization of 26 anatomical landmarks. The model is trained on multiple public benchmark datasets together with a clinically validated private dataset, improving robustness and generalization. The framework follows a two-stage coarse-to-fine strategy, combining a global heatmap-based detector with a local refinement module operating on high-resolution features. Experimental results show that CephAlignNet achieves an MRE of 1.21 mm and an SDR@2mm of 86.1%, demonstrating robust performance across heterogeneous imaging conditions. These results highlight the potential of the proposed approach for reliable AI-assisted cephalometric analysis.
Synthetic Data Generation, Multi-Dimensional Validation, and Clinically Actionable Risk Stratification in a FIL-DLBCL Clinical Cohort
Authors: Aaron Costa, Stefano Nera, Riccardo Francia, Luca Piovesan, Simone Pernice, Francesco Merli, Annalisa Arcari, Emanuele Cencini, Stefania Montani, Simone Ferrero, Francesca Cordero, Marco Beccuti
Student paper: YES
Keywords: Synthetic Data Generation; Flow Matching; AutoML; Risk Stratification; Diffuse Large B-Cell Lymphoma
Abstract: Standardized risk assessment in Diffuse Large B-Cell Lymphoma, such as the Elderly Prognostic Index, relies on fixed Cox-derived coefficients from historical cohorts. This work investigates whether high-fidelity synthetic data can refine such tools within the Fondazione Italiana Linfomi clinical cohort. We introduce a Generative Framework centered on our main contribution, a Transformer-based Flow Matching generator designed to preserve multivariate clinical geometry, and benchmark it against domain-adapted Large Language Model, Variational Autoencoder, and state-of-the-art tabular synthesis methods. To move beyond statistical fidelity, we evaluate the generated cohorts through a Clinical Actionability Framework assessing their utility for downstream prognostic modelling. A fixed downstream predictive pipeline is used to generate mortality probabilities, which are discretized for survival-based risk stratification. Through Train-on-Synthetic-Test-on-Real and augmentation protocols, we show that high-fidelity synthetic data improve risk separation on held-out real patients while preserving interpretability. SHAP analysis suggests that the refined models capture clinically meaningful patient signals not fully represented by the EPI, helping explain where their predictions diverge from conventional risk stratification.
Patient-Centred Adaptive Healthcare Scheduling with Explainable Probabilistic Mental and Physical Actions
Authors: Stefania Costantini, Valentina Pitoni, Lorenzo De Lauretis
Student paper: NO
Keywords: Artificial Intelligence for Healthcare; Cognitive Agents; Patient-centred Healthcare Scheduling; Epistemic Logic; Probabilistic Reasoning; Explainable AI; Multi-Agent Systems
Abstract: Current healthcare scheduling systems primarily optimize organizational objectives, whereas patient-centred scheduling requires reasoning about patient-specific information, organizational constraints, trust relationships, and uncertainty, while providing transparent explanations for every scheduling decision. Building upon our previous evolution from ASP-based scheduling with Blueprint Personas to cognitive agents based on the epistemic logic L-DINF, this paper completes the next stage of this research programme by equipping cognitive agents with probabilistic reasoning capabilities. The resulting framework combines symbolic reasoning, probabilistic assessment, patient preferences, organizational constraints, trust relationships, and action reliability within a unified and explainable cognitive architecture. The framework has been fully operationalized in DALI2. A patient-centred healthcare case study demonstrates dynamic adaptation of a patient's first consultation according to eligibility, consent, modality preferences, tolerated delay, physician availability, and inter-clinic cooperation. A controlled synthetic evaluation illustrates the trade-off between decision reliability and scheduling coverage under different probability thresholds.
Omics-Enhanced Multimodal Representations for Cancer Prognosis and Staging
Authors: Sara De Luca, Federico D’Asaro, Paola Berchialla, Giuseppe Rizzo
Student paper: YES
Keywords: Multimodal Learning; Omics Integration; Large Language Models; Cancer Survival Prediction; Patient Stratification
Abstract: Comprehensive patient characterisation in oncology requires integrating heterogeneous data modalities, such as omics profiles, clinical records and radiology. Exploiting the complementarity of these information is key for accurate survival and staging prediction. We investigate whether exploiting omics-derived embeddings enriches the patient latent representation in a multimodal large language model and improves survival and staging prediction beyond unimodal and post-hoc fusion baselines. We present OmicsLaMed, which extends M3D-LaMed, a LLaMA-2-based model pre-trained on volumetric Computed Tomography scans and clinical text, with an omics connector that projects the embeddings generated with a Cox autoencoder (trained on more than 7000 RNA-seq and DNA methylation profiles) into the model's token embedding space. The model is fine-tuned via low-rank adaptation (LoRA) on three structured prediction tasks: vital status, survival interval classification, and cancer staging. Evaluation covers both a generative prediction paradigm and a representation paradigm (unsupervised clustering on the last hidden state of the model). The omics-enriched model achieves a mean balanced accuracy of 63.6% in the generative paradigm and 61.9% in the embedding clustering paradigm, outperforming baselines such as naive feature concatenation (58.4%), unimodal omics clustering (58.4%), late fusion (55.3%), and the zero-shot M3D-LaMed baseline (7.9% generative, 48.1% clustering). These results support the hypothesis that incorporating omics-derived representations alongside clinical reports and radiology imaging within a shared latent space yields richer, more clinically meaningful patient representations.
Automated Shift Scheduling in Radiology via Answer Set Programming
Authors: Aurora De Pieri, Francesco Filocamo, Marco Maratea, Cinzia Marte, Chiara Spadafora
Student paper: YES
Keywords: Answer Set Programming; Logic Programming; Digital Health
Abstract: Physician shift scheduling in hospital departments is a challenging combinatorial optimization problem with multiple constraints, including clinical skill coverage, on-call duty eligibility, mandatory rest periods, and a fair distribution of workload among staff. The generation of high-quality schedules is essential to maintain operational efficiency and promote physician well-being. This paper presents an Answer Set Programming approach for the automated generation of monthly physician rosters in the Radiology Department of the "Ferrari" Spoke Hospital in Castrovillari (Italy). The proposed model, evaluated on real-world monthly scheduling instances from the hospital department, encodes both national labor regulations and departmental requirements, comprising optimization criteria to improve workload fairness. Results show that the approach is able to generate balanced schedules in a short time, indicating that the ASP-based approach can serve as a decision-support tool for medical rostering in this hospital setting.
DUET: Dual-Objective Knowledge Distillation for Task-Aware ECG Representation Learning
Authors: Francesca Filice, Edoardo De Rose, Simone Bartucci, Francesco Calimeri, Simona Perri
Student paper: YES
Keywords: Foundation Models; Knowledge Distillation; Feature Selection; Electrocardiogram; Fairness; Model Compression
Abstract: Foundation models have proven to be remarkably effective in electrocardiogram analysis, by learning informative and transferable representations from large-scale datasets. Nonetheless, their large number of parameters and computational requirements may hinder deployment on resource-constrained and edge devices. To address this limitation, model compression approaches, and Knowledge Distillation in particular, aim to transfer relevant information from a high-capacity teacher model to a smaller and less computationally demanding student. Traditional Knowledge Distillation methods are often primarily focused on transferring predictive outputs for a specific downstream task, and may therefore fail to fully exploit the rich latent representations learned by foundation models. Feature-based approaches seek to overcome this limitation by aligning the internal representations of teacher and student models, but they still face challenges arising from heterogeneous architectures, dimensional mismatch, and the possible transfer of task-irrelevant information. In this work, we propose DUET, a feature-based Knowledge Distillation framework that combines latent representation alignment with downstream classification guidance. Starting from the embeddings produced by a frozen ECG foundation model, DUET learns a low-dimensional, task-oriented projection and employs an Exponential Moving Average projector to provide stable latent targets to heterogeneous student architectures. We evaluate multiple configurations of the framework, differing in the use of debiasing regularization and direct classification supervision for the student. Experiments on two ECG datasets and three rhythm-classification tasks show that the proposed approach consistently improves the corresponding stand-alone student models across datasets, tasks, and student architectures. These results suggest that task-guided latent distillation is a promising strategy for transferring knowledge from ECG foundation models to more compact models.
Federated Eligibility Criteria Optimization for Clinical Trial Emulation
Authors: Eriberto Andrea Franchi, Aldo Marzullo, Elena De Momi, Alberto Redaelli
Student paper: YES
Keywords: Clinical trial emulation; Eligibility criteria; Federated learning; Real-world data; Survival analysis; MIMIC-IV
Abstract: Clinical trials protect their participants with eligibility criteria, but the same criteria often exclude most of the patients who would receive the therapy in routine care, slowing recruitment and limiting how well trial results generalize. Recent frameworks such as Trial Pathfinder showed that many of these criteria can be safely relaxed: by re-emulating a completed trial on real-world data and measuring how much each criterion changes the estimated treatment effect, it flags the rules that can be dropped without worsening the outcome. This analysis, however, assumes that all the data sit in one database. Run at a single hospital it sees only that hospital's patients, and hospitals usually cannot pool their records because of privacy regulation. We show that Trial Pathfinder’s criterion analysis can be conducted across multiple institutions without transferring patient-level records and without compromising performance. Each site shares only aggregate, non-identifying summaries of its patients at risk; from these the coordinator rebuilds exactly the same survival model that centralized pooling would have produced. The federated result is therefore equal to pooling by construction. On synthetic data with a known correct answer, the federated analysis makes the right relaxation decisions where a single, biased hospital is fooled into dropping a criterion it should keep. On a sepsis study built from the MIMIC-IV database and split into two hospital services, it reproduces the pooled hazard ratio and every criterion decision to within 10^(-15), while the smaller service, analyzed alone, is about twice as uncertain and disagrees on one criterion. Exact federated criterion optimization thus lets institutions broaden trial eligibility together, privately, with no penalty relative to a centralized analysis.
Evaluating Automated Subcutaneous Adipose Tissue Segmentation on Longitudinal MRI Through Prostate Cancer Survival Prediction
Authors: Fabio Gatelli, Mattia Savardi, Matteo Riva, Alessandro Monti, Marco Ravanelli, Alfredo Berruti, Salvatore Grisanti, Alberto Signoroni
Student paper: YES
Keywords: Subcutaneous Adipose Tissue; Fat Fraction MRI; Volumetric Segmentation; Survival Prediction
Abstract: The analysis of adipose-tissue composition on longitudinal fat-fraction (FF) MRI can provide informative imaging biomarkers in prostate cancer. Although 2D procedures based on a single slice are widely adopted and have been used in statistical studies, they cannot capture the complete volumetric distribution of adipose tissue or its longitudinal heterogeneity. A 3D statistical analysis therefore provides a more stable and faithful representation of body-composition changes over time. Within this framework, subcutaneous adipose tissue (SAT) provides a suitable starting point for volumetric methodological validation, given its widespread use in body-composition research and its relatively clear anatomical boundaries. Our objective is to assess whether an automated 3D SAT segmentation procedure is sufficiently consistent with the clinical 2D L3 vertebra reference and whether the resulting volumetric descriptors provide a more robust basis for downstream survival modeling than conventional 2D measurements. To this end, we evaluate automated segmentation strategies against the available ground truth on the annotated L3 slice using overlap and surface metrics, select the most suitable strategy, and use its volumetric masks for FF-based feature extraction. Segmentation quality was assessed through overlap and surface-distance metrics, and the resulting 3D SAT masks were integrated into a longitudinal survival-analysis framework for progression-free survival (PFS) and overall survival (OS). The results indicate that automated 3D SAT analysis is a feasible extension of conventional 2D L3 assessment and can provide a more robust basis for prognostic modeling on longitudinal FF MRI. Overall, the proposed workflow demonstrates the feasibility of integrating automated volumetric SAT analysis into longitudinal MRI-based prognostic studies.
Leveraging Large Language Models for Electronic Health Record Integration from Semi-Structured Clinical Reports
Authors: Carmen Guzmán-García, Maria Chiara Colombi, Gianfranco Lombardo, Carlotta Mutti, Monica Mordonini, Stefano Cagnoni
Student paper: YES
Keywords: Multimodal AI; Clinical NLP; Large Language Models; Semi-Structured Medical Reports; Data Integration
Abstract: A large portion of clinically relevant information is embedded in semi-structured reports combining tabular measurements and free-text notes. Converting these heterogeneous documents into structured data suitable for analysis and decision support remains a challenging task. Traditional rule-based and Named Entity Recognition (NER) approaches often lack robustness to variability in language and formatting. This work presents an end-to-end framework that leverages lightweight Large Language Models (LLMs) to extract structured information from semi-structured cardio-respiratory monitoring (CRM) reports and integrates the extracted data into a structured research database (REDCap). The proposed framework combines document parsing, prompt-based information extraction, JSON schema enforcement, and automated upload into a REDCap-based data management environment. We experimentally compare two prompting strategies (field-wise Question Answering and Single-Prompt JSON generation) across two model sizes in the Qwen LLM family (0.6B and 1.7B) and evaluate the impact of explicit reasoning (“thinking”) modes. Results show that field-wise extraction provides superior structural robustness, while larger models and reasoning-enabled configurations improve semantic correctness at the cost of increased computation time. The proposed framework implements a scalable strategy for transforming semi-structured clinical documentation into structured datasets suitable for research analytics and clinical decision support.
From Clinical Text to Causal Graphs: How Much Supervision Is Needed to Induce Graph-Usable Symptom Variables?
Authors: Yaqiao Huang, Paloma Rabaey, Henri Arno, Thomas Demeester
Student paper: YES
Keywords: Clinical natural language processing; Weak supervision; Causal structure learning; Clinical variable induction; Electronic health records
Abstract: Constructing symptom variables from clinical text for causal graph learning is challenging when patient-level annotations are limited. We present a two-stage framework that compares unsupervised, weakly supervised, and fully supervised variable induction before downstream causal structure learning. Using the SimSUM benchmark, we evaluate how each supervision regime affects both symptom recovery and structural fidelity. Full supervision achieves the strongest overall performance, but requires substantial record-level annotation and shows diminishing gains as more labeled data are added. Weak supervision, in contrast, uses only a small set of symptom descriptions to align text-derived clusters with the target concepts, notably improving graph recovery over unsupervised induction without patient-level labels. These results show that semantic coherence alone is insufficient for constructing graph-usable variables, and position anchor-guided weak supervision as a practical, annotation-efficient alternative for clinical causal modeling.
What Counts? A Neuro-Symbolic Framework for Non-Anthropocentric Ethical Reasoning
Authors: Bianca Lerma, Gianluca Apriceno, Mauro Dragoni
Student paper: NO
Keywords: Healthcare Decision Support Systems; Medical AI; Non-Anthropocentric Ethics; Active Inference; Ethical Logic; Neuro-Symbolic AI; Global Expected Free Energy
Abstract: Artificial Intelligence is becoming an integral part of healthcare, assisting in clinical decisions and personalized treatment planning. Yet, many current medical AI systems implicitly rely on anthropocentric ethical assumptions, focusing primarily on immediate, patient-level outcomes while failing to adequately represent broader temporal, systemic, and cross-agent impacts of clinical interventions. This gap is particularly significant in healthcare environments where decisions can cascade across multiple clinical actors, time horizons, and even ecological contexts. In this work, we present a healthcare-focused instantiation of NAEL (Non-Anthropocentric Ethical Logic), a computational framework that treats ethics not as fixed rules or external constraints, but as an evolving inferential mechanism. Built on principles of active inference, NAEL frames ethical reasoning as minimizing global expected free energy, extending traditional agent-centered models to incorporate uncertainty arising from an agent’s actions on other agents and the surrounding environment. Within this formulation, ethical considerations are embedded in the agent’s generative model and directly influence policy selection. Through a conceptual case study in intensive care unit resource allocation, we demonstrate how NAEL-based decision policies diverge from standard short-term optimization approaches, underscoring their relevance to more system-aware and ethically grounded clinical decision-making.
An MLOps Framework for Reproducible Radiomics Experiments
Authors: Giulio Mallardi, Luigi Quaranta, Donato Boccuzzi, Filippo Lanubile, Giancarlo Logroscino, Benedetta Tafuri
Student paper: YES
Keywords: Reproducibility; MLOps; ML Engineering; Radiomics; Neuroimaging; Frontotemporal Dementia
Abstract: Translating machine learning (ML) models from research prototypes into clinical use requires development processes that are fully traceable and reproducible. Meeting these requirements is challenging for research laboratories, where radiomics experiments often rely on ad-hoc scripts and manual steps. Machine Learning Operations (MLOps) offers a set of practices and tools widely adopted in industry to scale ML workflows while ensuring end-to-end traceability and reproducibility. This study presents an MLOps-enabled experimentation framework assembled from open-source tools and specifically designed to support radiomics experiments. The framework integrates key MLOps practices, including data versioning, experiment tracking, workflow automation, and environment reproducibility. We evaluate it through a case study focused on automated frontotemporal dementia (FTD) detection from MRI brain scans. Across the diagnostic tasks considered, we achieved strong classification performance. The framework removed the manual effort of orchestrating, tracking, and re-running experiments across classification tasks and feature-selection strategies, while supporting deterministic re-execution under fixed configurations and containerized environments. Although demonstrated on FTD detection, the framework is dataset- and task-agnostic, and can be readily adapted to other radiomics studies seeking reproducible and automated experimentation.
Towards Optimizing Laser Acceleration of Protons for Medical Applications: Benchmarking Generative AI Against Machine Learning Models
Authors: Arantxa Ortega-Leon, Alessandro Zani, Indaco Biazzo, Lea Schuh, Sergio Consoli
Student paper: NO
Keywords: Laser-driven ion acceleration; Text-to-text regression; Machine learning; Proton therapy; Synthetic data; RegressLM
Abstract: Laser-driven ion acceleration in the Target Normal Sheath Acceleration (TNSA) regime is a promising route to compact proton sources for medical applications, but experimental exploration of its parameter space is slow and costly. Machine learning (ML) surrogates trained on synthetic data have been proposed to guide this search, and generative language models have recently been reported to match state-of-the-art regressors by reframing prediction as text-to-text regression. We provide the first assessment of this paradigm in laser-driven ion acceleration. RegressLM, trained either by direct fine-tuning or by pre-training followed by fine-tuning, is benchmarked against three ML baselines, Gaussian Process Regression, Support Vector Regression and a Neural Network, for the maximum, average and total proton energy under matched training-set sizes and metrics at seven noise levels up to 30%. The kernel methods are more accurate and considerably cheaper, most clearly on noise-free data, but the margin is narrow where it matters: from moderate noise onwards all methods operate at the accuracy floor imposed by the injected noise, and neither generative method is statistically distinguishable from the tuned neural baseline. The margin is narrowest on the maximum-energy cutoff, the quantity that sets the depth at which a therapeutic proton beam deposits its dose. Reaching this accuracy with neither architectural modification nor hyperparameter search indicates real potential, whose realization we argue requires finer numeric tokenization, richer serialized inputs, transfer across genuinely distinct tasks, and the inverse problem of recovering laser and target parameters from a prescribed proton spectrum.
LLM-as-a-judge for Evaluating the Quality of Retrieval-Augmented Generation Systems
Authors: Leonardo Sanna, Erica Solinas, Mauro Dragoni
Student paper: NO
Keywords: Retrieval-Augmented Generation; Large Language Models; LLM-as-a-judge; Pregnancy; Evaluation
Abstract: Retrieval-Augmented Generation (RAG) systems are increasingly used in high-stakes domains like healthcare, yet evaluating retrieval quality remains challenging. This study explores the use of LLMs as automated judges to assess document relevance in a medical RAG pipeline focused on maternity and pregnancy in Italian. Comparing llama-3.3-70B-instruct, mistral-small-24B-instruct-2501 and deepseek-r1-distill-llama-70 against a human-annotated gold standard and cross-encoder scores, we find that LLMs systematically reflect retrieval relevance, with Mistral and DeepSeek showing the closest alignment to human judgments. However, LLMs exhibit a bias toward labeling documents as partially relevant, underscoring the importance of model selection and human supervision to refine RAG pipelines. These results highlight the potential of LLMs as practical, scalable evaluators for retrieval quality in medical applications.
MAESTRO: a framework for trustworthy integration of LLMs in psychological digital interventions
Authors: Leonardo Sanna, Mattia Franzin, Simone De Carli, Marco Bolpagni, Simone Casazza, Silvia Rizzi, Claudio Eccher, Mauro Dragoni
Student paper: NO
Keywords: Multi-agent systems; Trustworthy AI; Large Language Models; Digital Therapeutics; Psychology
Abstract: The integration of Large Language Models (LLMs) into conversational systems has greatly enhanced their expressive capabilities but also introduced challenges in structure, control, and reproducibility. Current design approaches often fall between rigid, rule-based systems that are predictable but limited in scalability, and fully LLM-driven models that enable natural interaction but lack transparent logic. We present MAESTRO (Multi-Agent Enhanced System for Therapy Resources Orchestration), a framework for the trustworthy integration of LLMs in complex settings such as psychological digital interventions. Central to MAESTRO is a flow-based abstraction model that represents conversational logic as assemblies of modular, reusable blocks. This approach allows designers to construct explicit interaction flows in which LLMs can serve as bounded-reasoning components, preserving predictability while enabling controlled flexibility.
Structure-Aware Representation Learning of ATC Codes
Authors: Christel Sirocchi, Giorgia Roselli, Sara Montagna
Student paper: NO
Keywords: ATC Codes; Embedding; Electronic Health Records
Abstract: Drug therapies recorded in electronic health records are commonly represented through the Anatomical Therapeutic Chemical (ATC) classification, a hierarchical taxonomy that organizes medications by increasing levels of pharmacological specificity. Conventional approaches, such as one-hot and multi-hot encodings, treat pharmacologically related drugs as independent entities, failing to capture this structure. This study investigates structure-aware numerical representations of ATC codes that preserve their hierarchical semantics. We design and compare four embedding approaches: (a) a hierarchical embedding optimized via multi-level classification and batch-hard triplet loss, (b) an autoencoder-based compressed representation, (c) a discrete autoencoder incorporating hierarchical constraints, and (d) a Graph Attention Network (GAT) trained directly on the ATC taxonomy graph. The proposed approaches are evaluated against random embeddings and Node2Vec using a comprehensive framework that measures hierarchy-level reconstruction accuracy, global clustering quality, local neighborhood consistency, and intra- versus inter-class distances. All proposed methods capture meaningful ATC structure, substantially outperforming random representations and achieving competitive performance with Node2Vec. Specifically, the graph-based GAT embedding provides the strongest global organization of the ATC space, achieving superior clustering separation and compactness, whereas Node2Vec preserves local neighborhood similarity most effectively. These results show that incorporating taxonomy-aware learning objectives produces clinically meaningful drug representations that better reflect the organization of pharmacological knowledge than conventional, unstructured embeddings.
A Multi-Metric Validation Framework for GNN Explanations in Brain Tumor Segmentation
Authors: Alberto G. Valerio, Francesco Scialpi, Gennaro Vessio, Giovanna Castellano, Gabriella Casalino
Student paper: YES
Keywords: Graph Neural Networks; Explainable Artificial Intelligence; Explanation Validation; Brain Tumor Segmentation; Biomedical Imaging; BraTS
Abstract: Graph Neural Networks (GNNs) provide a flexible representation of volumetric biomedical images, but their adoption in high-stakes clinical applications requires reliable explanation validation. Existing evaluation approaches often rely on visual inspection or individual quantitative metrics, although explanation quality encompasses multiple complementary properties. We present a multi-metric framework for validating local post-hoc explanations of GNN-based brain tumor segmentation by jointly assessing causal faithfulness, distributional impact, spatial alignment, and robustness within a reproducible evaluation protocol. The framework is evaluated on a benchmark brain tumor segmentation task using multiple GNN architectures and state-of-the-art explanation methods. Results show that different evaluation criteria emphasize different explanation properties, leading to substantially different rankings of backbone-explainer configurations. Moreover, explanation behavior varies across lesion characteristics, highlighting the importance of evaluating explanations from multiple perspectives. The proposed framework provides a reproducible methodology for identifying trade-offs among complementary validation dimensions and supporting the evaluation and comparison of GNN explanations for graph-based biomedical image analysis.
Language Models for Labelling and Clinical Knowledge Extraction in Hypertension Management
Authors: Gianluca Aguzzi, Martino Pengo, Grzegorz Bilo, Sara Montagna
Student paper: NO
Keywords: Clinical Decision Support Systems; Hypertension; Expert Knowledge Extraction; Language Models
Abstract: Hypertension is a chronic disease that requires continuous monitoring and long-term management since blood pressure control remains suboptimal in many patients. Clinical Decision Support Systems could improve hypertension care. However, to instruct and support the development of such systems, relevant data contained in unstructured clinical documentation must be properly extracted, as well as complex physician reasoning must be inferred and formalised. In this context, recent advances in large language models offer a promising opportunity to bridge the gap between unstructured clinical narratives and computable decision-support knowledge. In this work, we focus on the automatic analysis of outpatient letters in the context of hypertension management. In particular, we investigate whether language models can automatically label patients as having controlled or uncontrolled blood pressure based solely on outpatient letters and a structured rubric that codifies expert clinical knowledge extracted via language models-assisted prompt engineering. Focusing on locally deployable models, we assess their ability to generate clinically relevant labels under privacy and deployment constraints, and find that they approach, without matching, the performance of much larger models. On a verification set of 100 letters, disjoint from the set used to induce the clinical rubric, the best models agree with the expert annotation on 85% of the cases and a locally deployable open-weight model follows six points behind.
Benchmarking Temporal Reasoning and Audience-Aware Explanation Generation of Modern Large Language Models on Clinical Data
Authors: Gianluca Apriceno, Tania Bailoni, Mauro Dragoni
Student paper: NO
Keywords: Large Language Models; Time-Aware Reasoning; Clinical Guideline Compliance; Audience-Tailored Explanations
Abstract: Large Language Models are increasingly applied to clinical tasks, yet their ability to reason over temporally dependent rules and generate explanations for diverse audiences remains underexplored. In this study, we evaluate seven recent LLMs on five real-world patient scenarios, comparing structured (Turtle) and unstructured (natural language) representations of clinical knowledge. We assess both the accuracy of guideline violation detection and the quality of audience-tailored explanations, based on expert qualitative evaluation of correctness, clarity, temporal traceability, and reasoning alignment. Our results show that LLMs handle unstructured prompts more reliably, while structured formats often lead to misinterpretations, particularly in scenarios involving longer temporal sequences. Expert evaluation indicates that while LLMs can generate coherent explanations, their consistency and alignment with intended audiences can vary, highlighting the importance of expert validation. These findings highlight both the potential and limitations of LLMs for temporally aware clinical decision support, pointing to areas for further development in reliable, patient-centered AI systems.
Towards LLM-Assisted Process Modelling from Healthcare Documents: A Framework and a BPMN Case Study
Authors: Sofia Benotti, Roberto Nai, Emilio Sulis, Vittoria M.S. Trifiletti, Guido Boella, Paolo Ferrero, Claudio Plazzotta, Elisa Sandri, Patricia Scioli
Student paper: YES
Keywords: Process modelling; LLM; Diagnostic Therapeutic Care Pathways
Abstract: Diagnostic Therapeutic Care Pathways are essential for standardising healthcare processes, yet they are typically documented in unstructured natural language, making their transformation into formal process models a time-consuming and expertise-intensive task. Recent advances in Large Language Models (LLMs) offer new opportunities to support knowledge extraction from textual documents and facilitate conceptual modelling activities. This paper investigates the adoption of LLMs to support the construction of BPMN models from clinical documentation. We focus on a subset of key process modelling tasks, e.g. activity identification, organisational resource extraction, and the identification of candidate BPMN elements while maintaining human oversight throughout the modelling process. The approach is evaluated through a case study in the healthcare sector, focusing on the complex workflow of a Breast Unit at a hospital in Turin, which is modelled using the standard BPMN notation. The results contribute to research on the potential and limitations of LLMs in supporting healthcare process modelling by addressing issues such as ambiguity, reliability, and traceability.
Machine Learning for Daily Rehabilitation Scheduling Capacity Prediction
Authors: Pierangela Bruno, Giuseppe Galatà, Cristian Loria, Matteo Mammoliti, Marco Maratea
Student paper: YES
Keywords: Machine Learning; Rehabilitation Scheduling; Healthcare Analysis; Regression models
Abstract: Estimating the daily number of rehabilitation assignments before schedule generation can provide valuable support for operational planning in rehabilitation centers. However, rehabilitation planning data are primarily collected to support the scheduling process and require extensive preprocessing before they can be exploited for predictive analytics. In this paper, we investigate whether historical rehabilitation records can be effectively exploited to predict the daily rehabilitation scheduling capacity, i.e., the total number of assignments of patients in a rehabilitation plan. To this end, we propose a Machine Learning (ML) pipeline that analyzes nested planning records, converting them into a structured tabular representation through feature extraction, data aggregation, and data cleaning. The resulting dataset is used to train and evaluate several regression models, including Random Forest, XGBoost, and TabPFN. Experimental results obtained from real-world data collected at two rehabilitation centers show that the approach can accurately estimate the daily number of assignments and, consequently, support rehabilitation planning generation.
Comparing Recurrent Neural Networks against Transformers for One-Step Molecular Dynamics Prediction
Authors: Alessia Dodaro, Rama Mhalla, Ibrahim Shkhis, Carlo Adornetto, Vincenzo Damico, Francesca Filice, Gianluigi Greco, Tiziana Marino, Mario Prejanò, Simone Ventrici
Student paper: YES
Keywords: Molecular Dynamics; Deep Learning; Transformer; Recurrent Neural Networks; Molecular Trajectory Prediction
Abstract: Accurate prediction of molecular dynamics (MD) trajectories is critical for accelerating molecular simulations, yet conventional MD remains computationally expensive. This work compares recurrent neural networks (RNNs) and Transformer-based architectures for one-step molecular dynamics prediction. The models predict atomic coordinates, velocities, and forces from sequential simulation data, and are evaluated using standard regression metrics. To investigate the role of molecular context, we perform an ablation study comparing Transformers that use only local atomic neighborhoods with models that additionally incorporate global molecular information. Experimental results show that Transformer-based models consistently outperform the recurrent architectures while remaining computationally efficient. The ablation study further suggests that one-step MD prediction is driven primarily by local spatiotemporal information, while the contribution of global molecular context appears limited in the considered setting. These findings highlight the effectiveness of Transformer architectures for molecular dynamics prediction and provide insights into the relative importance of local and global molecular information.
A Two-Stage Radiomics Pipeline for MRI-Based Multi-Class Subtyping of Primary CNS Lymphoma: Dataset and Proposed Method
Authors: Paolo Fiore, Simone Bartucci, Edoardo De Rose, Francesca Filice, Nicoletta Anzalone, Raffaello Bonacchi, Francesco Calimeri, Teresa Calimeri, Andrés J.M. Ferreri, Simona Perri
Student paper: YES
Keywords: Primary CNS lymphoma; Radiomics; MRI; Patient-level subtyping; Multi-center generalization; Machine learning; Human-centered AI; Interpretability
Abstract: Primary central nervous system lymphoma (PCNSL) can be radiologically subdivided into four distinct morphological subtypes that could potentially inform diagnosis and management; however, this subtyping currently relies on expert neuroradiological assessment, which is not readily available outside specialized centers. Building on our prior work defining a data-driven radiological taxonomy in a large PCNSL cohort, we present work in progress toward a two-stage pipeline for automated, patient-level subtyping of PCNSL based on lesion-level features. The pipeline first extracts a deterministic set of lesion-level MRI features consistent with expert assessment, then predicts the patient’s subtype using a supervised model trained on these features. We describe a two-center dataset of 142 patients from IRCCS San Raffaele Hospital, Milan, and 150 from the public UCSF PCNSL MRI dataset, with a protocol that validates the model internally before testing its generalization externally. At this stage, our contribution is the pipeline, the dataset, and the validation protocol; results are not yet available. The goal is a system that supports, rather than replaces, the neuroradiologist, extending expert-level subtyping to centers that lack it.
MIND-WOZ: A Multimodal Sentiment Analysis System for Detecting Mental Health Disorders in Conversations
Authors: Alessandra Grossi, Giulia Rizzi, Francesca Gasparini
Student paper: YES
Keywords: Speech Depression Recognition; Sentiment analysis; Multimodal Machine Learning; Mental Health
Abstract: Depression is one of the most common mental disorders in the world, affecting about 4% of the population. Artificial Intelligence systems capable of detecting early stages of depression can provide clinicians with valuable decision-support tools, enabling more timely and accurate diagnoses. In this context, Speech Depression Recognition (SDR) has emerged as a promising research area. However, automatic depression detection from speech remains challenging due to limited data availability and high subject heterogeneity. To address these challenges, this paper introduces MIND-WOZ - Multimodal Identification of Neuropsychiatric Disorders on DAIC-WOZ dataset, a multimodal approach based on pre-trained models for the identification of mental health disorders in conversations. Using the DAIC-WOZ dataset, two different configurations are investigated: Base MIND-WOZ, using pre-trained general networks, and Sentiment MIND-WOZ, which incorporates explicit sentiment information. Experimental results show that integrating explicit sentiment information improves classification performance, achieving F1-scores of more than 70%. Furthermore, an in-depth analysis of the dataset reveals critical benchmark limitations, including subject heterogeneity and target leakage from health-related interview questions. These findings highlight the need for data-centric improvements to develop more robust and generalizable speech-based depression recognition systems.
The Prosody of Clinical Speech: Annotation and Computational Analysis of Psychiatric Interviews
Authors: Claire Lawand, Haeeul Hwang, Jalal Al-Tamimi, Valeria Lucarini, Mahshid Hozhabr, Deok-Hee Kim-Dufor, Christophe Lemey, Julien Desclés, Motasem Alrahabi
Student paper: YES
Keywords: psychosis; speech analysis; silence; filled pauses; clinical NLP
Abstract: We present an exploratory analysis of silence and filled pause annotation in 39 semi-structured clinical interviews between psychiatrists and patients assessed for psychosis risk, drawn from the APAISE+ corpus. Participants were classified into three groups: at clinical high risk (AR), not at risk (NAR), and first-episode psychosis (FEP). We developed a manual annotation protocol using Praat TextGrid files, parsed 58,243 annotated intervals, and applied statistical methods to study silence distributions, stationarity, within-interview dynamics, and clinical group differences. Silence accounts for 15.5% of patient speaking time (median 1.01 s). Silence duration differs significantly across three clinical groups (AR, NAR, FEP; Kruskal-Wallis H = 289.9, p < 0.001), while filled pause rate does not (H = 0.59, p = 0.75). First-episode psychosis patients produce proportionally longer silences (>2 s) than at-risk patients, consistent with disrupted speech planning. These findings are invisible to automatic speech recognition, which suppresses over 90% of filled pauses and cannot represent silence, underscoring the irreplaceable role of manual annotation in clinical NLP.
Handling Data Heterogeneity for Clinical Outcome Forecasting in Rehabilitation Environments
Authors: Corrado Loglisci, Stefano Mazzoleni
Student paper: NO
Keywords: Rehabilitation; Clinical Outcome Forecasting; Longitudinal Data; Data Heterogeneity; Robotic Rehabilitation; Wearable Sensors
Abstract: Rehabilitation environments produce heterogeneous longitudinal data whose different acquisition protocols and temporal granularities hinder their direct integration into forecasting models. This work-in-progress paper proposes a predictive framework centered on unified interval representations delimited by consecutive clinical assessments. Timestamped robotic-rehabilitation and wearable sequences are summarized through explicit descriptors of distribution, temporal variation, and sampling, while assessment-aligned motion-capture measurements are incorporated directly. The resulting homogeneous longitudinal feature space is independent of the downstream predictor and supports both next-assessment and longer-horizon forecasting. LSTM and GRU networks, together with a Linear Mixed-Effects Model baseline, are considered to evaluate the information retained by the representation. The framework is being validated within the project PRIN PREDICTOR.
When Metrics Disagree: Format Sensitivity in ECG-Image Adaptation of Small Vision-Language Models
Authors: Alessandro Mancuso, Alessandro Quarta, Pierangelo Veltri, Francesco Calimeri
Student paper: YES
Keywords: Vision-language models; Electrocardiogram; ECG; Evaluation methodology; LLM-as-a-judge; Parameter-efficient fine-tuning
Abstract: Deploying small in-house Vision-Language Models (VLMs) in resource-constrained clinics requires reliable evaluation. However, evaluating free-text outputs against closed-form benchmarks can introduce metric-specific biases, especially when relying on LLM-as-a-judge frameworks or deterministic parsers. We analyze outputs from MedGemma-1.5-4B and DeepSeek-VL2-Tiny adapted on 10,000 ECG image-report pairs and evaluated on a 1,000-item ECG-QA cohort. Using a human semantic-consensus reference for parser-judge disagreements, we show that conclusions about fine-tuning effectiveness depend strongly on the evaluation method. Disagreements between evaluation modes reach up to 22.0 percentage points. Uncalibrated LLM judges exhibit verbosity and semantic-relaxation biases, accepting some fluent reports inconsistent with the reference, whereas parsers penalize legitimate reformulations.
Human Digital Twin and Human-Robot Interaction for Narrative-Enhanced Prediction of Quality of Life of Frail Individuals
Authors: Maria Grazia Miccoli, Berardina De Carolis, Corrado Loglisci, Giuseppe Palestra
Student paper: YES
Keywords: Human Digital Twin; Human-Robot Interaction; Narrative Medicine; Quality of Life; Frailty; Clinical NLP
Abstract: Frailty is a multidimensional condition characterized by reduced physiological reserve and increased vulnerability to adverse health outcomes. Standard frailty assessment tools rely predominantly on structured clinical and functional indicators, often overlooking the subjective, lived dimension of Quality of Life (QoL). In this context, we propose a conceptual framework that integrates three complementary components: (i) a human digital twin that continuously models the physiological, functional and behavioural state of a frail individual; (ii) a human-robot interaction layer, envisioned as a natural communication channel for longitudinal narrative elicitation in domestic or care settings; and (iii) a narrative medicine module that applies natural language processing to elicited personal narratives in order to extract qualitative indicators of perceived well-being. The central hypothesis underlying this framework is that narrative data captures dimensions of well-being that are complementary to standard clinical and functional measures, and that, once integrated with structured data within a unified digital twin, their inclusion can improve the accuracy and clinical relevance of QoL assessment for frail individuals. As a first step toward this goal, this paper presents and evaluates a preliminary, text-based prototype of the narrative medicine module, which estimates WHOQOL-BREF-aligned QoL scores directly from spontaneous conversational narratives, and discusses how this component is intended to be integrated with robot-mediated elicitation and a fuller digital-twin representation in future work.
Structural vs Psycholinguistic Evaluation of Foundation Model-Generated Aphasia Therapy Material
Authors: Mihir Mulye, Stefan Conrad, Stefan Knecht
Student paper: YES
Keywords: Generative models; Aphasia Therapy; Brysbaert Concreteness
Abstract: Automated evaluation metrics for generated aphasia therapy material assess structural qualities but do not explicitly target the psycholinguistic properties that forms the basis of therapy effectiveness. In aphasia rehabilitation, vocabulary concreteness is one of the strongest predictors of patient lexical retrieval performance, yet existing evaluation metrics are not entirely designed to capture it. In this work, we ask whether automated structural metrics for generated aphasia therapy material implicitly align with psycholinguistic norms. To this end, we generate 5,134 multiple-choice therapy items using six open-source foundation models and evaluate each using structural quality metrics, and independently look up the concreteness of its vocabulary using the Brysbaert psycholinguistic norms. We compute nine theoretically motivated Spearman correlations that reveal that the two frameworks are compatible but non-redundant with majority of the values ranging from (ρ = 0.032-0.440) indicating weak to moderate correlation. Among the structural metrics computed, choice concreteness and distractor validity emerge as the ones with most close alignment with psycholinguistic concreteness, whereas question clarity and answer uniqueness appear to be largely orthogonal. Our findings confirm that structural and psycholinguistic evaluation capture independent dimensions of the therapy content quality and could be applied together for potential deployment decisions.
Exploring Generative Augmentation for Rare Disease Pathology: a Hirschsprung Disease Case Study
Authors: Mara Pistellato, Miriam Duci, Nicola Galiazzo, Leonardo Maccari, Francesco Fascetti Leon, Francesca Uccheddu
Student paper: NO
Keywords: Generative AI; Hirschsprung disease; GAN; Diffusion models; detection; Rare Diseases
Abstract: While artificial intelligence (AI) has significantly transformed digital pathology in high-volume fields like oncology, its application in rare diseases remains constrained by data scarcity. Indeed, small patient cohorts increase the risks of overfitting and poor generalization across institutions, hindering the adoption of decision-support tools. In this paper we explore the potential of generative AI approaches for synthetic data augmentation to possibly overcome these limitations and develop novel methodological frameworks. Our study focuses on Hirschsprung disease (HD), a rare congenital neuropathology (affecting 1 in 5,000 births) where diagnosis depends on the identification of absent ganglion cells and hypertrophic nerve fibers. Although AI-supported assessment has shown feasibility, its performance is currently restricted by small dataset sizes. To overcome this, we exploit generative augmentation to expand training data while ensuring synthetic data remain clinically relevant without distorting morphological features or amplifying existing biases. We explore different approaches starting from different baselines, highlighting the potential of introducing generative AI samples in the diagnosis path.
The Model Understands, the Rules Decide: A Domain-Specific Language and Deterministic Runtime for Safe Conversational Triage Training
Authors: Marco Graziano, Stefano Vitali
Student paper: NO
Keywords: Clinical decision support systems; Conversational AI agents; Domain-specific languages; Emergency triage; Large language models; Patient safety
Abstract: Large language models (LLMs) understand colloquial patient language well enough to conduct a triage interview, but they cannot be trusted to assign the triage priority code: handed the verbatim decision tables of an official triage manual, a single-prompt LLM produced 42 dangerous under-triage errors in 100 critical cases. We present a system that separates these two competencies of understanding language and deciding the code, architecturally rather than behaviorally. The solution is built with LGDL, general-purpose software previously developed by the author for conversational AI agents whose behavior is governed rather than generated. The LGDL software comprises a domain-specific language (the Language-Game Definition Language, in which an author declares a conversation's moves, slots, guards, and actions) and a deterministic runtime that compiles and executes those declarations: matching user input through a single supervised LLM call per turn, conducting the declared interviews, invoking external capabilities, and recording every decision in a per-turn audit trail. The triage interview is one such LGDL game: symptom-card routing scopes, the questions of the official Manuale Regionale Triage Modello Lazio a cinque codici of the Italian Lazio Region, conditional slots, escalation semantics, while a deterministic rules engine, mechanically transcribed from the same manual, computes the priority code as min(vitals-band code, discriminator code) with raise-only severity floors; the LLM's only role is language comprehension, and it is architecturally unable to participate in the code decision. Uncertainty ("I don't know") and missing data can only raise priority, never lower it. On a 590-item labeled Italian corpus covering ten encoded symptom cards, the system reaches 93.0% chief-complaint routing accuracy (370 utterances), 99.0% final-code agreement (100 scripted interviews), 100% colloquial slot extraction (120 answers), and zero under-triage in every configuration and run, at about USD 0.002 per interview; the same LLM prompted end-to-end reaches only 54% agreement. The intended use is educational - a training simulator for triage nurses - and all encoded rules are pending review by clinical partners.
Flow Builder: no-code conversation design tool for digital therapeutics in psychology
Authors: Leonardo Sanna, Mattia Franzin, Mauro Dragoni, Claudio Eccher
Student paper: NO
Keywords: Conversational Agents; Digital Therapeutics; No-code Design Tool
Abstract: The effectiveness of Digital Therapeutics (DTx) relies on understanding and shaping human behavior, yet development is hindered by fragmented conversational scripts and repetitive chatbot implementation. Integrating advanced agentic components such as Large Language Models (LLMs) and multi-agent systems further complicates clinical safety, data provenance, and ethical risks. We present Flow Builder, a no-code platform within the MAESTRO (Multi-Agent Enhanced System for Therapy Routine Orchestration) framework, designed to standardize conversation design. Leveraging a modular block-based hierarchy with over seventeen specialized processors, from text delivery to semantic retrieval and multi-step LLM reasoning, Flow Builder enables clinicians to focus on therapeutic logic while ensuring safe, consistent, and scalable agentic execution.
SAFE-Info: A Human-Centered Protocol for Evaluating Generative AI-Produced Patient Information in Community Healthcare
Authors: Francesco Germini, Gianluca Fedelfranco
Student paper: NO
Keywords: generative artificial intelligence; patient information; health literacy; human-centered AI; evaluation protocol; community healthcare
Abstract: Generative artificial intelligence can rapidly draft patient-facing information about access to services, preparation for procedures, prevention, and care pathways. Fluency, however, does not establish factual correctness, source fidelity, completeness, safety, understandability, or actionability. This work-in-progress paper proposes SAFE-Info, a human-centered protocol for evaluating generative-AI-produced patient information before organizational release in community healthcare. SAFE-Info combines four domains: Source faithfulness; Accuracy and omission control; Fitness for understanding and action; and Equity, accessibility, and safety. The protocol specifies a reference packet, versioned generation record, claim-level review, severity taxonomy, independent human assessment, and a traffic-light release decision. It integrates established instruments such as the Patient Education Materials Assessment Tool and the CDC Clear Communication Index with healthcare-specific controls for unsupported claims, critical omissions, service-access errors, and reviewer workload. A prospective pilot is outlined across four classes of territorial-care information, comparing unconstrained and source-grounded generation. The framework does not authorize AI use or replace clinical, legal, privacy, or organizational governance; it evaluates a defined class of outputs after a use case has been approved for testing. Proposed thresholds are preliminary and require validation for reliability, feasibility, and citizen comprehension.
Where the Risk Lives: Human Junctions, Not Systems - Ten Categories of AI in Healthcare and the Non-Delegable Objective Function
Authors: Angelo Rossi Mori
Student paper: NO
Keywords: AI in healthcare; taxonomy; human oversight; objective function; agentic systems; clinical decision support; health technology assessment; accountability
Abstract: Debate on artificial intelligence in healthcare is impaired by a single label covering objects that differ in technique, risk profile and applicable law. This position paper proposes ten categories of systems and then argues a stronger claim: even after categories are distinguished, risk is not a property of the system but of its human junction, the point where machine elaboration meets human decision. The same decision-support system is governable in a hospital team with time to discuss and hazardous for an isolated, overloaded general practitioner. We characterise human junctions through the geometry of roles, a symmetric comparison of the human error a system corrects and the machine error a human would catch, and resilience under failure. We replace the rule that a human must decide last with a triad of placements - in, on, and before the loop - and a conservation principle: delegation never removes human decision, it relocates it from the act to the rule, and what is pathological is undeclared relocation. Ten categories cluster into six junction profiles, and agentic systems push all of them towards the human further upstream. The non-delegable core is ownership of the objective function and of its periodic revision, whose proper seat, where legitimate objectives are plural and conflicting, is deliberative and multi-stakeholder.
MIAO: A Mental Illness Analysis Ontology for Detecting Mental Health Conditions
Authors: Gianluca Apriceno, Sergio Muñoz, Tania Bailoni, Mauro Dragoni, Carlos Á. Iglesias
Student paper: NO
Keywords: Mental Health; Mental Illness Detection; Ontologies; Knowledge Representation
Abstract: Mental health disorders affect more than one billion people worldwide and represent one of the leading causes of long-term disability. Despite recent progress in artificial intelligence for mental health support, current systems continue to exhibit critical limitations, including the generation of unreliable or harmful outputs in scenarios where expert knowledge is essential. These shortcomings highlight the need for structured, explicit, and interoperable representations of domain knowledge to support safe and effective mental health detection. While several ontologies have been proposed in this domain, existing efforts often focus on specific disorders, lack integration, or overlook the detection process itself. In this work, we present the Mental Illness Analysis Ontology (MIAO), a novel ontology designed to model the mental illness detection process independently of any particular diagnostic framework or detection method. MIAO provides: (i) a general conceptual model that describes the key elements of mental illness detection, including the steps involved, the participants, and the types of evidence used across both human and AI-based approaches, and (ii) a clear separation between the detection process and the underlying mental illness representations, enabling seamless integration with diverse mental health ontologies. By offering a unified and extensible framework, MIAO aims to support the development of hybrid, knowledge-driven AI systems and facilitate interoperability across mental health applications.
Assessing Mental Health Therapeutic Capacity in AI Agents
Authors: Piergiorgio Maruotti, Mattia Rampazzo, Sergio Muñoz, Mauro Dragoni, Patrizio Bellan
Student paper: NO
Keywords: Agent Personality Framing; Mental Health; Empathy; Theory of Mind; Multi-Agent Systems
Abstract: Large Language Models (LLMs) have been widely investigated for conversational support in mental health. However, their therapeutic reliability in empathy and perspective-taking remains uncertain. We aim to fill this gap by introducing a comprehensive evaluation of the Pool of Experts framework, a multi-agent approach that enables role-specific identities without retraining the underlying model. Beyond its effectiveness, this framework provides a controlled test-bed to study whether personality framing induces measurable behavioral variation across roles and tasks. We systematically assess the Pool of Experts capability in empathy and theory-of-mind-oriented benchmarks through question-answering tasks, evaluating accuracy via strict and relaxed match against gold answers, and compare structured multi-agent orchestration with less structured conditions. Results demonstrate three main findings. First, architectural orchestration with deliberative aggregation consistently improves performance: a Final Decision Maker agent improves accuracy by up to 4.1 percentage points compared with individual experts. This aggregation mechanism also provides superior error recovery compared to majority voting. Second, such improvements generalize robustly across model families without requiring larger parameter scales. Compact and large-scale architectures achieve comparable benefits. Third, process-oriented frameworks yield the strongest gains, while theory-of-mind tasks benefit more reliably from structured orchestration than empathy tasks. This suggests perspective-taking is currently more tractable for optimization than affective alignment. These findings support structured multi-agent orchestration as a reliable way to improve socio-cognitive behaviors in mental-health-oriented LLM systems.
Cross-Continental Generalizability of EHR Foundation Model: Predicting Length of Stay in Italy Using US-Trained Embeddings
Authors: Giovanna Nicora, Francesca Negri, Valentina Tibollo, Alberto Malovini, Arianna Dagliati, Lucia Sacchi, Federico Sottotetti, Paola Baiardi, Riccardo Bellazzi
Student paper: NO
Keywords: Clinical Foundation Models; Electronic Health Records; Cross-Continental Generalizability; Length of Stay Prediction; OMOP Common Data Model; Medical Event Data Standard
Abstract: Background: Clinical Foundation Models (FMs) hold the potential to exhibit generalist medical AI capabilities. However, their evaluation often relies on locally-limited datasets, raising questions regarding their generalizability across diverse global healthcare systems, populations, and clinical practices. This study assesses the cross-continental transferability of a US-trained FM, CLMBR-T-base, applied to a large, distinct, Italian hospital cohort to predict Length of Stay (LoS) Methods: We utilized a retrospective EHR dataset from IRCCS Maugeri, Pavia, Italy, comprising 99,884 admissions for 63,030 unique patients (2012-2025). To ensure interoperability, raw data were harmonized using the OMOP Common Data Model and converted to the Medical Event Data Standard (MEDS). Patient embeddings were generated zero-shot using the pretrained CLMBR-T-Base model. We benchmarked supervised regression models (XGBoost, Multilayer Perceptron) trained on these embeddings against traditional count-based baselines to predict log-transformed LoS. We built different models trained on the complete hospital admissions set and stratified across four clinical macro-categories (Surgical, Rehabilitation, Oncology, Generic). Embedding cohesion was validated using latent space analysis (HDBSCAN clustering). Results: The best overall model utilizing FM embeddings achieved an R2 of 0.705. Performance varied significantly by department type, proving lower in Rehabilitation (R2=0.33) and Oncology (R2=0.17). Latent space analysis confirmed that the FM embeddings preserved high semantic transferability, autonomously organizing Italian patients into clinically coherent groups. Conclusion: Clinical FMs trained on US data exhibit significant generalizability, effectively capturing "universal" clinical semantics across distinct international healthcare architectures when mediated by rigorous data standardization pipelines (OMOP/MEDS).
Imprecise Probabilities for Privacy-Accuracy Trade-Offs in Bayesian Networks
Authors: Niccolò Rocchi, Fabio Stella, Cassio de Campos
Student paper: YES
Keywords: Bayesian Networks; Credal Networks; Imprecise Probabilities; Membership Inference Attacks; Privacy-Accuracy Trade-Offs; Model Privacy
Abstract: Bayesian networks are well-known probabilistic graphical models that enable explainable reasoning under uncertainty. In many domains, data are scarce and fragmented across institutions, motivating collaboration and the sharing of learned models rather than individual-level records to increase the available evidence and reduce bias. Releasing a model, however, can still reveal sensitive information: adversaries may mount membership inference attacks to determine whether a specific individual contributed to the training. A common mitigation strategy is to inject noise into the learned model before its release. This perturbation may compromise the quality and interpretability of subsequent inferences. We investigate how a Bayesian network can be effectively masked rather than perturbed to protect it against attacks by leveraging credal networks, the imprecise version of Bayesian networks, allowing us to employ sets of parameters instead of exact point estimates. We formalize and compare various masking and attack strategies and investigate how privacy leakage depends on these strategies. Finally, we analyze how privacy and utility can be traded off in credal naive Bayes models, comparing them with the common noise-injection baseline. Results indicate that credal masking provides principled protection, substantially reducing attack success while yielding fine-tuned privacy-accuracy trade-offs.
Standardizing Retrieval-Augmented Generation Pipelines in Medical Domains
Authors: Leonardo Sanna, Esin Ezgi Yildiz, Mauro Dragoni
Student paper: NO
Keywords: Retrieval-Augmented Generation; Information Retrieval; Reciprocal Rank Fusion
Abstract: The design and evaluation of Retrieval-Augmented Generation pipelines remain fragmented, with system components often tailored to specific datasets and tasks. In this paper, we investigate a simple yet effective RAG pipeline to provide a standardized baseline for system engineering and retrieval evaluation. Our experiments show that a hybrid pipeline combining a retriever, reranking at each step, and Reciprocal Rank Fusion (RRF), yields consistent improvements across datasets, outperforming retrieval-only and individually reranked outputs. These results highlight a robust and generalizable baseline configuration for medical Retrieval-Augmented Generation systems, enabling researchers to focus on task-specific optimizations such as query augmentation, model selection, and retrieval refinement.
TRUTH-Lung: A Multimodal Dataset of CT Scans, Expert Segmentations, and Molecular Biomarkers for Lung Nodule Radiogenomics
Authors: Andrea Santomauro, Luigi Portinale, Giorgio Leonardi, Francesca Ugo, Ivan Gallesio, Costanza Massarino, Alessia Francese, Annalisa Roveta, Antonio Maconi
Student paper: NO
Keywords: Lung cancer; Lung nodules; Computed tomography; Medical imaging dataset; Histopathology; Immunohistochemistry; Radiogenomics; Multimodal learning
Abstract: Background: Early detection and characterization of pulmonary nodules are critical for improving lung cancer outcomes. While thoracic computed tomography (CT) is the primary imaging modality for lung cancer screening, radiological assessment alone is often insufficient to reliably determine malignancy, histological subtype, or underlying molecular features. The development of artificial intelligence methods to support lung nodule analysis is limited by the scarcity of publicly available datasets that integrate imaging data with confirmed histopathological and molecular ground truth. Objective: The aim of this work is to introduce TRUTH-Lung, a publicly available multimodal dataset that links thoracic CT imaging with expert nodule annotations, histopathological diagnoses, and structured immunohistochemical biomarker information, enabling research in lung nodule analysis, radiogenomics, and multimodal machine learning. Methods: The dataset comprises 545 retrospective chest CT scans collected at a single academic hospital between 2021 and 2024, including 280 scans with lung nodules and 265 control scans without nodules. For nodule-positive cases, three-dimensional segmentation masks were manually generated by experienced radiologists. Histopathological confirmation was available for almost all nodules and provided as original pathology reports in PDF format. A custom rule-based pipeline was developed to parse unstructured histology reports and extract standardized diagnostic labels, immunohistochemical marker status, and quantitative biomarker scores. Imaging, anatomical, diagnostic, and molecular data were linked using anonymized unique identifiers. Results: TRUTH-Lung provides a balanced cohort for both detection and classification tasks and includes a wide spectrum of nodule sizes, histological subtypes, and molecular profiles. The dataset supports multiple benchmark tasks, including lung nodule detection, segmentation, malignancy classification, histological subtype prediction, and quantitative biomarker estimation. All data were anonymized and validated for internal consistency across modalities. The dataset is publicly released through Hugging Face under a Creative Commons Attribution 4.0 International license. Conclusions: TRUTH-Lung addresses key limitations of existing public lung nodule datasets by integrating thoracic CT imaging with confirmed histopathological and molecular ground truth. By enabling multimodal and radiogenomic research, this dataset provides a valuable resource for the development and evaluation of advanced artificial intelligence methods for lung cancer screening and diagnosis.
Please contact us at hc-aixia@googlegroups.com.