Multimodal architecture for interpretable image analysis.
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13677Keywords:
Model formulation, Identification and validation, Decision support and control, Biomedical system modeling, simulation and visualization, Quantification of physiological parameters for diagnosis and treatment assessmentAbstract
Artificial intelligence models have demonstrated remarkable performance, although their limited interpretability poses a challenge in fields where understanding the decisions made is crucial. This challenge is particularly significant in the clinical setting, such as in the analysis of breast cancer, a disease characterized by high variability in its manifestations. To address this limitation, we propose a multimodal approach that uses a simplified synthetic dataset combining images with a vector of additional data to guide the model’s learning. For interpretable analysis, the SHAP tool is applied to evaluate the contributions of both inputs: on the one hand, the image, and on the other, the vector variables. The results demonstrate the viability of the proposed approach, suggesting as future lines of work the optimization of the model to reduce classification errors and improve the detection of complex cases, as well as its validation using real-world data.
References
Alzamil, D., Alkhamees, B., Hassan, M. M., 2025. A systematic review of multimodal fusion and explainable ai applications in breast cancer diagnosis. Computer Modeling in Engineering & Sciences 145 (3), 2971.
Lundberg, S., 2018. Welcome to the shap documentation. https://shap.readthedocs.io/en/latest/, accessed: 2026-05-06.
Lundberg, S., 2024. Shap image interpretation examples. https://shap.readthedocs.io/en/latest/image_examples.html, accessed: 2026-03-31.
Ovalle, M. T. O., 2026. Inteligencia artificial explicable (xai) como marco para la transparencia en la ciencia de datos: Interpretabilidad, toma de decisiones etica y responsabilidad algor´ıtmica. Revista Cient´ıfica de Salud y Desarrollo Humano 7 (1), 258–283.
Perez, E., Strub, F., De Vries, H., Dumoulin, V., Courville, A., 2018. Film: Visual reasoning with a general conditioning layer. In: Proceedings of the AAAI conference on artificial intelligence. Vol. 32.
Pezzini, M. C., 2024. Inteligencia artificial explicable: an´alisis de metodologıas y aplicaciones. Ph.D. thesis, Universidad Nacional de La Plata.
Ribeiro, M. T., Singh, S., Guestrin, C., 2016. ”why should i trust you?.explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. pp. 1135–1144.
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D., 2017. Grad-cam: Visual explanations from deep networks via gradientbased localization. In: Proceedings of the IEEE international conference on computer vision. pp. 618–626.
Smithuis, R., Wijers, L., Dennert, I., 2026. Ultrasound of the breast. https://radiologyassistant.nl/breast/ultrasound/ultrasound-of-the-breast, accessed March 2026.
Soler, J. R., 2017. Herramientas computacionales basadas en imagen de ultrasonografıa para diagn´ostico m´edico. Ph.D. thesis, Universidad Miguel Hern´andez de Elche.
Stahlschmidt, S. R., Ulfenborg, B., Synnergren, J., 2022. Multimodal deep learning for biomedical data fusion: a review. Briefings in bioinformatics 23 (2), bbab569.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Sergio Lidon Calvo, Marina Poveda Pérez, Marta Nadal Herraiz, Juliana Manrique Córdoba, José María Sabater Navarro

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.