Cribado de CSAM basado en visión artificial

Autores/as

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13766

Palabras clave:

Aprendizaje automático, Criminalidad, Procesamiento de imagen

Resumen

La detección de material de abuso sexual infantil (CSAM) es un desafío relevante en informática forense y protección de menores, cuya investigación está limitada por restricciones legales, éticas y de privacidad. Este trabajo presenta un pipeline modular de visión artificial para el cribado de riesgo de CSAM basado en tres señales visuales proxy: detección facial, estimación de edad aparente y clasificación de contenido adulto. El sistema integra YOLOv12n-face, SwinFace y ViT-B/16, aplicando una regla de decisión que marca una imagen ´cuando coinciden evidencia de contenido adulto y presencia de  menores. Además, se proponen dos datasets proxy de tres clases, APD-2M-3C y NPDI-2K-3C, orientados a diferenciar imágenes pornográficas, no pornográficas sencillas y no pornográficas difíciles. Al no utilizar CSAM real, la evaluación se centra en el rendimiento temporal y los falsos positivos sobre datasets sin CSAM. Los resultados preliminares muestran una capacidad eficiente para procesar grandes volúmenes de imágenes, aunque se requiere validación adicional antes de un uso operativo.

Referencias

1| Amirgaliyev, B., Mussabek, M., Rakhimzhanova, T., Zhumadillayeva, A. (2025). A review of machine learning and deep learning methods for person detection, tracking and identification, and face recognition with applications. Sensors 25. DOI: 10.3390/S25051410

2| Avila, S., Thome, N., Cord, M., Valle, E., ArauJo, A. D. A. (2013). Pooling in image representation: The visual codeword point of view. Computer Vision and Image Understanding 117(5), 453–465.

3| Coelho, T., Ribeiro, L. S. F., Macedo, J., dos Santos, J. A., Avila, S. (2024). Transformers-based few-shot learning for scene classification in child sexual abuse imagery. XXXVII Conference on Graphics, Patterns and Images (SIBGRAPI Estendido 2024), pp. 8–14. DOI: 10.5753/sibgrapi.est.2024.31638

4| Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. ICLR 2021. URL: https://arxiv.org/pdf/2010.11929

5| Gangwar, A., Gonzalez-Castro, V., Alegre, E., Fidalgo, E. (2021). Attm-cnn: Attention and metric learning based cnn for pornography, age and child sexual abuse (csa) detection in images. Neurocomputing 445, 81–104. DOI: 10.1016/j.neucom.2021.02.056

6| Gangwar, A., Gonzalez-Castro, V., Alegre, E., Fidalgo, E. (2023). Triple-biggan: Semi-supervised generative adversarial networks for image synthesis and classification on sexual facial expression recognition. Neurocomputing 528, 200–216. DOI: 10.1016/j.neucom.2023.01.027

7| Gangwar, A., Gonzalez-Castro, V., Alegre, E., Fidalgo, E., Martínez-Mendoza, A. (2024). Deephsar: Semi-supervised fine-grained learning for multi-label human sexual activity recognition. Information Processing and Management 61. DOI: 10.1016/j.ipm.2024.103800

8| Kaewkorn, H., Zhou, L., Li, W. (2025). Rp-net: A robust polar transformation network for rotation-invariant face detection. Pattern Recognition 158. DOI: 10.1016/J.PATCOG.2024.111044

9| Lee, H.-E., Ermakova, T., Ververis, V., Fabian, B. (2020). Detecting child sexual abuse material: A comprehensive survey. Forensic Science International: Digital Investigation 34, 301022. DOI: 10.1016/j.fsidi.2020.301022

10| Lin, C., Gou, J., Fan, Z., Liao, Y. (2024). A feature fusion-based resnet using the pooling pyramid for age estimation. Proceedings of the International Joint Conference on Neural Networks. DOI: 10.1109/IJCNN60899.2024.10650545

11| Liu, X., Qiu, M., Zhang, Z., Shi, Y., Li, Z., Chen, X., Liu, Y., Chen, X., Yu, H. (2025). Enhancing facial age estimation with local and global multi-attention mechanisms. Pattern Recognition Letters 189, 71–77. DOI: 10.1016/J.PATREC.2025.01.005

12| Morais, P. (2025). Aprendizaje profundo en la clasificación automatizada de imágenes en la detección de abusos sexuales a menores. (unpublished)

13| Moreira, D., Avila, S., Perez, M., Moraes, D., Testoni, V., Valle, E., Goldenstein, S., Rocha, A. (2016). Pornography classification: The hidden clues in video space–time. Forensic Science International 268, 46–61. DOI: 10.1016/J.FORSCIINT.2016.09.010

14| Pereira, M., Dodhia, R., Anderson, H., Brown, R. (2020). Metadata-based detection of child sexual abuse material. arXiv preprint arXiv:2010.02387.

15| Qin, L., Wang, M., Deng, C., Wang, K., Chen, X., Hu, J., Deng, W. (2023). Swinface: A multi-task transformer for face recognition, expression recognition, age estimation and attribute estimation. IEEE Transactions on Circuits and Systems for Video Technology 34, 2223–2234. DOI: 10.1109/TCSVT.2023.3304724

16| Rautela, K., Sharma, D., Kumar, V., Kumar, D. (2024). Obscenity detection transformer for detecting inappropriate contents from videos. Multimedia Tools and Applications 83, 10799–10814. DOI: 10.1007/s11042-023-16078-2

17| Shou, Y., Cao, X., Liu, H., Meng, D. (2025). Masked contrastive graph representation learning for age estimation. Pattern Recognition 158. DOI: 10.1016/J.PATCOG.2024.110974

18| Steinebach, M. (2023). An analysis of PhotoDNA. Proceedings of the 18th International Conference on Availability, Reliability and Security (ARES 2023). ACM. DOI: 10.1145/3600160.3605048

19| Tian, Y., Ye, Q., Doermann, D. (2025). Yolov12: Attention-centric real-time object detectors. URL: https://arxiv.org/abs/2502.12524

20| UNICEF (2026). Artificial intelligence and child sexual abuse and exploitation. URL: https://www.unicef.org/reports/artificial-intelligence-and-child-sexual-abuse-and-exploitation

21| Westlake, B., Brewer, R., Swearingen, T., Ross, A., Patterson, S., Michalski, D., Hole, M., Logos, K., Frank, R., Bright, D., et al. (2022). Developing automated methods to detect and match face and voice biometrics in child sexual abuse videos. Trends and issues in crime and criminal justice (648), 1–15.

22| Yang, J. (2025). Facial recognition optimization based on adversarial sample generation in the field of artificial intelligence. Discover Artificial Intelligence 5. DOI: 10.1007/S44163-025-00287-9

23| Yu, S., Zhao, Q. (2025). Improving age estimation in occluded facial images with knowledge distillation and layer-wise feature reconstruction. Applied Sciences (Switzerland) 15. DOI: 10.3390/APP15115806

24| Yu, W., Zhou, P., Yan, S., Wang, X. (2025). Inceptionnext: When inception meets convnext. URL: https://arxiv.org/abs/2303.16900

25| Zhang, J., Hu, N. (2025). Accuracy and robustness evaluation of deep learning algorithms in facial recognition systems. Systems and Soft Computing 7. DOI: 10.1016/J.SASC.2025.200252

26| Zhu, D., Shan, X., Wu, C., Yung, K., Ip, A. (2024). Multi frame obscene video detection with vit: An effective for detecting inappropriate content. International Journal on Semantic Web and Information Systems 20. DOI: 10.4018/IJSWIS.359768

Descargas

Publicado

01-09-2026

Número

Sección

Visión por Computador