A Conceptual Analysis of Face and Visual Speech Multimodal Fusion for Non-Vocal Biometric Authentication
DOI:
https://doi.org/10.71302/jbidai.v8i2.84Keywords:
multimodal biometrics, face recognition, score-level fusion, deepfake, liveness detectionAbstract
Biometric authentication systems still face fundamental limitations, particularly in unimodal approaches that are vulnerable to environmental variations and visual spoofing attacks. To address these challenges, multimodal biometrics integrating physiological and behavioral traits have become an increasingly relevant approach. This study presents an analytical review of recent research in visual multimodal biometrics, with a focus on score-level fusion strategies and the integration of static face recognition and dynamic lip movement analysis as a non-vocal authentication mechanism. The literature synthesis indicates that score-level fusion is the most flexible and stable approach for combining heterogeneous biometric modalities, especially when integrating static spatial features and dynamic temporal patterns. Furthermore, Transformer-based deep learning architectures are identified as having significant potential for modeling the temporal dependencies of lip movements. This study also highlights key security challenges, particularly presentation attacks and visual-only deepfakes, and emphasizes the importance of visual dynamics–based liveness detection as an integral component of biometric authentication systems. Based on these findings, the study formulates a conceptual framework for visual multimodal biometric authentication that integrates identity verification and liveness detection within a unified process, while also identifying future research opportunities, including self-supervised learning, model optimization for resource-constrained devices, and the design of more discriminative visual passphrases.
References
[1] A. Roihan et al., “Perancangan Purwarupa Sistem Keamanan Kunci Pintu Berbasis Pengenalan Wajah”, Journal of Innovation And Future Technology (IFTECH), vol. 6, no. 2, pp. 234–242, Aug. 2024, doi: 10.47080/iftech.v6i2.3415.
[2] A. Ross and A. K. Jain, “Multimodal biometrics: An overview,” Proc. 12th European Signal Processing Conference, 2004. doi: 10.1109/EUSIPCO.2004.7075379.
[3] F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A unified embedding for face recognition and clustering,” Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2015, pp. 815–823. doi: 10.1109/CVPR.2015.7298682.
[4] R. Raghavendra and C. Busch, “Presentation attack detection methods for face recognition systems: A comprehensive survey,” IEEE Access, vol. 7, pp. 100–132, 2019. doi: 10.1109/ACCESS.2019.2929433.
[5] J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3444–3453. doi: 10.1109/CVPR.2017.367.
[6] T. Afouras, J. S. Chung, and A. Zisserman, “Deep lip reading: A comparison of models and an online application,” IEEE/ACM Trans. Audio, Speech, and Language Processing, vol. 26, no. 12, pp. 2246–2260, 2018. doi: 10.1109/TASLP.2018.2810139.
[7] J. Yang, A. Waibel, and A. Jain, “Biometric recognition: Challenges and opportunities,” IEEE Computer, vol. 44, no. 1, pp. 74–80, 2011. doi: 10.1109/MC.2010.384.
[8] P. K. Ratha, K. Ricanek, and M. Savvides, “Score level fusion of multimodal biometrics using triangular norms,” Pattern Recognition Letters, vol. 32, no. 14, pp. 1843–1850, 2011. doi:10.1016/j.patrec.2011.06.029.
[9] O. N., Kadhim, M. H., Abdulameer, Y. M. H., Al-Mayali, “A multimodal biometric system for iris and face traits based on hybrid approaches and score level fusion,” In BIO Web of Conferences, vol. 97, pp. 00016, 2024. doi: 10.1051/bioconf/20249700016.
[10] S. R. B. Kisku, J. S. Chang, and A. Kumar, “Multimodal biometrics: Weighted score level fusion based on non-ideal iris and face images,” Expert Systems with Applications, vol. 41, no. 11, pp. 5390–5404, 2014. doi:10.1016/j.eswa.2014.02.051.
[11] M. He et al., “Performance evaluation of score level fusion in multimodal biometric systems,” Pattern Recognition, vol. 43, no. 5, pp. 1789–1800, 2010. doi:10.1016/j.patcog.2009.11.018.
[12] A. Naseem et al., “Robust multimodal biometric system based on optimal score level fusion model,” Expert Systems with Applications, vol. 116, pp. 364–376, 2019. doi:10.1016/j.eswa.2018.08.036.
[13] S. Tharewal et al., “Score-Level Fusion of 3D Face and 3D Ear for Multimodal Biometric Human Recognition,” Computational Intelligence and Neuroscience, 2022, Art. no. 3019194. doi:10.1155/2022/3019194.
[14] S. N. Garg, R. Vig, and S. Gupta, “A Survey on Different Levels of Fusion in Multimodal Biometrics,” Indian Journal of Science and Technology, vol. 10, no. 44, pp. 1–11, 2017. doi:10.17485/ijst/2017/v10i44/120575.
[15] F. Wang and J. Han, “Multimodal biometric authentication based on score level fusion using support vector machine,” Opto-Electronics Review, vol. 17, no. 1, pp. 59–64, 2009. doi:10.2478/s11772-008-0054-8.
[16] F., Shafizadegan, A. R., Naghsh-Nilchi, E. Shabaninia, “Multimodal vision-based human action recognition using deep learning: a review,” Artificial Intelligence Review, vol. 57, 178, 2024. doi:10.1007/s10462-024-10730-5
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.







