NollySum-Net: A Multimodal Transformer Framework for Automatic Summarization of Nollywood Movies

Multimodal Transformer Framework for Automatic Summarization of Nollywood Movies

Authors

  • Tumenayu Ofut Ogar University of Cross River State , University of Cross River State image/svg+xml Author
    Competing Interests

    No conflict of interest.

DOI:

https://doi.org/10.67254/ucrjoset

Keywords:

African cinema; Cross-modal attention; Deep learning; Multimodal transformers; NollySum-Net; Nollywood; video summarization;

Abstract

The Nigerian film industry, Nollywood, is the second-largest film-producing industry globally, after Bollywood, producing over 2,500 films annually. Despite this huge volume of production, there is still a lack of automated tools for content analysis and summarisation in the specific context of Nollywood films with unique narrative structures, multilingual dialogue (English, Yoruba, Hausa, Igbo and Pidgin), and culturally peculiar visual attributes. In this paper, we introduce NollySum-Net, a novel multimodal deep learning model for automatic summarisation of Nollywood movies. Our method combines vision transformers for visual feature extraction, wav2vec 2.0 for audio processing in multiple languages, and a cross-modal attention mechanism to combine visual, audio, and textual features. To address these challenges, we present a new dataset, named NollySum-1K, consisting of 1000 annotated Nollywood movie clips with shot-level importance scores, scene boundaries and multilingual transcripts.  We provide a detailed analysis of the problems faced by Nollywood movies, like non-linear narratives, long dialogues and culturally referenced scene shifts. To the best of our knowledge, this work constitutes the first large-scale study of automatic summarization specifically focused on African cinema, establishing a baseline for future research in this domain

References

Adejunmobi, M. (2012, March 22). Nollywood and New Templates for Minor Transnational Film [Conference presentation]. Society for Cinema and Media Studies Conference, Boston, MA, United States.

Apostolidis, K., et al. (2021). Video summarization using deep neural networks: A systematic review. ACM Computing Surveys, 53(4), 1–36.

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. Proceedings of the International Conference on Learning Representations (ICLR).

Ekwuazi, H. (2011). Film in Nigeria. Africa World Press.

Gong, Yihong & Liu, Xin. (2001). Summarizing Video By Minimizing Visual Content Redundancies. Proceedings - IEEE International Conference on Multimedia and Expo. 10.1109/ICME.2001.1237793.

Gygli, M., Grabner, H., Riemenschneider, H., & Van Gool, L. (2014). Creating summaries from user videos. Proceedings of the European Conference on Computer Vision (ECCV), 505-520.

Mahasseni, B., Lam, M., & Todorovic, S. (2017). Unsupervised video summarisation with adversarial LSTM networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 202-211.

Ofut, O. T., Bakpo, F. S., & Ebem, D. U. (2022). Nollywood movie sequences summarization using a modified recurrent neural network model. Quest Journals: Journal of Software Engineering and Simulation, 8(1), 60-72.

Omoniyi, Tope. (2014). A borderlands' perspective of language and globalization. International Journal of the Sociology of Language. 2014. 10.1515/ijsl-2013-0085.

Park, J., Lee, J. & Sohn, K. Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization. International Journal Computer Vision 133, 8617–8641 (2025). https://doi.org/10.1007/s11263-025-02577-2.

Ye, Linwei & Rochan, Mrigank & Liu, Zhi & Wang, Yang. (2019). Cross-Modal Self-Attention Network for Referring Image Segmentation. 10.48550/arXiv.1904.04745.

Vaswani, Ashish & Shazeer, Noam & Parmar, Niki & Uszkoreit, Jakob & Jones, Llion & N.Gomez, Aidan & Kaiser, Lukasz & Polosukhin, Illia. (2025). Attention Is All You Need. 10.65215/r5bs2d54.

Yuan, Y., Li, X., & Zhu, Y. (2019). Multimodal attention network for video summarization. Proceedings of the 27th ACM International Conference on Multimedia (MM), 1235-1243.

Zhang H.J., Wu J., Zhong D., Smoliar S.W. An integrated system for content-based video retrieval and browsing. Pattern Recognit. 1997;30:643–658. doi: 10.1016/S0031-3203(96)00109-4

Zhang, Ke & Chao, Wei-Lun & Sha, Fei & Grauman, Kristen. (2016). Video Summarization with Long Short-Term Memory. 10.1007/978-3-319-46478-7_47.

Zhou, Kaiyang & Qiao, Yu & Xiang, Tao. (2018). Deep Reinforcement Learning for Unsupervised Video Summarization with Diversity-Representativeness Reward. Proceedings of the AAAI Conference on Artificial Intelligence. 32. 10.1609/aaai.v32i1.12255.

Downloads

Published

2026-09-16

How to Cite

NollySum-Net: A Multimodal Transformer Framework for Automatic Summarization of Nollywood Movies: Multimodal Transformer Framework for Automatic Summarization of Nollywood Movies. (2026). Unicross Journal of Science, Engineering and Technology, Formerly Called Crutech Journal of Science, Engineering and Technology, 1(2), 9-18. https://doi.org/10.67254/ucrjoset

Similar Articles

1-10 of 12

You may also start an advanced similarity search for this article.