AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP

This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022

Bibliografiske detaljer
Main Authors: Chowdhury, Salman Mostafiz, Rohid, Ali Ahammed, Hussain, Rizwan, Mostafa, Chowdhury Sujana
Andre forfattere: Mostakim, Moin
Format: Thesis
Sprog:en_US
Udgivet: Brac University 2022
Fag:
Online adgang:http://hdl.handle.net/10361/17114
id 10361-17114
record_format dspace
spelling 10361-171142022-08-23T21:01:38Z AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP Chowdhury, Salman Mostafiz Rohid, Ali Ahammed Hussain, Rizwan Mostafa, Chowdhury Sujana Mostakim, Moin Department of Computer Science and Engineering, Brac University Speech Recognition Wav2vec2 BERT AVSpeech Keywords extraction Neural Networks NLP Automatic speech recognition Neural networks (Computer science) This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022 Cataloged from PDF version of thesis. Includes bibliographical references (pages 39-43). Audiovisual data is the most extensively used and abundantly distributed type of data on the internet in today’s information and communication age. However, the necessary audiovisual data is challenging to retrieve because the majority of them are not correctly categorized. As a result, it is difficult to locate the necessary au diovisual data in times of need, and as a result, a great amount of potentially useful information does not reach users in a timely manner. A piece of data is only as good as the time frame in which it was acquired. Additionally, because audiovisual files such as lecture notes, recorded classes, and recorded conversations are quite large, skimming through a large amount of audiovisual data for the necessary information can be time consuming, and the likelihood of not receiving the appropriate infor mation on time is relatively high. As a result, we propose a novel model that will take any audio or video file as input and label it according to its content utilizing Convolutional Neural Networks, BERT, and several machine learning techniques. Our proposed model accepts any audiovisual file as input and extracts features from the contents and uses convolutional neural network and transformer to recognize and transcript the speeches of the conversations. Using BERT models and cosine similarity keywords and phrases are extracted from the transcript and the input file will be labeled with the key phrases and keywords that are most similar to the con text of the content. Finally, the input file will be appropriately labeled with these key phrases and keywords so that anyone in the future in need of similar information can quickly locate this audiovisual file. Salman Mostafiz Chowdhury Ali Ahammed Rohid Rizwan Hussain Chowdhury Sujana Mostafa B. Computer Science 2022-08-23T04:42:05Z 2022-08-23T04:42:05Z 2022 2022-01 Thesis ID: 17101149 ID: 17101361 ID: 17301092 ID: 21101107 http://hdl.handle.net/10361/17114 en_US Brac University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. 43 Pages application/pdf Brac University
institution Brac University
collection Institutional Repository
language en_US
topic Speech Recognition
Wav2vec2
BERT
AVSpeech
Keywords extraction
Neural Networks
NLP
Automatic speech recognition
Neural networks (Computer science)
spellingShingle Speech Recognition
Wav2vec2
BERT
AVSpeech
Keywords extraction
Neural Networks
NLP
Automatic speech recognition
Neural networks (Computer science)
Chowdhury, Salman Mostafiz
Rohid, Ali Ahammed
Hussain, Rizwan
Mostafa, Chowdhury Sujana
AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP
description This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022
author2 Mostakim, Moin
author_facet Mostakim, Moin
Chowdhury, Salman Mostafiz
Rohid, Ali Ahammed
Hussain, Rizwan
Mostafa, Chowdhury Sujana
format Thesis
author Chowdhury, Salman Mostafiz
Rohid, Ali Ahammed
Hussain, Rizwan
Mostafa, Chowdhury Sujana
author_sort Chowdhury, Salman Mostafiz
title AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP
title_short AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP
title_full AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP
title_fullStr AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP
title_full_unstemmed AVCL: Audio Video clustering for learning Conversation labeling using Neural Network and NLP
title_sort avcl: audio video clustering for learning conversation labeling using neural network and nlp
publisher Brac University
publishDate 2022
url http://hdl.handle.net/10361/17114
work_keys_str_mv AT chowdhurysalmanmostafiz avclaudiovideoclusteringforlearningconversationlabelingusingneuralnetworkandnlp
AT rohidaliahammed avclaudiovideoclusteringforlearningconversationlabelingusingneuralnetworkandnlp
AT hussainrizwan avclaudiovideoclusteringforlearningconversationlabelingusingneuralnetworkandnlp
AT mostafachowdhurysujana avclaudiovideoclusteringforlearningconversationlabelingusingneuralnetworkandnlp
_version_ 1814308004173971456