Volume : VIII, Issue : III, March - 2019

Audio Description of Image

Parth Zalavadia, Vrutti Patel, Ronaksingh Sahni, Siddhesh Bhagat

Abstract :

Audio Description of Image is a cornerstone in computer vision, a field which has seen much technical advancement in recent years. Describing the contents and audio of images is a challenge for the machine to achieve it. It requires not only accurate recognition of object and human, but also their attributes and relationship as well as scene information. Audio Description of Image makes initial attempts to deal with the above challenges to produce multi sentence natural language description of image contents and producing the audio of the same. It takes a local region based on the approach to extract regional image details and by combines multiple technique including attribute learning and deep learning through the use of machine learned features to create high level labels that can generate detailed audio description of real–world images. Audio Description of Image will contain the function of scene classification, object detection and classification, attribute learning, relationship detection, sentence generation and audio generation. Currently caption generation systems do exist, however they are shå resources to create the RoI (Region of Interest) layer. LSTM (Long Short Term Memory) provides a better approach to meaningful sentence creation using the different objects detected in the image using Convolutional Neural Network (CNN). By blending CNN and LSTM, it is possible to considerably reduce the sum of time and resources consumed by the system, also the caption generated by Audio Description of Image will be semantically more accurate.

Keywords :

CNN   RNN   LSTM   RoI   technology.  

Article: Download PDF   DOI : 10.36106/ijsr  

Cite This Article:

AUDIO DESCRIPTION OF IMAGE, Parth Zalavadia, Vrutti Patel, RonakSingh Sahni, Siddhesh Bhagat INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH : Volume-8 | Issue-3 | March-2019


Number of Downloads : 696


References :