{"title":"A Systematic Literature Review of Recognition of Compound Facial Expression of Emotions","authors":"Sana Ullah, Wenhong Tian","doi":"10.1145/3447450.3447469","DOIUrl":"https://doi.org/10.1145/3447450.3447469","url":null,"abstract":"Recently, emotion recognition has gained increasing attention in various fields. Human facial expression plays key role in daily interaction among people. The understanding of different categories of facial expression is crucial in perceptual interfaces, computational models, and human cognition and even in every field of life. The basic facial expression of emotions is categorized into seven groups which are including neutral, disgust, fear, anger, surprise, sadness and happiness. The compound facial expression produced from the combination of some of the basic emotions. However, there has been no survey article on Compound Facial Expression of Emotions. In this article, we present survey on recognition, research methodologies and applications of Compound Facial Expression of Emotions. Diverse types of datasets/databases for the recognition of compound facial expressions of emotion with their pros and cons have been discussed on detail. Furthermore, different research methods along with advantages, disadvantages and possible future improvements are also deliberated. In the last, opportunities and challenges are enlisted.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"2016 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"121407551","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"On the Performance Evaluation of State-of-the-art Rate Control Algorithms for Practical Video Coding and Transmission Systems","authors":"Wei Gao","doi":"10.1145/3447450.3447479","DOIUrl":"https://doi.org/10.1145/3447450.3447479","url":null,"abstract":"The existing rate control (RC) algorithms are actually evaluated using single objectives, however many aspects can jointly affect the performances of practical video coding and transmission systems. In this paper, we first identify the key factors for RC performance assessment, and then introduce the joint RC performance (JRCP) evaluation method for comprehensive consideration of involved many objectives. Second, we investigate the effective strategies to normalize each item from the comparative perspective to make them aggregated for fair comparison. Finally, we conduct experiments based on the latest H.265/HEVC reference software to evaluate the performances of different state-of-the-art RC algorithms, and demonstrate the usage and applicability of proposed method. Additionally, the weighting strategy can be easily and flexibly configured according to the particular demands of practical video applications. Based on the proposed method, we also would like to draw the attention from the video coding and communication community, and the video encoder and transmission systems can be better optimized for visual experience enhancement in practical systems.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"28 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"114837088","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Hanjiang Chen, Zhengang Yang, Jinsong Liu, Kejia Wang
{"title":"Real-time Millimeter-wave 3D SAR Imaging with Orthogonal Bistatic Model","authors":"Hanjiang Chen, Zhengang Yang, Jinsong Liu, Kejia Wang","doi":"10.1145/3447450.3447483","DOIUrl":"https://doi.org/10.1145/3447450.3447483","url":null,"abstract":"To satisfy the real-time demand, both data sampling and data processing are required to be high efficient in a millimeter-wave 3D SAR imaging system. For fast data sampling, antenna array is applied, which requires multiple transmit antennas and receiving antennas. In this paper, an orthogonal bistatic model is adopted for 2D and 3D imaging, and this model can significantly decrease the number of antennas, the complexity and cost of the hardware are also reduced. For fast data processing, an image reconstruction algorithm based on wavenumber domain transform is proposed, compared to BP algorithm, the proposed algorithm has much higher compute efficiency. A 2D imaging simulation with the frequency of 55GHz is demonstrated, and the imaging time is only 0.0177s. At the same time, a 3D imaging simulation with frequencies from 50GHz to 60GHz is also demonstrated, the imaging time is only 0.371s, and all of the result images can be clearly shown.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"24 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"122915879","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Semantic Image Synthesis with Trilateral Generative Adversarial Networks","authors":"Benyuan Li, Yue Yu, Menglan Wang","doi":"10.1145/3447450.3447484","DOIUrl":"https://doi.org/10.1145/3447450.3447484","url":null,"abstract":"We present a new method that can synthesize photorealistic images from semantic label maps using GANs with a trilateral generator. Due to the limitation of network capacity, it is difficult for scene image generation networks to consider resolution and image fidelity at the same time. The solution is to use multi-path structure instead of the traditional single-path structure. In this work, we propose a novel trilateral Generative Adversarial Network (trilateral GAN), which has fewer parameters than other recent methods to synthesize 512*1024 images with high fidelity. Moreover, we improve the semantic consistency loss and feature matching loss by using the features before activation, which can make the synthetic images have sharper edges and richer textures. Finally, we replace all the Batch-Normalization (BN) layers with Spectral Normalization (SN) in the network and use Conditional Normalization Blocks (CNB) to avoid difficult convergence during training and make the synthesized images more realistic. The proposed network has a better performance and gets higher mIoU scores on COCO-Stuff, ADE20K and Cityscapes than the state-of-the-art methods.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"4 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"131803139","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"A Study of Student Learning Status Classification Based on the Detection of Key Objects within the Visual Field","authors":"Qiubo Huang, Yixuan Hua","doi":"10.1145/3447450.3447456","DOIUrl":"https://doi.org/10.1145/3447450.3447456","url":null,"abstract":"In order to improve students' concentration in solitary learning, this paper proposes a method to detect students' learning status. The students are first photographed, and then a Faster RCNN model is used to classify the learning status of the students in the photos. In order to improve the classification accuracy, the system detects the 2D face key points and matches them with the 3D face model to calculate the head angle based on the conversion between coordinates to determine the visual field of the eyes. Based on the visual field, it is possible to detect key objects that students may be focusing on, such as books, computers, cell phones. Based on these key objects, it can improve the accuracy of classifying students' learning status. The system can better identify learning status and help guardians to manage the students.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"95 Suppl 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"133818084","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Multi-attribute Recognition of Vehicles Based on the Multi-task Convolutional Neural Network","authors":"Y. Qi, Haorui Liu","doi":"10.1145/3447450.3447472","DOIUrl":"https://doi.org/10.1145/3447450.3447472","url":null,"abstract":"Recently, automatic vehicle identification system has great significance in practical life and with modern technology now widely available, the research on automatic vehicle recognition based on computer vision is becoming increasingly thorough and extensive. However, automatic vehicle identification also faces significant challenges, such as changes in lighting, weather conditions and distortion of captured images in natural road scenes. Since 2012, as a deep learning in image classification, the convolution neural network has become a major subject in many fields, so this study would focus on the vehicle identification in terms of the convolutional neural network. The Cars196 data set and an improved Convolutional Neural Network (CNN) network capable of multi-attribute recognition are applied to study the optimization network problem of the convolutional neutral network and to realize the multi-attribute classification and identification of automobile brands, models and colors. The data analysis shows that the improved CNN recognition network has a recognition accuracy in automatic vehicle recognition, but the overall recognition network is affected by many factors. According to the results of data analysis, in practice, alongside the improvement of traditional convolutional neural network, the improved new neural network structure should also be used for vehicle recognition, so as to improve the accuracy of vehicle identification to better serve the community.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"65 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"132646568","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"MI-FGSM on Faster R-CNN Object Detector","authors":"Zhenghao Liu, Wenyu Peng, Junhao Zhou, Zifeng Wu, Jintao Zhang, Yunchun Zhang","doi":"10.1145/3447450.3447455","DOIUrl":"https://doi.org/10.1145/3447450.3447455","url":null,"abstract":"The adversarial examples show the vulnerability of deep neural networks, which makes adversarial attacks widely concerned. However, most of the attack methods are based on image classification model. In this paper, we use Momentum Iterative Fast Gradient Sign Method (MI-FGSM), which stabilize optimization and escape from poor local maxima, to generate adversarial examples on the Faster R-CNN object detector. We have made some improvements on the previous object detection attack methods. The best current attack method, Project Gradient Descent (PGD) on object detection, starts from a random value, resulting in the uncertainty of the attack result. In contrast, our attacks are more stable and powerful in both white-box attacks and black-box attacks, and can better adapt to various neural network architectures. Experiment on Pascal VOC2007 shows that, under same setting of white-box attack, PGD has 0.23% mean average precision (mAP) on Faster R-CNN with VGG16, while our method achieves 0.17%. In addition, we analyze the difference between classification and detection attacks, and find that in addition to misclassification, the adversarial examples produced by detection models can also lead to mislocation.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"1 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"128843938","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Real time lane detection model based on Lightweight","authors":"Guangkun Zhai","doi":"10.1145/3447450.3447453","DOIUrl":"https://doi.org/10.1145/3447450.3447453","url":null,"abstract":"Many cars now have functions to assist the driver in driving, such as keeping the lane. This function can keep the vehicle in a proper position between lanes. This function can ensure the safety of the vehicle. The traditional lane detection method relies on image processing and has poor robustness. At the same time, due to the need for post-processing technology, it will generate a lot of time-consuming calculations, which is not conducive to use in changing road scenes. Based on this, this paper proposes a two-stage lane line detection method based on deep neural network, which divides lane detection into two stages, lane edge segmentation and lane line classification. In the first stage, a semantic segmentation network based on a lightweight network performs pixelated segmentation on the edge of lane lines. In the second stage, a long-term storage network based on GRU is used to detect lane lines based on lane edges. Practice has proved that the lane line detection method proposed in this paper is very reliable in various driving scenarios. At the same time, due to the use of various lightweight networks, the network model in this paper can be deployed on vehicle systems with powerful real-time performance.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"1 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"129667832","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"A String Prediction-Based Algorithm to AVS3 Screen Content Coding","authors":"Dong-Jin Jiang, Zhu Hong, Cheng Fang, Feiyang Zeng, Jucai Lin, Jun Yin","doi":"10.1145/3447450.3447480","DOIUrl":"https://doi.org/10.1145/3447450.3447480","url":null,"abstract":"String Prediction (SP) is a special intra prediction mode for screen content coding, which can provide better coding efficiency while keeping lower coding complexity. In this paper, we present a string prediction-based improved algorithm to third-generation of the Audio Video Coding Standard (AVS3) screen content coding (SCC). The main content is to use the string vector (SV) of history-based block vector prediction (HBVP) as string vector prediction (SVP), and obtain the best SV by comparing the rate-distortion (RD) cost, which comes from CU-level estimation search and pixel-level motion estimation search. SV is equal to the sum of SVP and string vector difference (SVD). While the encoder only needs to signal the SVD and the HBVP index to the decoder. The experiments show that with little coding complexity increases, the algorithm presented in this paper achieves an average Y BD-rate of -5.02% (TGM)/ -3.06% (MC), U BD-rate of -5.10% (TGM)/ -3.39% (MC), V BD-rate of -5.08% (TGM)/ -3.66% (MC) from AVS SCC common test condition (CTC) test suite in all intra configuration.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"1 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"129726561","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Tianbo Liu, Lihui Su, S. Yuan, Gong Cheng, Feng Zhang
{"title":"Research of single object tracking method based on Siamese Network and Level Set","authors":"Tianbo Liu, Lihui Su, S. Yuan, Gong Cheng, Feng Zhang","doi":"10.1145/3447450.3447477","DOIUrl":"https://doi.org/10.1145/3447450.3447477","url":null,"abstract":"In this paper, we propose a single object tracking method based on a Siamese network and level set to increase the accuracy of object tracking and segmentation while maintaining real-time tracking. First of all, feature extraction and similarity measurement are performed on the object template image and the search image by using the Siamese backbone network based on ResNet. Then the top-down object tracking contour refinement is performed on the feature map using the level set method. Finally, the loss function of the algorithm is defined, and the proposed algorithm is evaluated and compared on the VOT and DAVIS dataset. The experimental results illustrate that the proposed algorithm maintains the real-time performance, and at the same time the accuracy has been significantly improved.","PeriodicalId":389770,"journal":{"name":"Proceedings of the 2020 4th International Conference on Video and Image Processing","volume":"35 1","pages":"0"},"PeriodicalIF":0.0,"publicationDate":"2020-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"122145435","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}