Wednesday, March 23, 2016

Topic modeling in sentiment analysis: A systematic review

 

Abstract

With the expansion and acceptance of Word Wide Web, sentiment analysis has become progressively popular research area in information retrieval and web data analysis. 

Due to the huge amount of user-generated contents over blogs, forums, social media, etc., sentiment analysis has attracted researchers both in academia and industry, since it deals with the extraction of opinions and sentiments. 

In this paper, we have presented a review of topic modeling, especially LDA-based techniques, in sentiment analysis. 

We have presented a detailed analysis of diverse approaches and techniques, and compared the accuracy of different systems among them. 

The results of different approaches have been summarized, analyzed and presented in a sophisticated fashion. 

This is the really effort to explore different topic modeling techniques in the capacity of sentiment analysis and imparting a comprehensive comparison among them.

References


Liu, B., Sentiment Analysis and Opinion Mining, Synthesis Lectures on Human Language Technologies, 5(1), pp. 1-167, 2012.


Pang, B. & Lee, L., Opinion Mining and Sentiment Analysis, Foundations and Trends in Information Retrieval, 2(1-2), pp. 1-135, 2008.


Hu, M. & Liu, B., Mining Opinion Features in Customer Reviews, in Proceedings of the Nineteenth National Conference on Artificial Intelligence (AAAI-04), San Jose, USA, vol. 4, pp. 755-760, July 2004.


Hu, M. & Liu, B., Mining and Summarizing Customer Reviews, in Proceedings of the tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-04), Washington, USA, pp. 168-177, ACM, Aug. 2004.


Zhang, L. & Liu, B., Aspect and Entity Extraction for Opinion Mining, Data Mining and Knowledge Discovery for Big Data, pp. 1-40, Springer Berlin Heidelberg, 2014.


Kitchenham, B. A. & Mendes, E., A Comparison of Cross-Company and Within-Company Effort Estimation Models for Web Applications, in Proceedings of the 8th International Conference on Empirical Assessment in Software Engineering (EASE-04), Edinburgh, Scotland, UK, pp. 47-55, May 2004.


Hofmann, T., Unsupervised Learning by Probabilistic Latent Semantic Analysis, Machine Learning, 42(1-2), pp. 177-196, 2001.


Blei, D. M., Ng, A. Y. & Jordan, M. I., Latent Dirichlet Allocation, The Journal of Machine Learning Research, 3, pp. 993-1022, 2003.


Fang, L. & Huang, M., Fine Granular Aspect Analysis Using Latent Structural Models, in Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics, Jeju, South Korea: Short Papers-Volume 2, pp. 333-337, Association for Computational Linguistics, July 2012.


Lin, Z., Jin, X., Xu, X., Wang, W., Cheng, X. & Wang, Y., A Cross-Lingual Joint Aspect/Sentiment Model for Sentiment Analysis, in Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management (CIKM-14), Shanghai, China, pp. 1089-1098, ACM, Nov. 2014.


Xueke, X., Xueqi, C. ,Songbo, T., Yue, L. & Huawei, S., Aspect-Level Opinion Mining of Online Customer Reviews, China Communications, 10(3), pp. 25-41, 2013.


Zhai, Z., Liu, B., Xu, H. & Jia, P., Constrained LDA for Grouping Product Features in Opinion Mining, Advances in knowledge discovery and data mining, pp. 448-459, Springer, 2011.


Moghaddam, S. & Ester, M., ILDA: Interdependent LDA Model for Learning Latent Aspects and their Ratings from Online Product Reviews, in Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval(SIGIR-11), Beijing, China, pp. 665-674, ACM, July 2011.


Brody, S. & Elhadad, N., An Unsupervised Aspect-Sentiment Model for Online Reviews, in Human Language Technologies: in Proceedings of the 11th Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT-10), Los Angeles, USA, pp. 804-812, Association for Computational Linguistics, June 2010.


Zhao, W.X., Jiang, J., Yan, H. & Li, X., Jointly Modeling Aspects and Opinions with a Maxent-LDA Hybrid, in Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing (EMNLP-10), Massachusetts, USA, pp. 56-65, Association for Computational Linguistics, Oct. 2010.


Jo, Y. & Oh, A. H., Aspect and Sentiment Unification Model for Online Review Analysis, in Proceedings of the Fourth ACM International Conference on Web Search and Data Mining (WSDM-11), Hong Kong, pp. 815-824, ACM, Feb. 2011.


Xu, X., Tan, S., Liu, Y., Cheng, X. & Lin, Z., Towards Jointly Extracting Aspects and Aspect-Specific Sentiment Knowledge, in Proceedings of the 21st ACM International Conference on Information and Knowledge Management (CIKM-12), Maui Hawaii, USA, pp. 1895-1899, ACM, Oct. 2012.


Mukherjee, A. & Liu, B., Aspect Extraction through Semi-Supervised Modeling, in Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics, Jeju, South Korea: Long Papers-Volume 1, pp. 339-348, Association for Computational Linguistics, July 2012.


Kim, S., Zhang, J., Chen, Z., Oh, A. H. & Liu, S., A Hierarchical Aspect-Sentiment Model for Online Reviews, in Proceedings of the Twenty-Seventh AAAI conference on Artificial Intelligence(AAAI-13), Washington, USA, July 2013.


[Bagheri, A., Saraee, M. & De Jong, F., ADM-LDA: An Aspect Detection Model Based on Topic Modelling Using the Structure of Review Sentences, Journal of Information Science, 40(5), pp. 621-636, 2014.


Gruber, A., Weiss, Y. & Rosen-Zvi, M., Hidden Topic Markov Models, in Proceeding of the 11th International Conference on Artificial Intelligence and Statistics (AISTATS-07), San Juan, Puerto Rico, pp. 163-170, Mar. 2007.


Wang, T., Cai, Y., Leung, H.-f., Lau, R. Y., Li, Q. & Min, H., Product Aspect Extraction Supervised with Online Domain Knowledge, Knowledge-Based Systems, 71, pp. 86-100, 2014.


Chen, Z., Mukherjee, A., Liu, B., Hsu, M., Castellanos, M. & Ghosh, R., R., Leveraging Multi-Domain Prior Knowledge in Topic Models, in Proceedings of the Twenty-Third international joint conference on Artificial Intelligence (IJCAI-13), Beijing, China, pp. 2071-2077, AAAI Press, Aug. 2013.


Chen, Z., Mukherjee, A., Liu, B., Hsu, M., Castellanos, M. & Ghosh, R., Discovering Coherent Topics Using General Knowledge, in Proceedings of the 22nd ACM international conference on Conference on information & knowledge management (CIKM-13), San Francisco, USA, pp. 209-218, ACM, Oct. 2013.


Chen, Z., Mukherjee, A., Liu, B., Hsu, M., Castellanos, M. & Ghosh, R., Exploiting Domain Knowledge in Aspect Extraction, in Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP-13), Seattle, USA, pp. 1655-1667, Oct. 2013.


Chen, Z., Mukherjee, A. & Liu, B., Aspect Extraction with Automated Prior Knowledge Learning, in Proceedings of the 52nd Annual Meeting of the Association of Computational Linguistics (ACL-214), Baltimore, USA, pp. 347-358, June 2014.


Han, J., Cheng, H., Xin, D. & Yan, X., Frequent Pattern Mining: Current status and Future Directions, Data Mining and Knowledge Discovery, 15(1), pp. 55-86, 2007.


Rosen-Zvi, M., Chemudugunta, C., Griffiths, T., Smyth, P. & Steyvers, M., Learning Author-Topic Models from Text Corpora, ACM Transactions on Information Systems (TOIS), 28(1), pp. 1-38, 2010.


Chen, Z. & Liu, B., Topic Modeling Using Topics from Many Domains, Lifelong Learning and Big Data, in Proceedings of the 31st International Conference on Machine Learning (ICML-14), Beijing, China, pp. 703-711, June 2014.


Chen, Z. & Liu, B., Mining Topics in Documents: Standing on the Shoulders of Big Data, in Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-14), New York City, USA, pp. 1116-1125, ACM, Aug. 2014.





DOI: http://dx.doi.org/10.5614%2Fitbj.ict.res.appl.2016.10.1.6

http://journals.itb.ac.id/index.php/jictra/article/view/1442

Monday, December 7, 2015

Emoji Sentiment Ranking


.
Abstract
There is a new generation of emoticons, called emojis, that is increasingly being used in mobile communications and social media. In the past two years, over ten billion emojis were used on Twitter. Emojis are Unicode graphic symbols, used as a shorthand to express concepts and ideas. In contrast to the small number of well-known emoticons that carry clear emotional contents, there are hundreds of emojis. But what are their emotional contents? We provide the first emoji sentiment lexicon, called the Emoji Sentiment Ranking, and draw a sentiment map of the 751 most frequently used emojis. The sentiment of the emojis is computed from the sentiment of the tweets in which they occur. We engaged 83 human annotators to label over 1.6 million tweets in 13 European languages by the sentiment polarity (negative, neutral, or positive). About 4% of the annotated tweets contain emojis. The sentiment analysis of the emojis allows us to draw several interesting conclusions. It turns out that most of the emojis are positive, especially the most popular ones. The sentiment distribution of the tweets with and without emojis is significantly different. The inter-annotator agreement on the tweets with emojis is higher. Emojis tend to occur at the end of the tweets, and their sentiment polarity increases with the distance. We observe no significant differences in the emoji rankings between the 13 languages and the Emoji Sentiment Ranking. Consequently, we propose our Emoji Sentiment Ranking as a European language-independent resource for automated sentiment analysis. Finally, the paper provides a formalization of sentiment and a novel visualization in the form of a sentiment bar.
.
http://kt.ijs.si/data/Emoji_sentiment_ranking/index.html

Thursday, October 1, 2015

Cardiff Castle Panoramic View

360° panorama of the grounds of Cardiff Castle, showing (l to r) the interpretation centre, the barbican and South Gate, the Black Tower, the Clock Tower and the main range, the reconstructed Roman Wall, the shell keep on the motte, and the Norman banked earth defences

Cardiff Castle (Welsh: Castell Caerdydd) is a medieval castle and Victorian Gothic revival mansion located in the city centre of Cardiff, Wales. The original motte and bailey castle was built in the late 11th century by Norman invaders on top of a 3rd-century Roman fort. The castle was commissioned either by William the Conqueror or by Robert Fitzhamon, and formed the heart of the medieval town of Cardiff and the Marcher Lord territory of Glamorgan. In the 12th century the castle began to be rebuilt in stone, probably by Robert of Gloucester, with a shell keep and substantial defensive walls being erected. Further work was conducted by Richard de Clare, 6th Earl of Gloucester, in the second half of the 13th century. Cardiff Castle was repeatedly involved in the conflicts between the Anglo-Normans and the Welsh, being attacked several times in the 12th century, and stormed in 1404 during the revolt of Owain Glyndŵr.

Thursday, May 28, 2015

Classification of Sentiment Analysis on Tweets using Machine Learning Techniques [Problem Statement]



.

1.3 Problem Statement

Given a set of tweets containing multiple features and varied opinions, the objective is to extract expressions of opinion describing a target feature and classify it as positive or negative.


1.4 Motivation
Sentiment analysis and opinion minion [14] is open research field with manifold real life applications. Blogs, Forum, Twitter, Facebook and other resources on internet are put to use by humans for expressing their opinions. The social media has bought the people around the world closer; communication is one click away. Before social media there was expensive short messaging service (SMS) provided by telecommunication companies with domestic and international charges. Today the short messaging has evolved from just sending messages to single person to sending messages to multiple people at cheapest price. This service is provided by many websites but Twitter was the one which pioneered it. Today twitter has hundreds of millions users who post nearly half a billions tweets every day i.e. approximately thousands of tweets for every second. Tweets are not only posted in English language but also in different local languages of the world. These data are precious to business intelligence where the company wants to know "why isn't consumer buying our laptops?", "why the competitors products are outselling our products". Thus a concrete system to process above mentioned queries is the need of the hour.


1.5 Objective:

Classify every tweet in either as positive sentiment or negative sentiment using different Machine Learning techniques and check which classifier performs the best.

.

https://core.ac.uk/download/pdf/80148155.pdf

Tuesday, March 31, 2015

Sentiment Analysis: Text Pre-Processing, Reader Views and Cross Domains [Problem Statement]

 


.

1.1 BACKGROUND AND PROBLEM DEFINITION

.

In computational linguistics, sentiment analysis is considered to be a classification problem. It involves natural language processing (NLP) on many levels, and inherits its challenges. There exists a wide variety of applications that could benefit from its results, such as news analytics, marketing, question answering, knowledge bases and so on. The challenge of this field is to improve the machine’s ability to understand texts in the same way as human readers are able to. Taking advantages from the huge amount of opinions expressed on the internet especially from social media blogs is vital for many companies and institutions, whether it is in terms of product feedback, public mood, or investor opinions.

.

The present thesis searches into different possibilities to improve sentiment classification performance. To address this problem, three different key issues are investigated. The first issue is to improve sentiment classification through text preprocessing. The second issue is to improve it through utilising text properties. The third issue is to improve it through inferring sentiment from one domain to another. These issues are explained in the following.

.

...

.
1.2 AIM AND OBJECTIVES
.
The main aim of this thesis is to explore key ways of improving sentiment classification performance. To achieve this, there are three distinctive objectives. The first objective aims to improve sentiment prediction through text pre-processing. A wide variety of pre-processing methods is presented and an appropriate feature selection method is selected for the analysis. Document level sentiment classification is performed along with the focus on products reviews and the use of movie reviews as an example.
.
The second objective intends to improve sentiment classification through delving into various text properties. The example here is financial news that has two properties. Firstly, the financial news contains announcements of financial events that could be utilised in the sentiment prediction. A model that employs news events in sentiment classification is proposed. Secondly, financial news allows for capturing the investors (reader) opinions through stock market returns. It is argued in this thesis that in some tasks such as financial forecasting, it is the sentiment expressed in the responses of content readers (for instance, through trading behaviour) that may be more useful as a means of creating predictive models. A new model that is built to predict financial news sentiment based on a novel method to capture reader sentiment is presented.
.
Furthermore, the financial news covers a wide variety of different domains such as economics, accounting, law, etc. Therefore, the third objective aims to improve sentiment classification through investigating the case of cross-domain sentiment analysis. A method for selecting domain dependent and independent words is proposed, and a new model for cross domain sentiment analysis is evaluated against other approaches.
.
.

Monday, March 9, 2015

Cardiff change back from red to blue


Cardiff City have unveiled a new badge that will be worn on their kits from the 2015-16 season.

Owner Vincent Tan gave the go-ahead for the Championship club's home shirts to change back from red to blue and to make the Bluebird more prominent on the badge after consulting supporters.

Sian Branson, founder of the Bluebirds Unite group, which campaigned for the colour change, welcomed the move.

"At least I know I'm supporting CCFC when I look at this badge," she said.

"The future's blue and we don't have to feel as detached from our club any more."

Branson added there was "still plenty that needs to be done" and hoped the fans and club could continue to work together.

The club's new crest features an oriental dragon based on the one featured at Cardiff City Hall.


Source:
http://www.bbc.com/sport/football/31795873

http://www.walesonline.co.uk/sport/football/football-news/cardiff-citys-new-crest-revealed-8799490