Credibility Analysis of User‐Designed Content Using Machine Learning Techniques

Milind Gayakwad*, Suhas Patil, Amol Kadam, Shashank Joshi, Ketan Kotecha, Rahul Joshi*, Sharnil Pandya, Sudhanshu Gonge, Suresh Rathod, Kalyani Kadam, Maya Shelke

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

14 Citations (Scopus)
3 Downloads (Pure)

Abstract

Content is a user‐designed form of information, for example, observation, perception, or review. This type of information is more relevant to users, as they can relate it to their experience. The research problem is to identify the credibility and the percentage of credibility as well. Assessment of such content is important to convey the right understanding of the information. Different techniques are used for content analysis, such as voting the content, Machine Learning Techniques, and manual assessment to evaluate the content and the quality of information. In this research article, content analysis is performed by collecting the Movie Review dataset from Kaggle. Features are extracted and the most relevant features are shortlisted for experimentation. The effect of these features is analyzed by using base regression algorithms, such as Linear Regression, Lasso Regression, Ridge Regression, and Decision Tree. The contribution of the research is designing a heterogeneous ensemble regression algorithm for content credibility score assessment, which combines the above baseline methods. Moreover, these factors are also toned down to obtain the values closer to Gradient Descent minimum. Different forms of Error Loss, such as Mean Absolute Error, Mean Squared Error, LogCosh, Huber, and Jacobian, and the performance is optimized by introducing the balancing bias. The accuracy of the algorithm is compared with induvial regression algorithms and ensemble regression separately; this accuracy is 96.29%.

Original languageEnglish
Article number43
Number of pages15
JournalApplied System Innovation
Volume5
Issue number2
DOIs
Publication statusPublished - 14 Apr 2022
Externally publishedYes

Keywords

  • content analysis
  • regression loss analysis
  • the credibility of content based on score

Cite this