skip to main content
10.1145/1718487.1718542acmconferencesArticle/Chapter ViewBasic AbstractPublication PageswsdmConference Proceedingsconference-collections
Several features on this page require Premium Access.
You are using the Basic Edition. Features requiring a subscription appear in grey.
research-article
Free access

Boilerplate detection using shallow text features

Published: 04 February 2010 Publication History
Additional metrics (Premium feature)

Abstract

In addition to the actual content Web pages consist of navigational elements, templates, and advertisements. This boilerplate text typically is not related to the main content, may deteriorate search precision and thus needs to be detected properly. In this paper, we analyze a small set of shallow text features for classifying the individual text elements in a Web page. We compare the approach to complex, state-of-the-art techniques and show that competitive accuracy can be achieved, at almost no cost. Moreover, we derive a simple and plausible stochastic model for describing the boilerplate creation process. With the help of our model, we also quantify the impact of boilerplate removal to retrieval performance and show significant improvements over the baseline. Finally, we extend the principled approach by straight-forward heuristics, achieving a remarkable detection accuracy.

Formats available

You can view the full content in the following formats:

References

[1]
G. Altmann. Quantitative Linguistics -- an international handbook, chapter Diversification processes. de Gruyter, 2005.
[2]
S. Baluja. Browsing on small screens: recasting web-page segmentation into an efficient machine learning framework. In WWW '06: Proceedings of the 15th international conference on World Wide Web, pages 33--42, New York, NY, USA, 2006. ACM.
[3]
Z. Bar-Yossef and S. Rajagopalan. Template detection via data mining and its applications. In WWW, pages 580--591, 2002.
[4]
M. Baroni, F. Chantree, A. Kilgarriff, and S. Sharoff. Cleaneval: a competition for cleaning web pages. In N. Calzolari, K. Choukri, B. Maegaard, J. Mariani, J. Odjik, S. Piperidis, and D. Tapias, editors, Proceedings of the Sixth International Language Resources and Evaluation (LREC'08), 2008.
[5]
D. Cai, S. Yu, J.-R. Wen, and W.-Y. Ma. Extracting content structure for web pages based on visual representation. In X. Zhou, Y. Zhang, and M.E. Orlowska, editors, APWeb, volume 2642 of LNCS, pages 406--417. Springer, 2003.
[6]
D. Chakrabarti, R. Kumar, and K. Punera. Page-level template detection via isotonic smoothing. In WWW '07: Proc. of the 16th int. conf. on World Wide Web, pages 61--70, New York, NY, USA, 2007. ACM.
[7]
D. Chakrabarti, R. Kumar, and K. Punera. A graph-theoretic approach to webpage segmentation. In WWW '08: Proceeding of the 17th international conference on World Wide Web, pages 377--386, New York, NY, USA, 2008. ACM.
[8]
Y. Chen, W.-Y. Ma, and H.-J. Zhang. Detecting web page structure for adaptive viewing on small form factor devices. In WWW '03: Proceedings of the 12th international conference on World Wide Web, pages 225--233, New York, NY, USA, 2003. ACM.
[9]
S. Debnath, P. Mitra, N. Pal, and C.L. Giles. Automatic identification of informative sections of web pages. IEEE Trans. on Knowledge and Data Engineering, 17(9):1233--1246, 2005.
[10]
S. Evert. A lightweight and efficient tool for cleaning web pages. In N. Calzolari, K. Choukri, B. Maegaard, J. Mariani, J. Odjik, S. Piperidis, and D. Tapias, editors, Proceedings of the Sixth International Language Resources and Evaluation (LREC'08), Marrakech, Morocco, may 2008. European Language Resources Association (ELRA). http://www.lrec-conf.org/proceedings/lrec2008/.
[11]
D. Fernandes, E.S. de Moura, B. Ribeiro-Neto, A.S. da Silva, and M.A. Gonçalves. Computing block importance for searching on web sites. In CIKM '07, pages 165--174, 2007.
[12]
A. Ferraresi, E. Zanchetta, M. Baroni, and S. Bernardini. Introducing and evaluating ukwac, a very large web-derived corpus of english. In Proceedings of the WAC4 Workshop at LREC 2008.
[13]
A. Finn, N. Kushmerick, and B. Smyth. Fact or fiction: Content classification for digital libraries. Joint DELOS-NSF Workshop on Personalisation and Recommender Systems in Digital Libraries (Dublin), 2001.
[14]
D. Gibson, K. Punera, and A. Tomkins. The volume and evolution of web page templates. In WWW'05, pages 830--839, New York, NY, USA, 2005. ACM.
[15]
J. Gibson, B. Wellner, and S. Lubar. Adaptive web-page content identification. In WIDM '07: Proceedings of the 9th annual ACM international workshop on Web information and data management, pages 105--112, New York, NY, USA, 2007. ACM.
[16]
K. Hofmann and W. Weerkamp. Web corpus cleaning using content and structure. In Building and Exploring Web Corpora, pages 145--154. UCL Presses Universitaires de Louvain, September 2007.
[17]
H.-Y. Kao, J.-M. Ho, and M.-S. Chen. Wisdom: Web intrapage informative structure mining based on document object model. Knowledge and Data Engineering, IEEE Transactions on, 17(5):614--627, May 2005.
[18]
C. Kohlschütter. A densitometric analysis of web template content. In WWW '09: Proc. of the 18th intl. conf. on World Wide Web, New York, NY, USA, 2009. ACM.
[19]
C. Kohlschütter and W. Nejdl. A Densitometric Approach to Web Page Segmentation. In ACM 17th Conf. on Information and Knowledge Management (CIKM 2008), 2008.
[20]
I. Ounis, C. Macdonald, M. de Rijke, G. Mishne, and I. Soboroff. Overview of the trec 2006 blog track. In E.M. Voorhees and L.P. Buckland, editors, TREC, volume Special Publication 500--272. National Institute of Standards and Technology (NIST), 2006.
[21]
J. Pasternack and D. Roth. Extracting article text from the web with maximum subsequence segmentation. In WWW '09: Proceedings of the 18th international conference on World wide web, pages 971--980, New York, NY, USA, 2009. ACM.
[22]
C.E. Shannon. A mathematical theory of communication. Bell system techn. journal, 27, 1948.
[23]
M. Spousta, M. Marek, and P. Pecina. Victor: the web-page cleaning tool. In WaC4, 2008.
[24]
K. Vieira, A.S. da Silva, N. Pinto, E.S. de Moura, a.M.B.C. Jo and J. Freire. A fast and robust method for web page template detection and removal. In CIKM '06: Proc. 15th ACM int. conf. on Information and knowledge management, pages 258--267, 2006.
[25]
R. Vulanovic and R. Köhler. Quantitative Linguistics -- An international Handbook, chapter Syntactic units and structures, pages 274--291. de Gruyter, 2005.
[26]
L. Yi, B. Liu, and X. Li. Eliminating noisy information in web pages for data mining. In KDD '03: Proc. of the 9th ACM SIGKDD int. conf. on Knowledge discovery and data mining, pages 296--305, 2003.

Cited By

View all
  • (2026)Dripper: Token-Efficient Main HTML Extraction with a Lightweight LMProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.210.1145/3770855.3817915(3258-3269)Online publication date: 9-Aug-2026
  • (2026)HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training DataProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.110.1145/3770854.3783937(2196-2207)Online publication date: 9-Aug-2026
  • (2026)Modeling Completion Time in Mathematics Formative Assessments: Content-Based Prediction of Time VariationArtificial Intelligence in Education10.1007/978-3-032-29755-6_21(309-323)Online publication date: 25-Jun-2026
  • Show More Cited By

Index Terms

Index Terms (Premium feature)
  1. Boilerplate detection using shallow text features

    Recommendations

    Comments

    Comments (Premium feature)