Crawling on the World Wide Web

Wang, Li and Fox, Edward A. (2002) Crawling on the World Wide Web. Technical Report TR-02-10, Computer Science, Virginia Tech.

Full text available as:
PDF - Requires Adobe Acrobat Reader or other PDF viewer.
LiWangReportAccept.pdf (268273)

Abstract

As the World Wide Web grows rapidly, a web search engine is needed for people to search through the Web. The crawler is an important module of a web search engine. The quality of a crawler directly affects the searching quality of such web search engines. Given some seed URLs, the crawler should retrieve the web pages of those URLs, parse the HTML files, add new URLs into its buffer and go back to the first phase of this cycle. The crawler also can retrieve some other information from the HTML files as it is parsing them to get the new URLs. This paper describes the design, implementation, and some considerations of a new crawler programmed as an learning exercise and for possible use for experimental studies.

Item Type:	Departmental Technical Report
Subjects:	Computer Science > Information Retrieval
ID Code:	572
Deposited By:	User, Eprints
Deposited On:	27 June 2002