The article by Henzinger et al., Challenges in Web Search Engines, examines some problems with search engines that I never considered before. Of course, it's obvious with all the advertisements popping up on the query result pages that money runs the operation, and the more you have, the more people will see your website. I was unaware of the term "spam" being applied to search engines, however. As with some algorithms, the most popular web pages appear first in the query list. With others, it appears authors and publishers of these sites have bought their ticket to popularity, which will ultimately lead to more links to that page by all other algorithms.
Part One of the Hawking article focuses on generic terms related to web searches. While I found some of his diagrams to be a little confusing, he supports them with several definitions, such as URL, crawling, indexes, spamming, and hashing function as they apply to search engines. He spends a great deal of time examining web crawling and various forms of crawling, which filters previously generated search results from the same and related queries. The second part of this article examines the process indexing by crawling by search engines, namely Google, Yahoo, and Microsoft (GYM). Typically, there are two steps to indexing by algorithms. First, the "indexer" scans the text of documents. Next, it sorts the numbered results by "inversion" in term number order, with the document secondarily numbered. Hawking also describes processes later in the indexing process, such as scaling and compression, for ease of storage.
No comments:
Post a Comment