资源列表
UpdateAddrIndex
- 电信行业,编写的地址搜索引擎的代码,功能是更新索引库的类-Telecommunications industry, to write the address of the search engine code update the index library class
GenWordlib
- 电信行业,编写的地址搜索引擎的代码,功能是产生词典,用于分词。-Telecommunications industry, the address written in the code of the search engine, the function is to generate dictionary for word.
GenAddrIndexAll
- 电信行业,编写的地址搜索引擎,此类是用于建立地址库的源代码-Telecommunications industry, write the address search engine, such is used to establish the source code of the address database
GenAddrSegmIndex
- 电信行业,地址搜索的程序,此代码功能是根据区域,对更新索引库-Telecom industry, Address Search program, this code function is based on the region, the index is updated library
TokenTest
- 电信行业,此代码是地址搜索程序的一部分,该代码的功能是分词的测试程序。-Telecommunications industry, address search program, the function of this code is written in the sub-word test.
MemCache
- memcache缓存使用,能够减轻数据库压力,接口非常简单-memcache cache use database can reduce pressure, the interface is very simple.
this-is-search-engine
- 一本关于搜索引擎的书籍,强调原理而不纠缠技术细节-Books of a search engine, stressed that the principle of not entangled technical details
Spider-Java
- 网络爬虫的简要介绍及一点源代码,分享给想要学习爬虫的人-The web crawler brief introduction and point-source code
wordbag
- 根据一个人物名单文件,查找wekipedia上相应网页,读取网页文本,并统计每个人物在每个网页上出现的次数,最终形成word bag,人物500人,运行时间6分钟左右。-from a namelist making a word bag
GetWeb
- 以下是一个Java爬虫程序,它能从指定主页开始,按照指定的深度抓取该站点域名下的网页并维护简单索引。-The following is a Java reptiles, it can start from the specified Home to crawl pages under the domain name of the site in accordance with the specified depth and maintain a simple index.
LoalaSam
- 功能强大地搜索引擎,实现简单,源程序风,封装好的-The powerful search engine, to achieve a simple, good the source wind, package
SearchEngin_NewSOO
- 一个垂直搜索引擎例子希望对大家有所帮助。包含了主体核心代码。-A vertical search engine example, we want to help. Contains the main core code.