nutch杂记

lc87624

浏览: 144696 次
性别:
来自: 北京

最近访客更多访客>>

mingletxt

gk1987

不爱不见

bestchenwu

博主相关

博客

微博

相册

留言

关于我

文章分类

社区版块

存档分类

博客分类：

nutch

nutch

1. 如何绕过目标站点的robots.txt限制
多数站点都是只允许百度、google等搜索引擎抓取的，所以会在robots.txt里限制其他爬虫。
nutch自然是会遵循robots协议的，但是我们可以通过修改nutch源码来绕过限制。
相关代码位于（nutch版本1.5.1，其他版本未测试）：
org.apache.nutch.fetcher.Fetcher的run方法.
找到以下几行代码并注释掉就OK了。

if (!rules.isAllowed(fit.u)) {
                // unblock
                fetchQueues.finishFetchItem(fit, true);
                if (LOG.isDebugEnabled()) {
                  LOG.debug("Denied by robots.txt: " + fit.url);
                }
                output(fit.url, fit.datum, null, ProtocolStatus.STATUS_ROBOTS_DENIED, CrawlDatum.STATUS_FETCH_GONE);
                reporter.incrCounter("FetcherStatus", "robots_denied", 1);
                continue;
              }

2. url掉转导致html parse不成功的问题
在抓取百度百科的景点数据时，发现部分页面不会走html parse部分的逻辑，而我的plugin是基于HtmlParserFilter扩展点的，因而没有生效。
后来发现请求部分页面的链接返回的http状态为301，跳转之后才会到真正页面，而nutch默认是不会抓取跳转后的页面的.这时需要修改nutch-site.xml，加入以下配置即可，nutch-default.xml里的默认值是0，我们这里改成一个大于0的值，nutch就会继续抓取跳转后的页面了。

<property>
  <name>http.redirect.max</name>
  <value>2</value>
  <description>The maximum number of redirects the fetcher will follow when
  trying to fetch a page. If set to negative or 0, fetcher won't immediately
  follow redirected URLs, instead it will record them for later fetching.
  </description>
</property>

3. 抽取的过程中发现某些属性老是抽不到，而在不使用nutch抓取的情况下是能抽到的，进而怀疑nutch抓取的页面不全。于是去google了一下"nutch content limit"，发现nutch有这么一个配置项：

<property>
  <name>http.content.limit</name>
  <value>65536</value>
  <description>The length limit for downloaded content using the http://
  protocol, in bytes. If this value is nonnegative (>=0), content longer
  than it will be truncated; otherwise, no truncation at all. Do not
  confuse this setting with the file.content.limit setting.
  </description>
</property>

用来限制抓取内容的大小，放大10倍后，问题解决。
需要注意的是nutch还有一个很容易混淆的配置项：

<property>
  <name>file.content.limit</name>
  <value>65536</value>
  <description>The length limit for downloaded content using the file://
  protocol, in bytes. If this value is nonnegative (>=0), content longer
  than it will be truncated; otherwise, no truncation at all. Do not
  confuse this setting with the http.content.limit setting.
  </description>
</property>

两个配置用于的协议不同，前者是http协议，后者是file协议，我一开始就配置错了，折腾了半天。。。

PS：最后推荐两篇介绍nutch的文章，在官方文档不那么给力的情况下，这两篇文章给了我不小的帮助，感谢下作者。
http://www.atlantbh.com/apache-nutch-overview/ 对nutch的整体流程做了介绍
http://www.atlantbh.com/precise-data-extraction-with-apache-nutch/ 用实际例子介绍了nutch plugin的开发和部署

分享到：

linux下如何将命令行输出通过pipe直接copy ... | 记公司邮件组里的一次sql优化讨论

2012-08-08 18:25
浏览 7087
评论(3)
分类:开源软件
查看更多

3 楼 qdj6679 2013-05-10

我想请问一下楼主，如果不遵守robot.txt协议。那势必会找到封锁，比如说可能会跳转到输入验证码的页面，这个如何解析呢？

2 楼 qdj6679 2013-05-09

谢谢你的经验总结，解我燃眉之急啊

1 楼冷色日光 2012-09-13

谢谢你的文章，对我很有帮助啊，我把那个file.content.limit和http.content.limit弄混了，开了它才明白，解决了一个困扰了半天的困难。

发表评论

您还没有登录,请您登录后再发表评论

最近访客更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论

nutch杂记

评论

发表评论

相关推荐

最近访客 更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论

nutch杂记

评论

发表评论

相关推荐

nutch inject源码阅读笔记

最近访客更多访客>>