网页抓取之新方法 (在java程序中使用jQuery)

Mybeautiful

浏览: 301296 次
性别:
来自: 武汉

最近访客更多访客>>

choulisa

wgx19830922

zpf82118

lgl4223939

博主相关

博客

微博

相册

留言

关于我

文章分类

社区版块

存档分类

博客分类：

Java 综合

爬虫网页抓取 Rhino javascript

你想要的任何信息，基本上在互联网上存在了，问题是如何把它们整理成你所需要的，比如在某个行业网站上抓取所有相关公司的的名字，联系电话，Email等，然后存到Excel里面做分析。网页信息抓取变得原来越有用了。

一般传统的网页，web服务器直接返回Html，这类网页很好抓，不管是用何种方式，只要得到html页面，然后做Dom解析就可以了。但对于需要Javascript生成的网页，就不那么容易了。张瑜目前也没有找到好办法解决此问题。各位有抓javascript网页经验的朋友，欢迎指点。

所以今天要谈的还是传统html网页的信息抓取。虽然前面说了，没有技术难度，但是是否能有相对更容易的方法呢？用过jQuery等js框架的朋友，可能都会觉得javascript貌似抓取网页信息的天然助手，而且其出生就是为了网页解析而存在的。当然现在有更多的应用了，如Server端的javascript应用，NodeJs.

如果能在我们的应用程序，如java程序中，能使用jQuery去抓网页，绝对是件激动人心的事情。确实有现成的解决方案，一个Javascript引擎，一个能支撑jQuery运行的环境就可以了。

工具 : java, Rhino, envJs. 其中 Rhino是Mozzila提供的开源Javascript引擎，envJs是一个模拟浏览器额环境，如Window等。代码如下，

package stony.zhang.scrape;


import java.io.FileNotFoundException;
import java.io.FileReader;
import java.io.IOException;
import java.lang.reflect.InvocationTargetException;

import org.mozilla.javascript.Context;
import org.mozilla.javascript.ContextFactory;
import org.mozilla.javascript.Scriptable;
import org.mozilla.javascript.ScriptableObject;

/**
 * @author MyBeautiful
 * @Emal: zhangyu0182@sina.com
 * @date Mar 7, 2012
 */
public class RhinoScaper {
	private String url;
	private String jsFile;

	private Context cx;
	private Scriptable scope;

	public String getUrl() {
		return url;
	}

	public String getJsFile() {
		return jsFile;
	}

	public void setUrl(String url) {
		this.url = url;
		putObject("url", url);
	}

	public void setJsFile(String jsFile) {
		this.jsFile = jsFile;
	}

	public void init() {
		cx = ContextFactory.getGlobal().enterContext();
		scope = cx.initStandardObjects(null);
		cx.setOptimizationLevel(-1);
		cx.setLanguageVersion(Context.VERSION_1_5);

		String[] file = { "./lib/env.rhino.1.2.js", "./lib/jquery.js" };
		for (String f : file) {
			evaluateJs(f);
		}
		
		try {
			ScriptableObject.defineClass(scope, ExtendUtil.class);
		} catch (IllegalAccessException e1) {
			e1.printStackTrace();
		} catch (InstantiationException e1) {
			e1.printStackTrace();
		} catch (InvocationTargetException e1) {
			e1.printStackTrace();
		}
		ExtendUtil util = (ExtendUtil) cx.newObject(scope, "util");
		scope.put("util", scope, util);
	}

	protected void evaluateJs(String f) {
		try {
			FileReader in = null;
			in = new FileReader(f);
			cx.evaluateReader(scope, in, f, 1, null);
		} catch (FileNotFoundException e1) {
			e1.printStackTrace();
		} catch (IOException e1) {
			e1.printStackTrace();
		}
	}

	public void putObject(String name, Object o) {
		scope.put(name, scope, o);
	}

	public void run() {
		evaluateJs(this.jsFile);
	}
}

测试代码：

package stony.zhang.scrape;

import java.util.HashMap;
import java.util.Map;

import junit.framework.TestCase;

public class RhinoScaperTest extends TestCase {

	public RhinoScaperTest(String name) {
		super(name);
	}

	public void testRun() {
		RhinoScaper rs = new RhinoScaper();
		rs.init();
		rs.setUrl("http://www.baidu.com");
		rs.setJsFile("test.js");
//		Map<String, String> o = new HashMap<String, String>();
//		rs.putObject("result", o);
		rs.run();
//		System.out.println(o.get("imgurl"));
	}

}

test.js文件，如下

$.ajax({
  url: "http://www.baidu.com",
  context: document.body,
  success: function(data){
 //   util.log(data);
    
    var result =parseHtml(data);
    
    var $v= jQuery(result);
 //   util.log(result);
    $v.find('#u a').each(function(index) {
         util.log(index + ': ' + $(this).attr("href"));
  //        arr.add($(this).attr("href"));
    });
  }
});


 function parseHtml(html) {
       //Create an iFrame object that will be used to render the HTML in order to get the DOM objects
        //created - this is a far quicker way of achieving the HTML to DOM conversion than trying
        //to transform the HTML objects one-by-one
         var oIframe = document.createElement('iframe');
     //Hide the iFrame from view
         oIframe.style.display = 'none';
         if (document.body)
            document.body.appendChild(oIframe);
        else
            document.documentElement.appendChild(oIframe);
        
        //Open the iFrame DOM object and write in our HTML
        oIframe.contentDocument.open();
        oIframe.contentDocument.write(html);
        oIframe.contentDocument.close();
    
        //Return the document body object containing the HTML that was just
        //added to the iFrame as DOM objects
        var oBody = oIframe.contentDocument.body;
    
        //TODO: Remove the iFrame object created to cleanup the DOM
    
        return oBody;
    }

我们执行Unit Test，将会在控制台打印从网页上抓取的三个baidu的连接，

0: http://www.baidu.com/gaoji/preferences.html
1: http://passport.baidu.com/?login&tpl=mn
2: https://passport.baidu.com/?reg&tpl=mn

测试成功，故证明在java程序中用jQuery抓取网页是可行的.

----------------------------------------------------------------------

张瑜，Mybeautiful , zhangyu0182@sina.com

Java学习这七年  如何阅读源代码 我应该做的更差吗？

Rhino-test.zip (2 MB)
下载次数: 397

查看图片附件

4
顶

3
踩

分享到：

如何抓取需要验证码的网页？ | 节日重定义

2012-03-07 13:57
浏览 11716
评论(8)
分类:编程语言
查看更多

8 楼 Mybeautiful 2014-11-04

hanjiangit 写道

青峰大辉写道

你好，整个工程直接运行报错：
Exception in thread "main" org.mozilla.javascript.EvaluatorException: uncaught JavaScript runtime exception: ReferenceError: "util" is not defined. (./js/pair.js#3)

受累看下。

同问，楼主

我刚才测试了下，没有发现你们说的问题；附上我测试图片。我用的jdk1.7;不知是否有关。

7 楼 hanjiangit 2014-11-04

青峰大辉写道

同问，楼主

6 楼青峰大辉 2014-07-02

5 楼 Mybeautiful 2013-03-19

sbear 写道

楼主可以提供一下源码吗

389331837 写道

代码有错 ExtendUtil 这个类是在那里定义的呢？

yxzkm 写道

嗯，不错！不过，请看一下jsoup，似乎在服务端就能解决dom的遍历问题

对不起没有及时回复，已经把整个项目附上了，大家试试看。

4 楼 yxzkm 2013-02-05

嗯，不错！不过，请看一下jsoup，似乎在服务端就能解决dom的遍历问题

3 楼 sbear 2013-01-23

楼主可以提供一下源码吗

2 楼 389331837 2012-11-21

代码有错 ExtendUtil 这个类是在那里定义的呢？

1 楼 Mybeautiful 2012-03-09

补充一下，
经过研究，如果Rhino能结合jsdom那将能解决javascript的问题，就如同node.js一样。有相关经验的朋友，提示一下。

发表评论

您还没有登录,请您登录后再发表评论

最近访客更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论