[笔记]hadoop tutorial - Reducer -

GQM

浏览: 25224 次
性别:
来自: 上海

最近访客更多访客>>

wafer1021

melin

萝__卜

leoeco2000

博主相关

博客

微博

相册

留言

关于我

文章分类

社区版块

存档分类

[笔记]hadoop tutorial - Reducer

博客分类：

hadoop

引用

Reducer reduces a set of intermediate values which share a key to a smaller set of values.

Reducer的数量
可通过以下方法设置

JobConf.setNumReduceTasks(int);

可以修改mapred.reduce.tasks参数，默认值为1。
官网推荐计算方法

0.95 * NUMBER_OF_NODES * mapred.tasktracker.reduce.tasks.maximum
1.75 * NUMBER_OF_NODES * mapred.tasktracker.reduce.tasks.maximum

其中选择0.95时，所有的reducer task都会在map task结束时立即启动；选择1.75时有部分reducer task需要等到第二轮执行（从先结束的节点开始执行），这样可以更好的做到负载均衡。

引用

It is legal to set the number of reduce-tasks to zero if no reduction is desired.
In this case the outputs of the map-tasks go directly to the FileSystem, into the output path set by setOutputPath(Path). The framework does not sort the map-outputs before writing them out to the FileSystem.

Reducer阶段的工作

引用

Reducer has 3 primary phases: shuffle, sort and reduce.

Shuffle

Input to the Reducer is the sorted output of the mappers. In this phase the framework fetches the relevant partition of the output of all the mappers, via HTTP.

Sort

The framework groups Reducer inputs by keys (since different mappers may have output the same key) in this stage.
The shuffle and sort phases occur simultaneously; while map-outputs are being fetched they are merged.

Secondary Sort

If equivalence rules for grouping the intermediate keys are required to be different from those for grouping keys before reduction, then one may specify a Comparator via JobConf.setOutputValueGroupingComparator(Class). Since JobConf.setOutputKeyComparatorClass(Class) can be used to control how intermediate keys are grouped, these can be used in conjunction to simulate secondary sort on values.
辅助排序可用于values需要排序的场合，如果将key对应的所有values直接在内存中排序可能会造成OOM问题，通过辅助排序将key和value封装成复合key，JobConf.setOutputKeyComparatorClass(Class)控制排序；JobConf.setOutputValueGroupingComparator(Class)控制分组；同时需要自定义Partitioner，将每个group下的数据分到同一个reducer中，而每个group会形成一个part文件，在key很多的场景下可能会需要文件的拼接。（我的理解一般key对应的value比较少的场景直接在reducer中排序，如果key对应的value非常多，需要使用辅助排序。）

Reduce

In this phase the reduce(WritableComparable, Iterator, OutputCollector, Reporter) method is called for each <key, (list of values)> pair in the grouped inputs.

分享到：

[实验]hadoop例子 trackinfo数据清洗的改 ... | [实验]hadoop例子 trackinfo数据清洗

2013-09-03 10:15
浏览 747
评论(0)
分类:编程语言
查看更多

发表评论

您还没有登录,请您登录后再发表评论

最近访客更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论

[笔记]hadoop tutorial - Reducer

评论

发表评论

相关推荐

最近访客 更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论

[笔记]hadoop tutorial - Reducer

评论

发表评论

相关推荐

[实验]avro与non-avro的mapred例子-wordcount改写

[实验]hadoop例子 trackinfo数据清洗的改写

[实验]hadoop例子 trackinfo数据清洗

[环境] hadoop 开发环境maven管理

[笔记]avro 介绍及官网例子

[实验]hadoop例子 在线用户分析

[笔记]hadoop mapred InputFormat分析

[笔记]hdfs namenode FSNamesystem分析

[笔记]hdfs namenode FSImage分析1

[实验]集群hadoop配置

[实验]单机hadoop配置

[问题解决]hadoop eclipse plugin

最近访客更多访客>>

[实验]hadoop例子在线用户分析