Windows11下Hadoop MapReduce Reduce阶段停滞(0%)求助
问题
在Windows 11系统中使用Python开发Hadoop MapReduce程序,Mapper函数已执行完成(Map 100%),但Reducer函数始终未运行,一直停留在Reduce 0%状态,等待许久无变化。
Mapper代码
import sys for line in sys.stdin: line = line.strip() words = line.split(", ") try: print('%s|%s' % (words[0].strip(), words[1])) except: pass
Reducer代码
from operator import itemgetter import sys current_word = None current_good = "" word = None for line in sys.stdin: line = line.strip() word, good = line.split('|') if current_word == word: current_good = current_good + ", " + good else: if current_word: print('%s|%s' % (current_word, current_good)) current_good = good current_word = word if current_word == word: print('%s|%s' % (current_word, current_good))
输入TXT数据
ani,pensil budi,pensil ani,buku dodi,penggaris
Reduce阶段停滞在0%,任务无任何进展。
解决步骤
修复键值对分隔符问题
Hadoop MapReduce默认要求Mapper输出的键值对用**制表符(\t)**分隔,而非自定义的|。如果使用自定义分隔符,Hadoop会将整行内容当作键,值为空,导致Reducer无法接收有效输入,进而停滞。
修改Mapper输出为制表符分隔:print('%s\t%s' % (words[0].strip(), words[1]))同步修改Reducer的拆分逻辑:
word, good = line.split('\t')验证Shuffle阶段状态
Map完成后,Hadoop需要对输出数据做排序、洗牌才能分发到Reducer。如果测试数据量极小,Hadoop可能延迟启动Reducer;也可能是缓存或节点通信问题导致Shuffle卡住。可以:- 增加输入数据量,测试是否触发Reducer运行;
- 查看Hadoop任务日志(
$HADOOP_HOME/logs目录下的任务跟踪日志),排查Shuffle阶段的报错信息。
确保脚本权限与环境兼容
- 在Mapper和Reducer脚本开头添加Shebang行,指定Python环境:
#!/usr/bin/env python3; - 本地执行
chmod +x mapper.py reducer.py赋予脚本可执行权限,再上传到Hadoop集群运行; - 确认集群所有节点安装了对应版本的Python。
- 在Mapper和Reducer脚本开头添加Shebang行,指定Python环境:
检查集群资源配置
若集群内存、CPU资源不足,Reducer可能无法分配到容器启动。通过ResourceManager页面(默认端口8088)查看集群资源使用情况,确认是否有足够容器分配给Reducer任务。本地测试Reducer逻辑
在本地模拟Hadoop的数据流,验证Reducer逻辑是否正常:cat input.txt | python mapper.py | sort | python reducer.py如果本地运行正常,说明问题出在Hadoop集群配置或运行流程;如果本地也卡住,检查Reducer代码是否有未处理的异常(比如拆分失败)或逻辑漏洞。
内容的提问来源于stack exchange,提问作者Aurell Layalia
相关产品推荐
相关产品推荐

