Python newspaper库新闻爬取项目输入URL后运行报错求助
新闻爬取应用报错修复方案
已定位的错误点
- 文件操作未正确关闭:代码中写入
article_content.txt后仅写了f.close,未加括号执行关闭操作,导致内容未实际刷入磁盘,后续Shell脚本读取到空文件 - 方法调用遗漏括号:判断单词末尾是否为小写时,
islower是方法对象,需加()调用才会返回布尔值,原代码逻辑完全失效 - 数值类型错误:统计词频时将频次存为字符串,后续排序时会按字符串规则排序而非数值大小,排序结果错误
- 资源占用与权限问题:执行Shell脚本前先打开了
keyword_count.txt未关闭,导致Shell写入文件失败;脚本中包含sudo命令,无权限时会直接执行失败 - 除零风险:如果NLP返回的关键词列表为空,计算相似度时会触发除以零报错
- 冗余逻辑错误:
keywords_变量在循环中重复赋值,完全没有必要,直接取切片即可
修复步骤
1 修复主程序代码错误
(1)修复文件关闭逻辑
原代码:
f= open("article_content.txt","w+") f.write(article.text) f.close f = open("keyword_count.txt","w+") subprocess.call(['./bda_script.sh'])
修改为:
# 写入文章内容后正确关闭文件 with open("article_content.txt","w+", encoding="utf-8") as f: f.write(article.text) # 提前关闭所有占用keyword_count.txt的句柄再执行脚本 subprocess.call(['./bda_script.sh'])
(2)修复islower调用错误
原代码:
if line.split()[0][-1].islower == False:
修改为:
if not line.split()[0][-1].islower():
(3)修复词频存储类型错误
原代码:
word_count[word] = line.split()[1]
修改为:
word_count[word] = int(line.split()[1])
(4)修复相似度计算除零问题
原代码:
print ('-->SIMILARITY: ',((len(set(keywords_) & set(article.keywords))/len(article.keywords))*100)+20,'%')
修改为:
if len(article.keywords) > 0: similarity = ((len(set(keywords_) & set(article.keywords))/len(article.keywords))*100)+20 else: similarity = 0 print ('-->SIMILARITY: ', similarity,'%')
(5)简化keywords_赋值逻辑
原代码:
for r in sorted_keys: keywords_ = sorted_keys[:len(article.keywords)]
修改为:
keywords_ = sorted_keys[:min(len(sorted_keys), len(article.keywords))]
2 修复Shell脚本问题
原脚本包含sudo无必要,修改为:
#!/bin/sh cat article_content.txt | ./mapper.py | ./reducer.py > keyword_count.txt
额外注意事项
- 需确保
mapper.py、reducer.py和主程序在同一目录,且有可执行权限,可执行chmod +x mapper.py reducer.py bda_script.sh添加权限 - 依赖库缺失的话可执行
pip install newspaper3k pillow安装依赖
内容的提问来源于stack exchange,提问作者Monu Patil
相关产品推荐
相关产品推荐

