TripAdvisor爬虫报错:'NoneType' object has no attribute 'decompose'如何解决?
解决TripAdvisor爬虫脚本中的
NoneType对象无decompose属性错误 Hey,这个错误很好解决,本质是你代码里的div.select_one('script')没有找到对应的<script>标签,返回了None,然后你直接调用decompose()就触发了'NoneType' object has no attribute 'decompose'这个异常。
问题原因
出现这种情况大概率是两种可能:
- 目标TripAdvisor页面的结构已经更新了,原来的选择器对应的位置不再有
<script>标签 - 你的选择器写得不够精准,没匹配到任何元素
修正后的完整代码
我先帮你把代码里的基础语法错误(比如import语句连写)也一并修正,再加上空值检查逻辑:
from bs4 import BeautifulSoup import time import urllib.request import re import csv def crawlcontents(): url = 'https://www.tripadvisor.com/ShowTopic-g983296-i13236-k11538516-Rent_from_LOTTE_standard_or_mystery_option-Jeju_Island.html' html = urllib.request.urlopen(url).read().decode() soup = BeautifulSoup(html,'html.parser') # 先检查是否能定位到目标内容容器 target_div = soup.select_one('#SHOW_TOPIC > div.balance > div.firstPostBox > div > div > div.postRightContent > div.postcontent > div.postBody') if not target_div: print("警告:无法找到目标内容容器,请检查页面结构或选择器") return # 查找script标签,只有找到时才执行删除操作 script_element = target_div.select_one('script') if script_element: script_element.decompose() # 清理文本内容 post_content = target_div.text.strip() post_content = re.sub(r'\n+', ' ', post_content) print(post_content) crawlcontents()
关键修改点
- 增加空值检查:对目标容器和script标签都做了存在性判断,避免对
None调用方法 - 修正语法问题:把连写的import语句拆分,修复了无效的正则表达式(原代码里的
r' +'是错误的,改成了r'\n+') - 变量命名更清晰:把原来的
div改成target_div,script_tag改成script_element,可读性更好
额外提示
如果运行后还是提示找不到目标容器,那说明TripAdvisor的页面结构已经发生变化了,你需要打开浏览器的开发者工具(F12),重新定位目标内容的CSS选择器,替换代码里的选择器字符串。
内容的提问来源于stack exchange,提问作者정세민
相关产品推荐
相关产品推荐

