Python使用BeautifulSoup提取TripAdvisor网页子节点内容问题求助
修正方案
节点定位逻辑说明
你需要的父容器div.eeqnt下第7个section(直接子节点)是评论列表所在区域,BeautifulSoup定位第N个直接子节点有两种方式:
- 调用
find_all获取所有直接子section后按索引取值,注意索引从0开始计数,第7个对应下标为6 - 直接用CSS选择器
div.eeqnt > section:nth-child(7)匹配,语法更简洁
可运行的修正代码
from bs4 import BeautifulSoup as bs import inspect # 你的其他初始化、driver加载页面代码保持不变 source = driver.page_source data = bs(source, 'html.parser') tempData = [] # 定位评论所在的第7个section # 方式1:按索引取 parent_div = data.find('div', class_='eeqnt') # recursive=False 只匹配直接子节点,避免选中section内嵌的其他section target_section = parent_div.find_all('section', recursive=False)[6] # 方式2:CSS选择器(可替换上面两行) # target_section = data.select_one('div.eeqnt > section:nth-child(7)') # 提取所有评论内容 if logging: print(f"line {inspect.getframeinfo(inspect.currentframe()).lineno}: beginScrape(): Get all data in comment...") comment_list = target_section.find_all('span', class_='QewHA H4 _a') for comment_span in comment_list: currData = comment_span.get_text(strip=True) tempData.append(currData)
后续维护提示
TripAdvisor的动态类名会随站点更新变化,若后续抓取失效可以重新审查评论元素,替换对应类名即可。
内容的提问来源于stack exchange,提问作者shin1234
相关产品推荐
相关产品推荐

