使用BeautifulSoup抓取VLR.gg同class元素遇报错,求解决方案
VLR.gg电竞数据抓取问题排查与解决
问题背景
尝试用BeautifulSoup抓取vlr.gg的比赛数据,目标从以下HTML结构中提取剩余时间(如18h 24m)和比赛格式(如Bo3):
<div class="match-header-vs-score"> <div class="match-header-vs-note"> <span class="match-header-vs-note mod-upcoming">18h 24m </span> </div> <div class="match-header-vs-placeholder"> – </div> <div class="match-header-vs-note">Bo3 </div> </div>
用户原始代码:
box5=soup.find_all("div",class_="match-header-vs-score") for p in box5: matchtdetails=p.find("div",class_="match-header-vs-note").get_text(strip=True) print(" In: ",matchtdetails) box6=soup.find_all("div",class_="match-header-vs-note") for q in box6: if q.find("div",class_="match-header-vs-note"): matchdetails1=q.find("div",class_="match-header-vs-note").get_text(strip=True) print(matchdetails1)
遇到的核心问题:
- box6无任何返回结果
- 调用
find_all("match-header-vs-note")时频繁触发NoneType Attribute Error - 循环内使用
find_all()存列表时,报错AttributeError: 'NavigableString' object has no attribute 'find_all'
问题原因与解决方法
1. box6无返回结果的原因
box6的逻辑是遍历每个match-header-vs-note元素,然后尝试在其内部查找子级的同class元素,但目标HTML中match-header-vs-note的子节点是<span>,没有嵌套的同class div,因此q.find(...)永远返回None,自然无输出。
2. NoneType Attribute Error的原因
当p.find("div",class_="match-header-vs-note")找不到目标元素时,会返回None,此时直接调用.get_text()就会触发错误。必须先判断元素是否存在,再提取文本。
3. NavigableString错误的原因
如果遍历过程中不小心接触到文本节点(比如标签之间的空白字符),这些是NavigableString对象,没有find_all方法。需确保只遍历Tag类型的元素。
正确实现代码
方案1:按索引提取目标元素
利用容器内match-header-vs-note的顺序,直接按索引获取剩余时间和比赛格式:
from bs4 import BeautifulSoup # 假设soup已完成HTML解析 box5 = soup.find_all("div", class_="match-header-vs-score") for container in box5: # 获取当前容器下所有match-header-vs-note元素 note_elements = container.find_all("div", class_="match-header-vs-note") # 确保至少有2个目标元素再处理 if len(note_elements) >= 2: # 第一个元素:提取span内的剩余时间文本 time_left = note_elements[0].find("span", class_="match-header-vs-note mod-upcoming").get_text(strip=True) # 第二个元素:直接提取比赛格式文本 match_format = note_elements[1].get_text(strip=True) print(f" In: {time_left}") print(f" Format: {match_format}")
方案2:通过相邻节点定位
利用中间的match-header-vs-placeholder节点,定位其后续的比赛格式元素:
from bs4 import BeautifulSoup box5 = soup.find_all("div", class_="match-header-vs-score") for container in box5: # 提取剩余时间 time_note = container.find("div", class_="match-header-vs-note") if time_note: time_span = time_note.find("span") if time_span: time_left = time_span.get_text(strip=True) print(f" In: {time_left}") # 提取比赛格式:找到占位符后,取其下一个兄弟节点 placeholder = container.find("div", class_="match-header-vs-placeholder") if placeholder: format_note = placeholder.find_next_sibling("div", class_="match-header-vs-note") if format_note: match_format = format_note.get_text(strip=True) print(f" Format: {match_format}")
关键注意事项
- 避免链式调用风险:每次调用
find()后,先判断返回值是否为None,再执行后续操作,防止触发NoneType错误。 - 明确元素层级:不要在
match-header-vs-note内部查找同class子元素,目标HTML中不存在这种嵌套结构。 - 过滤非Tag元素:如果需要遍历节点的子元素(如
.contents或.children),通过isinstance(item, bs4.element.Tag)过滤掉文本节点。
内容的提问来源于stack exchange,提问作者Priyanshu Lahiri
相关产品推荐
相关产品推荐

