Scrapy爬取网球比分遇AttributeError及解析方案问询
网球比赛比分爬取问题
我想要爬取某网站表格中第三场网球比赛的比分,需要从指定HTML元素解析单场比赛的各盘比分,目标HTML片段如下:
html = ( ''' <td class="day-table-score"> <a data-override-transition="" data-ga-category="" data-ga-action="Click" data-ga-label="" data-use-ga="true" class="not-in-system"> <!-- Determine set tie break score --> 76 <sup>8</sup> <!-- Determine set tie break score --> 61 </a> </td> ''' )
我用以下代码模拟生成HtmlResponse对象:
from scrapy.http.response.html import HtmlResponse response = HtmlResponse("xyz", body=html, encoding="utf-8")
通过response.xpath("normalize-space()").get()可以获取比分原始文本,但为了识别<sup>标签中的抢七比分,我尝试遍历response.xpath("//text()")返回的Selector对象并访问其attrib属性,代码如下:
text_lines = response.xpath("//text()") for tl in text_lines: print(tl.attrib)
结果出现报错:
Exception has occurred: AttributeError 'str' object has no attribute 'attrib'
虽然type(tl)显示tl是Selector对象而非字符串,现在咨询两个问题:
- 该报错原因是什么?是否与parsel包有关?
- 这种遍历文本节点构建比分字典(如
{"set_1_score": 76, "set_1_tiebreak_score": 8, "set_2_score": 61, "set_2_tiebreak_score": None})的方式是否为最优方案?
问题解答
1. 报错原因解析
你调用tl.attrib时出错,是因为你遍历的text_lines里的Selector对应的是文本节点,而非元素节点。文本节点本身没有属性(attrib是元素节点的专属属性),当你尝试访问tl.attrib时,Scrapy/Parsel内部会自动将文本节点转为字符串处理,进而触发'str' object has no attribute 'attrib'的错误,这和parsel包的底层转换逻辑直接相关。
2. 比分解析方案优化
遍历文本节点的方式不是最优方案,推荐直接针对HTML元素结构进行解析,更简洁可靠:
- 先定位到每个包含单盘比分的节点组(每个盘的比分由主比分文本+可选
<sup>抢七分组成) - 对每个盘的节点组分别提取主比分和抢七分
示例代码:
# 定位所有包含比分的节点(文本节点+sup节点) score_nodes = response.xpath("//td[@class='day-table-score']/a/node()") current_set = {} result_list = [] for node in score_nodes: # 提取节点文本并过滤空内容和注释 text = node.xpath("normalize-space(text())").get() if text and text.isdigit(): # 遇到主比分,先把上一个盘的结果存入列表(如果存在) if current_set: result_list.append(current_set) # 初始化当前盘的字典 current_set = {"score": int(text), "tiebreak": None} elif node.xpath("name()='sup'"): # 遇到sup标签,把抢七分关联到当前盘 current_set["tiebreak"] = int(node.xpath("text()").get()) # 存入最后一个盘的结果 if current_set: result_list.append(current_set) # 转换为要求的字典格式 final_result = {} for idx, set_data in enumerate(result_list, start=1): final_result[f"set_{idx}_score"] = set_data["score"] final_result[f"set_{idx}_tiebreak_score"] = set_data["tiebreak"] print(final_result)
输出:
{'set_1_score': 76, 'set_1_tiebreak_score': 8, 'set_2_score': 61, 'set_2_tiebreak_score': None}
这种方式直接利用节点结构关系解析,避免了遍历所有文本节点(包括冗余注释)的无效操作,可读性和可维护性更强,也更贴合HTML的结构逻辑。
内容的提问来源于stack exchange,提问作者Jossy
相关产品推荐
相关产品推荐

