You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取网球比分遇AttributeError及解析方案问询

网球比赛比分爬取问题

我想要爬取某网站表格中第三场网球比赛的比分,需要从指定HTML元素解析单场比赛的各盘比分,目标HTML片段如下:

html = (
    '''
    <td class="day-table-score">
        <a data-override-transition="" data-ga-category="" data-ga-action="Click" data-ga-label="" data-use-ga="true" class="not-in-system">
            <!-- Determine set tie break score --> 76 <sup>8</sup>
            <!-- Determine set tie break score --> 61
        </a>
    </td>
    '''
)

我用以下代码模拟生成HtmlResponse对象:

from scrapy.http.response.html import HtmlResponse

response = HtmlResponse("xyz", body=html, encoding="utf-8")

通过response.xpath("normalize-space()").get()可以获取比分原始文本,但为了识别<sup>标签中的抢七比分,我尝试遍历response.xpath("//text()")返回的Selector对象并访问其attrib属性,代码如下:

text_lines = response.xpath("//text()")
for tl in text_lines:
    print(tl.attrib)

结果出现报错:

Exception has occurred: AttributeError
'str' object has no attribute 'attrib'

虽然type(tl)显示tl是Selector对象而非字符串,现在咨询两个问题:

  1. 该报错原因是什么?是否与parsel包有关?
  2. 这种遍历文本节点构建比分字典(如{"set_1_score": 76, "set_1_tiebreak_score": 8, "set_2_score": 61, "set_2_tiebreak_score": None})的方式是否为最优方案?

问题解答

1. 报错原因解析

你调用tl.attrib时出错,是因为你遍历的text_lines里的Selector对应的是文本节点,而非元素节点。文本节点本身没有属性(attrib是元素节点的专属属性),当你尝试访问tl.attrib时,Scrapy/Parsel内部会自动将文本节点转为字符串处理,进而触发'str' object has no attribute 'attrib'的错误,这和parsel包的底层转换逻辑直接相关。

2. 比分解析方案优化

遍历文本节点的方式不是最优方案,推荐直接针对HTML元素结构进行解析,更简洁可靠:

  • 先定位到每个包含单盘比分的节点组(每个盘的比分由主比分文本+可选<sup>抢七分组成)
  • 对每个盘的节点组分别提取主比分和抢七分

示例代码:

# 定位所有包含比分的节点(文本节点+sup节点)
score_nodes = response.xpath("//td[@class='day-table-score']/a/node()")
current_set = {}
result_list = []

for node in score_nodes:
    # 提取节点文本并过滤空内容和注释
    text = node.xpath("normalize-space(text())").get()
    if text and text.isdigit():
        # 遇到主比分,先把上一个盘的结果存入列表(如果存在)
        if current_set:
            result_list.append(current_set)
        # 初始化当前盘的字典
        current_set = {"score": int(text), "tiebreak": None}
    elif node.xpath("name()='sup'"):
        # 遇到sup标签,把抢七分关联到当前盘
        current_set["tiebreak"] = int(node.xpath("text()").get())

# 存入最后一个盘的结果
if current_set:
    result_list.append(current_set)

# 转换为要求的字典格式
final_result = {}
for idx, set_data in enumerate(result_list, start=1):
    final_result[f"set_{idx}_score"] = set_data["score"]
    final_result[f"set_{idx}_tiebreak_score"] = set_data["tiebreak"]

print(final_result)

输出:

{'set_1_score': 76, 'set_1_tiebreak_score': 8, 'set_2_score': 61, 'set_2_tiebreak_score': None}

这种方式直接利用节点结构关系解析,避免了遍历所有文本节点(包括冗余注释)的无效操作,可读性和可维护性更强,也更贴合HTML的结构逻辑。


内容的提问来源于stack exchange,提问作者Jossy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 13:55:20