You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取问题:如何提取<td>标签内的目标文本(排除子元素)

提取包含子元素的标签内的目标文本

当标签包含多个子节点(文本+其他标签)时,.string会返回None——因为这个方法仅在标签有唯一子节点时才返回内容。要提取945.000,可以用以下几种方法(以BeautifulSoup为例):

  • 方法1:通过.stripped_strings获取首个文本片段
    .stripped_strings会遍历标签下所有非空白文本片段,直接取第一个即可:

    from bs4 import BeautifulSoup
    
    html = '<td>945.000 (<font color="blue">+3.000</font>)</td>'
    soup = BeautifulSoup(html, 'html.parser')
    td_tag = soup.find('td')
    target_text = next(td_tag.stripped_strings)
    print(target_text)  # 输出:945.000
    
  • 方法2:通过.contents提取首个文本节点并处理
    .contents返回标签的直接子节点列表,第一个就是包含目标的文本节点,再拆分去掉多余字符:

    target_text = td_tag.contents[0].strip().split('(')[0].strip()
    print(target_text)  # 输出:945.000
    
  • 方法3:合并所有文本后拆分
    先用.get_text()合并标签下所有文本,再按分隔符拆分提取目标:

    full_text = td_tag.get_text(strip=True)
    target_text = full_text.split('(')[0]
    print(target_text)  # 输出:945.000
    

内容的提问来源于stack exchange,提问作者Willl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 19:10:31