You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup解析HTML注释表格时带*文本显示None的解决方法

问题根因

带标记的球员姓名返回None,核心原因是这类姓名所在的<td>标签内存在多个文本子节点,直接调用.string属性仅能获取标签下唯一子节点的文本内容,多文本节点场景下会直接返回None,和字符本身无关。

可行实现方案

按照需求原地替换pitching_tbl内所有*字符、同步更新DOM结构后再解析数据,可直接参考以下修正后的代码:

import requests
import pandas as pd
from bs4 import BeautifulSoup, Comment

page = BeautifulSoup(requests.get('https://www.baseball-reference.com/register/team.cgi?id=b0a9f9bc').text, features='lxml')

tbls = []
for comment in page.find_all(text=lambda text: isinstance(text, Comment)):
    if comment.find("<table ") > 0:
        comment_soup = BeautifulSoup(comment, 'lxml')
        table = comment_soup.find("table")
        tbls.append(table)

pitching_tbl = tbls[0]

# 遍历表格内所有文本节点,原地移除*标记
for text_node in pitching_tbl.find_all(text=True):
    if '*' in text_node:
        text_node.replace_with(text_node.replace('*', ''))

def parse_row(row):
    # 用get_text替代string,兼容多文本节点场景,避免返回None
    return [cell.get_text(strip=True) for cell in row.find_all('td')]

rows = pitching_tbl.find_all('tr')
# 提取表头作为DataFrame列名
header = [th.get_text(strip=True) for th in rows[0].find_all('th')]
# 跳过表头行解析实际数据
data = pd.DataFrame([parse_row(row) for row in rows[1:]], columns=header)
关键逻辑说明
  • 调用pitching_tbl.find_all(text=True)遍历表格下所有文本节点,检测到*字符时直接用replace_with方法更新节点内容,该操作会直接修改pitching_tbl对应的DOM结构,满足原地更新HTML内容的要求
  • 原解析逻辑中的x.string替换为x.get_text(strip=True),无论标签下有多少个文本子节点都能正确拼接拿到完整文本,从根源避免返回None的问题
  • 补充了表头提取逻辑,原实现会将表头行误存入数据区、且DataFrame无有效列名,修正后数据结构更规范

内容的提问来源于stack exchange,提问作者Jensen Holm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 02:12:07