BeautifulSoup解析HTML注释表格时带*文本显示None的解决方法
问题根因
带标记的球员姓名返回None,核心原因是这类姓名所在的<td>标签内存在多个文本子节点,直接调用.string属性仅能获取标签下唯一子节点的文本内容,多文本节点场景下会直接返回None,和字符本身无关。
可行实现方案
按照需求原地替换pitching_tbl内所有*字符、同步更新DOM结构后再解析数据,可直接参考以下修正后的代码:
import requests import pandas as pd from bs4 import BeautifulSoup, Comment page = BeautifulSoup(requests.get('https://www.baseball-reference.com/register/team.cgi?id=b0a9f9bc').text, features='lxml') tbls = [] for comment in page.find_all(text=lambda text: isinstance(text, Comment)): if comment.find("<table ") > 0: comment_soup = BeautifulSoup(comment, 'lxml') table = comment_soup.find("table") tbls.append(table) pitching_tbl = tbls[0] # 遍历表格内所有文本节点,原地移除*标记 for text_node in pitching_tbl.find_all(text=True): if '*' in text_node: text_node.replace_with(text_node.replace('*', '')) def parse_row(row): # 用get_text替代string,兼容多文本节点场景,避免返回None return [cell.get_text(strip=True) for cell in row.find_all('td')] rows = pitching_tbl.find_all('tr') # 提取表头作为DataFrame列名 header = [th.get_text(strip=True) for th in rows[0].find_all('th')] # 跳过表头行解析实际数据 data = pd.DataFrame([parse_row(row) for row in rows[1:]], columns=header)
关键逻辑说明
- 调用
pitching_tbl.find_all(text=True)遍历表格下所有文本节点,检测到*字符时直接用replace_with方法更新节点内容,该操作会直接修改pitching_tbl对应的DOM结构,满足原地更新HTML内容的要求 - 原解析逻辑中的
x.string替换为x.get_text(strip=True),无论标签下有多少个文本子节点都能正确拼接拿到完整文本,从根源避免返回None的问题 - 补充了表头提取逻辑,原实现会将表头行误存入数据区、且DataFrame无有效列名,修正后数据结构更规范
内容的提问来源于stack exchange,提问作者Jensen Holm
相关产品推荐
相关产品推荐

