You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup抓取对应指定表头文本的<td>内容?

解决维基百科信息框奖项爬取的循环终止问题

你的代码目前有两个核心问题:一是每次循环都调用table.find('tr'),这只会返回表格里的第一个元素,导致无限循环;二是没有处理表格中不存在"Awards"表头的情况,很容易引发报错。

这里给你两种更可靠的实现方式:

方法一:直接定位目标行

不用循环遍历,直接通过匹配逻辑找到表头文本为"Awards"的行,一步到位:

import requests
from bs4 import BeautifulSoup

testurl = "https://en.wikipedia.org/wiki/Alan_Turing"
page = requests.get(testurl)
page_content = BeautifulSoup(page.content, "html.parser")
table = page_content.find('table', attrs={'class':'infobox biography vcard'})

# 直接找到<th>文本为"Awards"的<tr>行
target_tr = table.find('tr', lambda tag: tag.th and tag.th.get_text(strip=True) == 'Awards')

if target_tr:
    td = target_tr.find('td')
    # 提取格式化后的奖项文本(处理换行和列表结构)
    awards = [item.get_text(strip=True) for item in td.find_all('li')]
    print("艾伦·图灵的奖项:", awards)
else:
    print("该页面信息框中未找到'Awards'表头")

方法二:安全遍历所有行

如果需要遍历所有行做更多扩展操作,可以用find_all('tr')获取所有行,逐个检查后自然终止循环,不会出现无限循环的问题:

import requests
from bs4 import BeautifulSoup

testurl = "https://en.wikipedia.org/wiki/Alan_Turing"
page = requests.get(testurl)
page_content = BeautifulSoup(page.content, "html.parser")
table = page_content.find('table', attrs={'class':'infobox biography vcard'})

td = None
# 遍历所有<tr>行
for tr in table.find_all('tr'):
    th = tr.find('th')
    # 先确保<th>存在,再检查文本是否匹配
    if th and th.get_text(strip=True) == 'Awards':
        td = tr.find('td')
        break

if td:
    # 处理奖项内容,提取每个列表项的纯文本
    awards_list = [li.get_text(strip=True) for li in td.find_all('li')]
    print("\n".join(awards_list))
else:
    print("未找到奖项信息")

关键优化说明:

  • 用get_text(strip=True)替代renderContents():renderContents()会返回带HTML标签的原始内容,而get_text()能直接提取纯文本,strip=True还能自动去除多余的空格和换行。
  • 增加空值判断:检查<th>是否存在,避免因某些行没有表头标签引发AttributeError。
  • 处理未找到目标的场景:通过if td:判断结果,避免后续操作因空值报错。

内容的提问来源于stack exchange,提问作者Hamuel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:02:38