如何用BeautifulSoup抓取对应指定表头文本的<td>内容?
解决维基百科信息框奖项爬取的循环终止问题
你的代码目前有两个核心问题:一是每次循环都调用table.find('tr'),这只会返回表格里的第一个元素,导致无限循环;二是没有处理表格中不存在"Awards"表头的情况,很容易引发报错。
这里给你两种更可靠的实现方式:
方法一:直接定位目标行
不用循环遍历,直接通过匹配逻辑找到表头文本为"Awards"的行,一步到位:
import requests from bs4 import BeautifulSoup testurl = "https://en.wikipedia.org/wiki/Alan_Turing" page = requests.get(testurl) page_content = BeautifulSoup(page.content, "html.parser") table = page_content.find('table', attrs={'class':'infobox biography vcard'}) # 直接找到<th>文本为"Awards"的<tr>行 target_tr = table.find('tr', lambda tag: tag.th and tag.th.get_text(strip=True) == 'Awards') if target_tr: td = target_tr.find('td') # 提取格式化后的奖项文本(处理换行和列表结构) awards = [item.get_text(strip=True) for item in td.find_all('li')] print("艾伦·图灵的奖项:", awards) else: print("该页面信息框中未找到'Awards'表头")
方法二:安全遍历所有行
如果需要遍历所有行做更多扩展操作,可以用find_all('tr')获取所有行,逐个检查后自然终止循环,不会出现无限循环的问题:
import requests from bs4 import BeautifulSoup testurl = "https://en.wikipedia.org/wiki/Alan_Turing" page = requests.get(testurl) page_content = BeautifulSoup(page.content, "html.parser") table = page_content.find('table', attrs={'class':'infobox biography vcard'}) td = None # 遍历所有<tr>行 for tr in table.find_all('tr'): th = tr.find('th') # 先确保<th>存在,再检查文本是否匹配 if th and th.get_text(strip=True) == 'Awards': td = tr.find('td') break if td: # 处理奖项内容,提取每个列表项的纯文本 awards_list = [li.get_text(strip=True) for li in td.find_all('li')] print("\n".join(awards_list)) else: print("未找到奖项信息")
关键优化说明:
- 用
get_text(strip=True)替代renderContents():renderContents()会返回带HTML标签的原始内容,而get_text()能直接提取纯文本,strip=True还能自动去除多余的空格和换行。 - 增加空值判断:检查
<th>是否存在,避免因某些行没有表头标签引发AttributeError。 - 处理未找到目标的场景:通过
if td:判断结果,避免后续操作因空值报错。
内容的提问来源于stack exchange,提问作者Hamuel
相关产品推荐
相关产品推荐

