You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复无法返回单元格值的维基百科表格网页爬虫?

问题原因

  • BeautifulSoup的find_all方法参数使用错误:你当前的写法row.find_all(['th'], ['td'])会将第二个列表['td']识别为标签属性过滤条件,而非要匹配的标签类型,因此无法找到符合要求的单元格元素,最终返回空列表。
  • 若要同时匹配<th>和<td>两种标签,需要将两类标签名放在同一个列表中作为find_all的第一个参数。

修复方案

仅需要修改循环中读取单元格的那一行代码即可,修正后的完整代码如下:

import requests
from bs4 import BeautifulSoup
import re
import dateutil

result = requests.get('https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population')
assert result.status_code==200
print(result.status_code)

src = result.content
document = BeautifulSoup(src, 'lxml')

table = document.find('table')
assert table.find('th').get_text() == "Rank"

rows = table.find_all('tr')

for row in rows[1:-1]:
    # 修正find_all的参数写法
    cells = row.find_all(['th', 'td'])
    # 增加strip参数可清理单元格内多余的换行、空格字符
    cells_text = [cell.get_text(strip=True) for cell in cells]
    print(cells_text)

内容的提问来源于stack exchange,提问作者Henriksokar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 00:18:03