如何修复无法返回单元格值的维基百科表格网页爬虫?
问题原因
BeautifulSoup的find_all方法参数使用错误:你当前的写法row.find_all(['th'], ['td'])会将第二个列表['td']识别为标签属性过滤条件,而非要匹配的标签类型,因此无法找到符合要求的单元格元素,最终返回空列表。- 若要同时匹配
<th>和<td>两种标签,需要将两类标签名放在同一个列表中作为find_all的第一个参数。
修复方案
仅需要修改循环中读取单元格的那一行代码即可,修正后的完整代码如下:
import requests from bs4 import BeautifulSoup import re import dateutil result = requests.get('https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population') assert result.status_code==200 print(result.status_code) src = result.content document = BeautifulSoup(src, 'lxml') table = document.find('table') assert table.find('th').get_text() == "Rank" rows = table.find_all('tr') for row in rows[1:-1]: # 修正find_all的参数写法 cells = row.find_all(['th', 'td']) # 增加strip参数可清理单元格内多余的换行、空格字符 cells_text = [cell.get_text(strip=True) for cell in cells] print(cells_text)
内容的提问来源于stack exchange,提问作者Henriksokar
相关产品推荐
相关产品推荐

