为何两段BeautifulSoup表格获取代码结果不同?一段返回None
问题:BeautifulSoup获取维基表格时两段代码的差异原因
我用两段代码尝试获取维基百科「美国州及领地列表」页面的表格,第一段返回None,第二段能正常拿到目标表格,想问问这个差异是不是和带连字符的类名有关?
第一段代码:
from bs4 import BeautifulSoup import requests url_link = "https://en.wikipedia.org/wiki/List_of_states_and_territories_of_the_United_States" source = requests.get(url_link).text soup = BeautifulSoup(source, "lxml") my_table = soup.find("table", class_ = "wikitable sortable plainrowheaders jquery-tablesorter")
第二段代码:
from bs4 import BeautifulSoup import requests url_link = "https://en.wikipedia.org/wiki/List_of_states_and_territories_of_the_United_States" source = requests.get(url_link).text soup = BeautifulSoup(source, "lxml") my_table = soup.find("table", class_ = "wikitable sortable plainrowheaders")
解答
这个差异和带连字符的类名完全无关,问题出在jquery-tablesorter这个类上:
- 用
requests爬取到的是页面的原始HTML代码,而jquery-tablesorter是页面加载后由前端JavaScript动态添加的类,原始HTML里根本没有这个类。 - 第一段代码指定了这个不存在的类,
soup.find()找不到匹配元素,所以返回None;第二段去掉了这个多余的类,只保留原始HTML中表格实际存在的wikitable sortable plainrowheaders类,因此能成功定位到目标表格。
内容的提问来源于stack exchange,提问作者Neha Rajput
相关产品推荐
相关产品推荐

