You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何两段BeautifulSoup表格获取代码结果不同?一段返回None

问题:BeautifulSoup获取维基表格时两段代码的差异原因

我用两段代码尝试获取维基百科「美国州及领地列表」页面的表格,第一段返回None,第二段能正常拿到目标表格,想问问这个差异是不是和带连字符的类名有关?

第一段代码:

from bs4 import BeautifulSoup
import requests
url_link = "https://en.wikipedia.org/wiki/List_of_states_and_territories_of_the_United_States"
source = requests.get(url_link).text
soup = BeautifulSoup(source, "lxml")
my_table = soup.find("table", class_ = "wikitable sortable plainrowheaders jquery-tablesorter")

第二段代码:

from bs4 import BeautifulSoup
import requests
url_link = "https://en.wikipedia.org/wiki/List_of_states_and_territories_of_the_United_States"
source = requests.get(url_link).text
soup = BeautifulSoup(source, "lxml")
my_table = soup.find("table", class_ = "wikitable sortable plainrowheaders")

解答

这个差异和带连字符的类名完全无关,问题出在jquery-tablesorter这个类上:

  • 用requests爬取到的是页面的原始HTML代码,而jquery-tablesorter是页面加载后由前端JavaScript动态添加的类,原始HTML里根本没有这个类。
  • 第一段代码指定了这个不存在的类,soup.find()找不到匹配元素,所以返回None;第二段去掉了这个多余的类,只保留原始HTML中表格实际存在的wikitable sortable plainrowheaders类,因此能成功定位到目标表格。

内容的提问来源于stack exchange,提问作者Neha Rajput

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:03:20