You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取维基百科Quartz属性表返回None的问题求助

搞定维基百科Quartz页面Infobox爬取问题+优化方案

为啥你的代码返回None?

最核心的问题是目标表格的类名已经更新了——维基百科的页面结构会不定期调整,你用的infobox ww_ib ww_ib_1是旧版的类名,现在矿物类的infobox通常用infobox或者infobox vcard这类更通用的类名。另外,初始化BeautifulSoup时最好指定解析器(比如html.parser),避免一些隐性的解析bug。

修正后的可用代码

下面是调整后的代码,能准确定位到目标表格,并提取你要的分类、氧化物、化学式信息:

from bs4 import BeautifulSoup
import requests

wiki_url = "http://en.wikipedia.org/wiki/Quartz"
# 用更标准的User-Agent,避免被维基反爬拦截
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}

# 发送请求,先确保请求成功
response = requests.get(wiki_url, headers=headers)
response.raise_for_status()  # 如果请求失败,直接抛出错误提示

# 初始化BeautifulSoup,指定解析器
soup = BeautifulSoup(response.content, 'html.parser')

# 用更通用的类名定位infobox表格
table = soup.find("table", class_="infobox")

if table:
    mineral_details = {}
    # 遍历表格每一行,提取目标字段
    for row in table.find_all('tr'):
        label_cell = row.find('th')
        value_cell = row.find('td')
        if label_cell and value_cell:
            label = label_cell.get_text(strip=True)
            value = value_cell.get_text(strip=True)
            # 匹配你需要的属性
            if label == "Category":
                mineral_details["分类"] = value
            elif label == "Oxide":
                mineral_details["氧化物"] = value
            elif label == "Formula":
                mineral_details["化学式"] = value
    
    print("提取到的矿物属性:")
    for key, val in mineral_details.items():
        print(f"{key}: {val}")
else:
    print("没找到目标表格哦,可能维基又更新页面结构了,得再检查下类名~")

爬维基百科的最优姿势:用官方API!

直接解析HTML最大的痛点就是页面结构不稳定,哪天维基改个类名或布局,你的代码就崩了。更靠谱的方式是用维基官方的MediaWiki API,直接拿结构化数据,不用折腾HTML解析。

给你个简单的示例,用API获取Quartz的属性:

import requests

api_endpoint = "https://en.wikipedia.org/w/api.php"
# 配置API参数,指定要获取的页面和数据类型
params = {
    "action": "query",
    "format": "json",
    "titles": "Quartz",
    "prop": "pageprops|revisions",
    "rvprop": "content",
    "rvslots": "main"
}

response = requests.get(api_endpoint, params=params)
data = response.json()

# 提取页面内容(这里可以用专门的库解析维基的wikitext格式,比如mwparserfromhell)
page_info = next(iter(data['query']['pages'].values()))
print(page_info['revisions'][0]['slots']['main']['*'])

如果想更省心,还可以用专门的Python库,比如wikipedia-api或者mwclient,这些库把API调用封装得很友好,直接调用方法就能拿数据,不用自己拼参数。

内容的提问来源于stack exchange,提问作者vineeth venugopal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:07:37