You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python批量抓取维基百科信息框并优化结果展示

问题解决方案

1. 从维基列表页批量提取子页面URL

直接用BeautifulSoup解析列表页HTML,筛选出符合规则的子页面链接,核心逻辑是定位列表中的条目链接,过滤掉非内容页面(如讨论页、文件页等):

import requests
from bs4 import BeautifulSoup
import pandas as pd

def extract_wiki_subpages(list_page_url):
    # 请求列表页
    response = requests.get(list_page_url)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 定位列表条目(维基列表通常用<ul>或<table>承载,可根据实际页面调整选择器)
    subpage_links = set()  # 用集合去重
    for a_tag in soup.find_all('a', href=True):
        href = a_tag['href']
        # 筛选内容页链接:以/wiki/开头,排除特殊命名空间页面
        if href.startswith('/wiki/') and not any(namespace in href for namespace in ['File:', 'Talk:', 'User:', 'Category:', 'Template:']):
            full_url = f"https://en.wikipedia.org{href}"  # 根据实际维基域名调整
            subpage_links.add(full_url)
    
    return list(subpage_links)

# 调用示例
wiki_list_url = "https://en.wikipedia.org/wiki/List_of_programming_languages"  # 替换为目标列表页
subpage_urls = extract_wiki_subpages(wiki_list_url)
print(f"提取到{len(subpage_urls)}个子页面URL")

注意:不同维基列表页的结构可能不同,若默认选择器不生效,可通过浏览器开发者工具定位条目对应的HTML标签(比如某些列表用class='mw-category-group'的容器),调整soup.find_all的参数即可。

2. 优化DataFrame展示格式

针对抓取的infobox数据,从列名规范、缺失值处理、可视化展示三个维度优化:

2.1 规范数据结构

抓取infobox时,统一列名格式(比如转为小写、替换空格为下划线),确保每条数据的字段对齐:

def parse_infobox(page_url):
    response = requests.get(page_url)
    soup = BeautifulSoup(response.text, 'html.parser')
    infobox = soup.find('table', class_='infobox')
    if not infobox:
        return None
    
    infobox_data = {}
    for row in infobox.find_all('tr'):
        header = row.find('th')
        value = row.find('td')
        if header and value:
            # 规范列名:去除特殊字符、转为小写、空格换下划线
            clean_key = header.get_text(strip=True).lower().replace(' ', '_')
            clean_value = value.get_text(strip=True)
            infobox_data[clean_key] = clean_value
    return infobox_data

# 批量抓取并构建DataFrame
infobox_list = []
for url in subpage_urls[:5]:  # 先测试前5条
    data = parse_infobox(url)
    if data:
        infobox_list.append(data)

df = pd.DataFrame(infobox_list)

2.2 优化展示效果

通过pandas的显示设置和样式工具,输出标准表格:

# 全局设置:避免列截断、换行
pd.set_option('display.max_columns', None)
pd.set_option('display.max_colwidth', 100)
pd.set_option('display.width', 1000)

# 输出Markdown格式的标准表格(适合文档或论坛展示)
print(df.to_markdown(index=False))

# 若在Jupyter环境中,可美化表格样式
styled_df = df.style \
    .set_table_styles([{'selector': 'thead th', 'props': [('background-color', '#f0f0f0'), ('font-weight', 'bold')]}]) \
    .highlight_null(null_color='#ffcccc')
display(styled_df)

2.3 处理缺失值

统一填充缺失字段,避免表格错位:

# 填充所有缺失值为指定内容(如"N/A")
df_filled = df.fillna("N/A")

内容的提问来源于stack exchange,提问作者zero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 22:07:41