如何用Python批量抓取维基百科信息框并优化结果展示
问题解决方案
1. 从维基列表页批量提取子页面URL
直接用BeautifulSoup解析列表页HTML,筛选出符合规则的子页面链接,核心逻辑是定位列表中的条目链接,过滤掉非内容页面(如讨论页、文件页等):
import requests from bs4 import BeautifulSoup import pandas as pd def extract_wiki_subpages(list_page_url): # 请求列表页 response = requests.get(list_page_url) soup = BeautifulSoup(response.text, 'html.parser') # 定位列表条目(维基列表通常用<ul>或<table>承载,可根据实际页面调整选择器) subpage_links = set() # 用集合去重 for a_tag in soup.find_all('a', href=True): href = a_tag['href'] # 筛选内容页链接:以/wiki/开头,排除特殊命名空间页面 if href.startswith('/wiki/') and not any(namespace in href for namespace in ['File:', 'Talk:', 'User:', 'Category:', 'Template:']): full_url = f"https://en.wikipedia.org{href}" # 根据实际维基域名调整 subpage_links.add(full_url) return list(subpage_links) # 调用示例 wiki_list_url = "https://en.wikipedia.org/wiki/List_of_programming_languages" # 替换为目标列表页 subpage_urls = extract_wiki_subpages(wiki_list_url) print(f"提取到{len(subpage_urls)}个子页面URL")
注意:不同维基列表页的结构可能不同,若默认选择器不生效,可通过浏览器开发者工具定位条目对应的HTML标签(比如某些列表用class='mw-category-group'的容器),调整soup.find_all的参数即可。
2. 优化DataFrame展示格式
针对抓取的infobox数据,从列名规范、缺失值处理、可视化展示三个维度优化:
2.1 规范数据结构
抓取infobox时,统一列名格式(比如转为小写、替换空格为下划线),确保每条数据的字段对齐:
def parse_infobox(page_url): response = requests.get(page_url) soup = BeautifulSoup(response.text, 'html.parser') infobox = soup.find('table', class_='infobox') if not infobox: return None infobox_data = {} for row in infobox.find_all('tr'): header = row.find('th') value = row.find('td') if header and value: # 规范列名:去除特殊字符、转为小写、空格换下划线 clean_key = header.get_text(strip=True).lower().replace(' ', '_') clean_value = value.get_text(strip=True) infobox_data[clean_key] = clean_value return infobox_data # 批量抓取并构建DataFrame infobox_list = [] for url in subpage_urls[:5]: # 先测试前5条 data = parse_infobox(url) if data: infobox_list.append(data) df = pd.DataFrame(infobox_list)
2.2 优化展示效果
通过pandas的显示设置和样式工具,输出标准表格:
# 全局设置:避免列截断、换行 pd.set_option('display.max_columns', None) pd.set_option('display.max_colwidth', 100) pd.set_option('display.width', 1000) # 输出Markdown格式的标准表格(适合文档或论坛展示) print(df.to_markdown(index=False)) # 若在Jupyter环境中,可美化表格样式 styled_df = df.style \ .set_table_styles([{'selector': 'thead th', 'props': [('background-color', '#f0f0f0'), ('font-weight', 'bold')]}]) \ .highlight_null(null_color='#ffcccc') display(styled_df)
2.3 处理缺失值
统一填充缺失字段,避免表格错位:
# 填充所有缺失值为指定内容(如"N/A") df_filled = df.fillna("N/A")
内容的提问来源于stack exchange,提问作者zero
相关产品推荐
相关产品推荐

