You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何正确爬取法语维基分类页下戏剧条目首句并生成DataFrame

法语维基分类页戏剧条目首句提取方案

实现思路

  • 放弃第三方wikipedia模块,直接用法语维基官方MediaWiki API请求数据,从根源避免特殊字符导致的重定向错误
  • 先抓取分类页下所有戏剧条目的标题,再通过API批量查询条目首句,效率比单页逐爬更高
  • 全程保留法语特殊字符的原生编码,不会出现转码错误导致的匹配失败问题

完整可运行代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 配置参数
CATEGORY_URL = "你持有的法语维基分类页地址"
WIKI_API_ENDPOINT = "https://fr.wikipedia.org/w/api.php"

# 第一步:抓取分类下所有戏剧条目标题
def get_category_pages(category_url):
    pages = []
    next_url = category_url
    while next_url:
        resp = requests.get(next_url)
        soup = BeautifulSoup(resp.text, "html.parser")
        # 提取条目链接
        for a in soup.select("div.mw-category-group li a"):
            title = a.get("title")
            # 跳过分类、帮助页等非条目内容
            if not title or ":" in title:
                continue
            pages.append({"title": title})
        # 处理分页,匹配法语维基的下一页按钮文本
        next_link = soup.select_one("a:contains('page suivante')")
        next_url = f"https://fr.wikipedia.org{next_link.get('href')}" if next_link else None
    return pages

# 第二步:批量调用API获取条目首句
def get_first_sentences(pages):
    titles = [p["title"] for p in pages]
    res = []
    # 维基API单次最多查询50个条目,拆分批量请求
    for i in range(0, len(titles), 50):
        batch_titles = titles[i:i+50]
        params = {
            "action": "query",
            "format": "json",
            "prop": "extracts",
            "exintro": 1,
            "explaintext": 1,
            "exsentences": 1,
            "titles": "|".join(batch_titles),
            "redirects": 1 # 自动处理条目内合法重定向
        }
        resp = requests.get(WIKI_API_ENDPOINT, params=params)
        data = resp.json()
        for page in data["query"]["pages"].values():
            if "extract" in page:
                res.append({"title": page["title"], "first_sentence": page["extract"].strip()})
            else:
                res.append({"title": page["title"], "first_sentence": ""})
    return res

# 第三步:输出为DataFrame
if __name__ == "__main__":
    pages = get_category_pages(CATEGORY_URL)
    pages_with_sentence = get_first_sentences(pages)
    df = pd.DataFrame(pages_with_sentence)
    # 查看结果
    print(df.head())
    # 导出为csv时用utf-8-sig编码,避免法语特殊字符乱码
    # df.to_csv("18世纪戏剧首句汇总.csv", index=False, encoding="utf-8-sig")

注意事项

  • 可以根据实际需要调整单次请求的批量大小,避免触发接口限流
  • 无摘要的空条目会默认保留空值,可自行过滤处理

内容的提问来源于stack exchange,提问作者jonas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 14:45:02