You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Wikipedia指定类别中随机获取带图主题摘要并打印?

问题描述

需求:从Wikipedia的指定类别中随机选取一个主题,打印该主题的摘要及对应配图。

已实现指定主题的获取代码如下:

import requests


def get_wikipedia_summary_and_image(topic):
    # Make a request to the Wikipedia API to get the summary and image of the given topic
    response = requests.get(f"https://en.wikipedia.org/api/rest_v1/page/summary/{topic}")

    # Check if the request was successful
    if response.status_code == 200:
        # Extract the summary and image from the response
        summary = response.json()['extract']
        image = response.json()['thumbnail']['source']

        # Print the summary and image
        print(summary)
        print(image)
    else:
        print("An error occurred while fetching the summary and image.")


# Example usage
get_wikipedia_summary_and_image("Double-slit experiment")

目前已能实现指定主题的摘要与配图获取,但不清楚如何实现指定类别下的主题随机化,寻求技术解决方法。

解决方案

要实现指定类别下随机选主题,需借助Wikipedia的MediaWiki API获取类别下的页面列表,再随机挑选主题,最后复用已有代码获取摘要和配图。具体实现如下:

步骤说明

  • 获取类别页面列表:调用MediaWiki API的query接口,通过categorymembers参数拉取指定类别下的页面标题。
  • 随机挑选主题:用Python的random模块从页面列表中随机选一个标题。
  • 兼容异常场景:处理无配图、类别无页面等异常情况,避免程序报错。

完整代码

import requests
import random


def get_wikipedia_summary_and_image(topic):
    response = requests.get(f"https://en.wikipedia.org/api/rest_v1/page/summary/{topic}")
    if response.status_code == 200:
        data = response.json()
        summary = data['extract']
        # 处理页面无缩略图的情况
        image = data.get('thumbnail', {}).get('source', '当前主题无可用配图')
        print(f"随机主题: {topic}")
        print("摘要内容:")
        print(summary)
        print("配图链接:")
        print(image)
    else:
        print(f"获取{topic}的摘要和配图时发生错误。")


def get_random_topic_from_category(category_name):
    api_params = {
        'action': 'query',
        'list': 'categorymembers',
        'cmtitle': f"Category:{category_name}",
        'cmlimit': '50',  # 单次请求最多返回50个页面,可按需调整,上限500
        'format': 'json'
    }
    response = requests.get("https://en.wikipedia.org/w/api.php", params=api_params)
    if response.status_code == 200:
        data = response.json()
        category_pages = data['query']['categorymembers']
        if not category_pages:
            print("指定类别下没有页面。")
            return None
        # 随机选择一个页面标题
        random_page = random.choice(category_pages)
        return random_page['title']
    else:
        print("获取类别页面列表时发生错误。")
        return None


# 示例:从"Physics experiments"类别中随机选取主题
if __name__ == "__main__":
    target_category = "Physics experiments"
    random_topic = get_random_topic_from_category(target_category)
    if random_topic:
        get_wikipedia_summary_and_image(random_topic)

注意事项

  • 如果类别下页面数量超过cmlimit上限,需要通过API返回的cmcontinue参数实现分页拉取,避免遗漏页面。
  • 部分维基页面可能没有缩略图,代码中已用get方法做容错处理。
  • 确保网络连接正常,API请求可能会因网络问题返回非200状态码。

内容的提问来源于stack exchange,提问作者Shounak Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 09:05:34