You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Wikipedia API抓取公司维基页面首图时偶发失效问题排查

问题原因与解决方案

为什么部分页面无法获取图片?

你的代码依赖维基百科API的pageimages属性,但这个属性仅返回被**标记为页面首图(pageimage)**的图片。像Binance、Coinbase这类页面,编辑者可能未将logo设置为pageimage属性,或是logo仅存在于信息框(Infobox)中但未关联到pageimage,导致API返回空结果。

改用prop=images获取页面所有图片,再通过文件名关键词筛选logo,或取页面首张图片作为备选。以下是修改后的代码:

import requests

API_ENDPOINT = "https://en.wikipedia.org/w/api.php"
title = "Binance"

# 第一步:获取页面所有图片列表
params = {
    "action": "query",
    "format": "json",
    "formatversion": 2,
    "prop": "images",
    "titles": title,
    "imlimit": 50  # 限制获取前50张图,足够筛选
}

response = requests.get(API_ENDPOINT, params=params)
data = response.json()

pages = data.get('query', {}).get('pages', [])
if not pages:
    print("页面不存在")
else:
    page = pages[0]
    images = page.get('images', [])
    if not images:
        print("页面无图片")
    else:
        # 筛选文件名含logo/Logo的图片
        logo_candidates = [img for img in images if 'logo' in img['title'].lower()]
        target_img_title = logo_candidates[0]['title'] if logo_candidates else images[0]['title']
        
        # 第二步:获取目标图片的原图链接
        img_params = {
            "action": "query",
            "format": "json",
            "formatversion": 2,
            "prop": "imageinfo",
            "iiprop": "url",
            "titles": target_img_title
        }
        img_response = requests.get(API_ENDPOINT, params=img_params)
        img_data = img_response.json()
        
        img_pages = img_data.get('query', {}).get('pages', [])
        if img_pages:
            img_info = img_pages[0].get('imageinfo', [])[0]
            print("图片URL:", img_info['url'])

代码说明

  1. 先调用API获取页面所有图片的标题列表,通过imlimit控制返回数量
  2. 优先筛选文件名包含logo的图片(大部分企业logo会用这个关键词命名)
  3. 若没有logo候选,直接取页面第一张图片
  4. 二次调用API获取图片的原图URL,因为images属性仅返回图片标题,不包含直接访问链接

内容的提问来源于stack exchange,提问作者Zumplo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 08:47:20