You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的BeautifulSoup和requests爬取Minecraft Wiki图片URL遇缩进错误

解决Python爬虫的IndentationError缩进错误问题

我尝试用Python的BeautifulSoup和requests库爬取Minecraft Wiki方块页面的图片URL,但在第35行代码for i in soup.find_all('img'):处遇到了IndentationError: unexpected indent错误,请求帮助定位并解决问题。

错误原因

Python是严格依赖缩进的编程语言,这段代码里的内层for循环缩进层级不匹配:它属于外层for block in blocks:循环的内部代码,但当前缩进比外层循环内的其他代码(比如soup = BeautifulSoup(response.text, 'html.parser'))多,导致解释器无法识别正确的代码块层级,触发缩进错误。

同时代码还有两个额外问题需要修正:

  1. 遍历img标签时,i本身就是img元素,不需要再用i.find('img', class_='mw-file-element'),这会导致找不到目标元素。
  2. 字典键写成了'image:'(带冒号),后续遍历images列表时用image['image']会触发KeyError。

修正后的完整代码

from bs4 import BeautifulSoup
import requests
import re

# URL of the website containing the block names and images
url = 'https://minecraft.wiki/w/Block'

# Send a GET request to the URL
response = requests.get(url)

# Parse the HTML content of the page
soup = BeautifulSoup(response.text, 'html.parser')

# Find all elements containing block names and images
blocks = []
for li in soup.find_all('li'):
    name_element = li.find('a', class_='mw-redirect')
    
    if name_element:
        name = name_element.text.strip()
        url = 'https://minecraft.wiki/w/' + re.sub(r'\s+', '_', name)
    
        blocks.append({'name': name, 'url': url})

for block in blocks:
    print(block['name'])
    print(block['url'])

images = []
for block in blocks:
    block_url = block['url']
    response = requests.get(block_url)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 修正缩进,和外层循环内代码保持同一层级
    for i in soup.find_all('img'):
        # 直接判断当前img元素的class是否符合要求
        if 'mw-file-element' in i.get('class', []):
            srcset = i.get('srcset', '')
            original_image_url = re.search(r'^(.+\.png)', srcset)

            if original_image_url:
                image_url = 'https://minecraft.wiki' + original_image_url.group(1)
                # 修正字典键,去掉多余的冒号
                images.append({'image': image_url})

for image in images:
     print(image['image'])

关键修改点

  • 缩进修正:将内层for i in soup.find_all('img'):及其内部代码的缩进调整为4个空格,和外层循环内的代码(如soup = BeautifulSoup(...))保持一致,确保代码块层级正确。
  • img元素判断优化:直接检查当前遍历到的img元素的class列表是否包含mw-file-element,避免无效的嵌套查找。
  • 字典键修正:把{'image:' : image_url}改为{'image': image_url},保证后续取值时不会出现KeyError。

内容的提问来源于stack exchange,提问作者giacomo abb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 06:12:05