You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python代码在Colab正常输出,VSCode报索引越界错误求助

问题描述

以下Python爬虫代码在Colab中运行可正常输出“Candy Crush Saga”,但在VSCode中执行时会抛出索引越界异常(index out of bound exception)。代码如下:

from urllib import request
from lxml import html
import requests

class AppCrawler:
    def __init__(self,starting_url,depth):
        self.starting_url=starting_url
        self.depth=depth
        self.apps=[]

    def crawl(self):
        self.get_app_from_link(self.starting_url)
        return 

    def get_app_from_link(self,link):
        startpage=requests.get(link)
        tree=html.fromstring(startpage.text)
        name= tree.xpath('//h1[@class="product-header__title app-header__title"]/text()')[0]
        print(name)


class App:
    def __init__(self,name,developer,price,links):
        self.name=name
        self.developer=developer
        self.price=price
        self.links=links

    def __str__(self):
        return("Name: "+ self.name.encode('UTF-8')+ "\n Developer: "+ self.developer.encode('UTF-8')+
        "\n Price: "+ self.price .encode('UTF-8'))

crawler=AppCrawler("https://apps.apple.com/us/app/candy-crush-saga/id553834731",0)
crawler.crawl()
问题原因与修复方案

核心原因

Colab环境的requests请求默认带有符合浏览器特征的User-Agent,而本地VSCode环境中,requests.get()的默认请求头被苹果服务器识别为非浏览器请求(爬虫),返回了不含目标h1标签的页面(比如反爬验证页或简化结构页),导致XPath查询结果为空列表,直接取索引[0]时触发索引越界异常。

修复步骤

  1. 添加模拟浏览器的请求头:给requests.get()设置User-Agent字段,让服务器返回正常的App Store页面。
  2. 增加异常判断逻辑:避免XPath查询结果为空时直接取索引导致崩溃。
  3. 修复字符串拼接问题:原代码__str__方法中的encode('UTF-8')会返回字节串,Python3中直接与字符串拼接会报错,需移除该操作。

修复后的代码示例:

from urllib import request
from lxml import html
import requests

class AppCrawler:
    def __init__(self,starting_url,depth):
        self.starting_url=starting_url
        self.depth=depth
        self.apps=[]
        # 模拟Chrome浏览器请求头
        self.headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
        }

    def crawl(self):
        self.get_app_from_link(self.starting_url)
        return 

    def get_app_from_link(self,link):
        startpage=requests.get(link, headers=self.headers)
        startpage.raise_for_status()  # 检查请求是否成功返回200状态码
        tree=html.fromstring(startpage.text)
        # 先获取查询结果列表,再判断是否存在目标内容
        name_list = tree.xpath('//h1[@class="product-header__title app-header__title"]/text()')
        if name_list:
            name = name_list[0].strip()
            print(name)
        else:
            print("未找到应用名称:可能页面结构变化或触发反爬拦截")


class App:
    def __init__(self,name,developer,price,links):
        self.name=name
        self.developer=developer
        self.price=price
        self.links=links

    def __str__(self):
        return("Name: "+ self.name + "\n Developer: "+ self.developer +
        "\n Price: "+ self.price)

crawler=AppCrawler("https://apps.apple.com/us/app/candy-crush-saga/id553834731",0)
crawler.crawl()

内容的提问来源于stack exchange,提问作者Roshan Scaria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 16:15:35