You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用Python爬虫函数时参数取值失败,触发NameError错误

爬虫函数调用报错及修复方案

问题概述

调用基于requests与BeautifulSoup编写的crawler函数时,传入publicationDetails[5]["link"]触发NameError,提示名称'publicationDetails'未定义,同时函数内部存在参数逻辑混乱、变量使用不当等问题。

原代码

import requests
from bs4 import BeautifulSoup
def crawler(seed,maxcount,publicationDetails):
  Q=[seed]
  count=0
  while(Q!=[] and count<maxcount):
    count+=1
    url=Q.pop(0)
    print(url)
    code=requests.get(url)
    plain=code.text
    s=BeautifulSoup(plain,"html.parser")
    for link in s.findAll("a"):
      newurl=link.get("href")
      if(newurl!=None and newurl!="/"):
        newurl=newurl.strip()
        if(newurl[0:5]!="http://" and newurl[0:6]!="http://"):
          if(url[len(url) -1] == '/'):
            newurl=url + newurl
          else:
            newurl=url + "/"+ newurl
        Q.append(newurl)

  publicationDetails=[]
  for page in range(0,1):
     
     page1=requests.get(newurl)
     soup=BeautifulSoup(page1.text,"html.parser")
     publications=soup.find_all("h3",{"class":"title"})
     for this_publications in publications:
            if(this_publications.find("a")):
                publicationDetails+=[{"name":this_publications.string,
                              "link":this_publications.find("a")["href"]
                              }]
crawler("https://pureportal.coventry.ac.uk/en/organisations/school-of-economics-finance-and-accounting/persons/",12,publicationDetails[5]["link"])

错误信息

NameError                                 Traceback (most recent call last)
<ipython-input-23-924a9dc1f5b3> in <module>()
----> 1 crawler("https://pureportal.coventry.ac.uk/en/organisations/school-of-economics-finance-and-accounting/persons/",12,publicationDetails[5]["link"])

NameError: 名称'publicationDetails'未定义

问题分析

  1. 调用参数未定义:调用函数时传入的publicationDetails[5]["link"]中,publicationDetails从未提前定义,直接使用必然触发NameError。
  2. 函数参数逻辑混乱:函数定义时接收publicationDetails参数,但内部直接将其重新赋值为空列表publicationDetails=[],导致传入的参数完全无效。
  3. URL判断错误:原代码中判断URL是否为绝对路径时,两次判断都是http://,漏掉了https://的情况。
  4. 变量使用风险:后续爬取出版物时使用的newurl是前面广度遍历中的临时变量,其值不确定,可能指向无效页面。
  5. 无意义循环:for page in range(0,1)仅执行一次,完全没必要使用循环结构。

修复后的代码

import requests
from bs4 import BeautifulSoup

def crawler(seed, maxcount):
    Q = [seed]
    count = 0
    # 广度遍历爬取链接
    while Q and count < maxcount:
        count += 1
        url = Q.pop(0)
        print(url)
        try:
            # 捕获请求异常,避免程序崩溃
            response = requests.get(url)
            response.raise_for_status()
            soup = BeautifulSoup(response.text, "html.parser")
            
            for link in soup.find_all("a"):
                new_url = link.get("href")
                if new_url and new_url != "/":
                    new_url = new_url.strip()
                    # 修复绝对URL判断逻辑,支持http和https
                    if not (new_url.startswith("http://") or new_url.startswith("https://")):
                        # 处理相对URL拼接
                        if url.endswith('/'):
                            new_url = url + new_url
                        else:
                            new_url = f"{url}/{new_url}"
                    Q.append(new_url)
        except requests.exceptions.RequestException as e:
            print(f"请求{url}失败: {str(e)}")
            continue

    publication_details = []
    # 爬取目标页面的出版物信息(此处以seed页面为例,可根据需求调整)
    try:
        response = requests.get(seed)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")
        publications = soup.find_all("h3", class_="title")
        
        for pub in publications:
            a_tag = pub.find("a")
            if a_tag:
                publication_details.append({
                    "name": pub.string.strip() if pub.string else "",
                    "link": a_tag["href"]
                })
    except requests.exceptions.RequestException as e:
        print(f"获取出版物信息失败: {str(e)}")
    
    return publication_details

# 调用函数,接收返回的出版物列表
publication_details = crawler(
    "https://pureportal.coventry.ac.uk/en/organisations/school-of-economics-finance-and-accounting/persons/",
    12
)

# 安全访问第5个元素的link
if len(publication_details) > 5:
    print(publication_details[5]["link"])
else:
    print("出版物数据不足,无法获取第5个元素的链接")

修复说明

  • 删除了无效的publicationDetails参数,改为函数返回出版物列表,避免参数传递混乱。
  • 修复了绝对URL的判断逻辑,新增对https://的支持。
  • 添加了请求异常捕获,处理网络请求失败的情况,避免程序直接崩溃。
  • 去掉了无意义的单循环结构,简化代码逻辑。
  • 调用时先获取返回的出版物列表,再判断长度后访问索引,避免索引越界错误。
  • 规范了变量命名(如newurl改为new_url),提升代码可读性。

内容的提问来源于stack exchange,提问作者Arjun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 10:06:22