调用Python爬虫函数时参数取值失败,触发NameError错误
爬虫函数调用报错及修复方案
问题概述
调用基于requests与BeautifulSoup编写的crawler函数时,传入publicationDetails[5]["link"]触发NameError,提示名称'publicationDetails'未定义,同时函数内部存在参数逻辑混乱、变量使用不当等问题。
原代码
import requests from bs4 import BeautifulSoup def crawler(seed,maxcount,publicationDetails): Q=[seed] count=0 while(Q!=[] and count<maxcount): count+=1 url=Q.pop(0) print(url) code=requests.get(url) plain=code.text s=BeautifulSoup(plain,"html.parser") for link in s.findAll("a"): newurl=link.get("href") if(newurl!=None and newurl!="/"): newurl=newurl.strip() if(newurl[0:5]!="http://" and newurl[0:6]!="http://"): if(url[len(url) -1] == '/'): newurl=url + newurl else: newurl=url + "/"+ newurl Q.append(newurl) publicationDetails=[] for page in range(0,1): page1=requests.get(newurl) soup=BeautifulSoup(page1.text,"html.parser") publications=soup.find_all("h3",{"class":"title"}) for this_publications in publications: if(this_publications.find("a")): publicationDetails+=[{"name":this_publications.string, "link":this_publications.find("a")["href"] }] crawler("https://pureportal.coventry.ac.uk/en/organisations/school-of-economics-finance-and-accounting/persons/",12,publicationDetails[5]["link"])
错误信息
NameError Traceback (most recent call last) <ipython-input-23-924a9dc1f5b3> in <module>() ----> 1 crawler("https://pureportal.coventry.ac.uk/en/organisations/school-of-economics-finance-and-accounting/persons/",12,publicationDetails[5]["link"]) NameError: 名称'publicationDetails'未定义
问题分析
- 调用参数未定义:调用函数时传入的
publicationDetails[5]["link"]中,publicationDetails从未提前定义,直接使用必然触发NameError。 - 函数参数逻辑混乱:函数定义时接收
publicationDetails参数,但内部直接将其重新赋值为空列表publicationDetails=[],导致传入的参数完全无效。 - URL判断错误:原代码中判断URL是否为绝对路径时,两次判断都是
http://,漏掉了https://的情况。 - 变量使用风险:后续爬取出版物时使用的
newurl是前面广度遍历中的临时变量,其值不确定,可能指向无效页面。 - 无意义循环:
for page in range(0,1)仅执行一次,完全没必要使用循环结构。
修复后的代码
import requests from bs4 import BeautifulSoup def crawler(seed, maxcount): Q = [seed] count = 0 # 广度遍历爬取链接 while Q and count < maxcount: count += 1 url = Q.pop(0) print(url) try: # 捕获请求异常,避免程序崩溃 response = requests.get(url) response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") for link in soup.find_all("a"): new_url = link.get("href") if new_url and new_url != "/": new_url = new_url.strip() # 修复绝对URL判断逻辑,支持http和https if not (new_url.startswith("http://") or new_url.startswith("https://")): # 处理相对URL拼接 if url.endswith('/'): new_url = url + new_url else: new_url = f"{url}/{new_url}" Q.append(new_url) except requests.exceptions.RequestException as e: print(f"请求{url}失败: {str(e)}") continue publication_details = [] # 爬取目标页面的出版物信息(此处以seed页面为例,可根据需求调整) try: response = requests.get(seed) response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") publications = soup.find_all("h3", class_="title") for pub in publications: a_tag = pub.find("a") if a_tag: publication_details.append({ "name": pub.string.strip() if pub.string else "", "link": a_tag["href"] }) except requests.exceptions.RequestException as e: print(f"获取出版物信息失败: {str(e)}") return publication_details # 调用函数,接收返回的出版物列表 publication_details = crawler( "https://pureportal.coventry.ac.uk/en/organisations/school-of-economics-finance-and-accounting/persons/", 12 ) # 安全访问第5个元素的link if len(publication_details) > 5: print(publication_details[5]["link"]) else: print("出版物数据不足,无法获取第5个元素的链接")
修复说明
- 删除了无效的
publicationDetails参数,改为函数返回出版物列表,避免参数传递混乱。 - 修复了绝对URL的判断逻辑,新增对
https://的支持。 - 添加了请求异常捕获,处理网络请求失败的情况,避免程序直接崩溃。
- 去掉了无意义的单循环结构,简化代码逻辑。
- 调用时先获取返回的出版物列表,再判断长度后访问索引,避免索引越界错误。
- 规范了变量命名(如
newurl改为new_url),提升代码可读性。
内容的提问来源于stack exchange,提问作者Arjun
相关产品推荐
相关产品推荐

