Python递归网页爬虫函数执行完成后返回额外None值问题求助
问题修复方案
None值产生原因
- 原函数仅在
results不为空时显式返回结果,若当前页面无符合规则的URL,Python会为没有显式return语句的函数默认返回None - 代码中直接使用
append接收递归调用的返回值,就会将无结果时返回的None追加到列表末尾
递归逻辑问题
- 递归调用返回的是子页面爬取到的URL列表,用
append会将整个列表作为单个元素存入当前列表,产生嵌套结构,不符合预期的平铺URL存储需求
修复方案
修改recursiveLinkSearch函数的两处核心逻辑:
- 移除
if results != []的返回判断,无论结果是否为空都显式返回results - 将递归调用的
append改为extend,合并子页面的爬取结果
修复后完整代码
def curlURL(url): # beautify with BS soup = BeautifulSoup(requests.get(url, timeout=3).text, "html.parser") return soup def recursiveLinkSearch(soup, url, layer, depth): results = [] # 提前判断层数,达到最大深度直接返回空列表减少无效遍历 if layer >= depth: return results # 扫描当前页面所有带href属性的标签 for a in soup.find_all(href=True): try: # 校验URL规则 if any(stringStartsWith in a.get('href')[0:4] for stringStartsWith in ["http", "https", "HTTP", "HTTPS"]) \ and a.get('href') != url: print(f"Found URL: {a.get('href')}") print(f"LOG: {colors.yellow}Current Layer: {layer}{colors.end}") results.append(a.get('href')) # 用extend合并子页面爬取结果 results.extend(recursiveLinkSearch(curlURL(a.get('href')), a.get('href'), layer+1, depth)) # 异常处理逻辑保持不变 except requests.exceptions.InvalidSchema: print(f"{a.get('href')}") print(f"{colors.bad}Invalid Url Detected{colors.end}") except requests.exceptions.ConnectTimeout: print(f"{a.get('href')}") print(f"{colors.bad}Connection Timeout. Passing...") except requests.exceptions.SSLError: print(f"{a.get('href')}") print(f"{colors.bad}SSL Certificate Error. Passing...") except requests.exceptions.ReadTimeout: print(f"{a.get('href')}") print(f"{colors.bad}Read Timeout. Passing...") # 无论结果是否为空都显式返回 return results
内容的提问来源于stack exchange,提问作者Gabriel C.
相关产品推荐
相关产品推荐

