You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python维基百科爬虫函数输出无法写入文件的技术求助

问题解决:递归维基爬虫结果无法写入文件

你的代码核心问题是**scrapeWikiArticle函数没有返回值**——它只是用print把标题输出到控制台,所以output = scrapeWikiArticle(...)得到的是None,而file.write(None)会直接报错,因为文件写入需要字符串类型的内容。

另外原代码还有个隐藏问题:没有递归终止条件,会无限爬下去直到栈溢出,必须一起处理。

修改方案

  1. 让函数收集所有爬取到的标题并返回,而不是直接打印
  2. 添加递归终止条件,避免无限递归
  3. 写入文件时指定编码,防止特殊字符乱码

修改后的完整代码:

import requests
from bs4 import BeautifulSoup
import random
import sys

# 设置递归深度上限,避免无限递归导致崩溃
sys.setrecursionlimit(10)

def scrapeWikiArticle(url):
    response = requests.get(url=url)
    soup = BeautifulSoup(response.content, 'html.parser')
    title = soup.find(id="firstHeading")
    current_title = title.text + "\n"  # 加换行符,让每个标题占一行
    
    allLinks = soup.find(id="bodyContent").find_all("a")
    random.shuffle(allLinks)
    linkToScrape = None
    
    # 过滤出有效的维基条目链接,排除特殊页面
    for link in allLinks:
        if link['href'].startswith("/wiki/") and not link['href'].startswith("/wiki/Special:"):
            linkToScrape = link
            break
    
    # 找到有效链接就递归爬取,拼接结果;没找到就终止递归
    if linkToScrape:
        next_url = "https://en.wikipedia.org" + linkToScrape['href']
        return current_title + scrapeWikiArticle(next_url)
    else:
        return current_title

# 调用函数获取所有标题文本
output = scrapeWikiArticle("https://en.wikipedia.org/wiki/Web_scraping")

# 写入文件,指定utf-8编码避免乱码
with open('output.txt', 'w', encoding='utf-8') as file:
    file.write(output)

关键修改点说明

  • 函数不再用print输出,而是把当前标题和后续递归的结果拼接后返回,这样output就能拿到所有标题的字符串
  • 用sys.setrecursionlimit(10)限制递归次数,你可以根据需求调整这个数值
  • 优化了链接过滤逻辑,跳过Special类的非条目页面,避免爬取无效内容
  • 写入文件时指定encoding='utf-8',确保包含特殊字符的标题能正确写入

内容的提问来源于stack exchange,提问作者Derrick Crash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 16:42:53