You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中是否有判断变量是否有值的方法?爬虫遇None写入错误求助

解决爬虫中TypeError: write() argument must be str, not None问题

你的问题核心是部分<a>标签没有href属性,导致link.get('href')返回None,而文件的write()方法仅接受字符串类型。Python里判断变量是否有效(非None)有两种常用方式:

  • 直接判断:if href:(会排除None、空字符串等所有假值)
  • 精准判断:if href is not None:(仅排除None,保留空字符串等特殊场景)

针对你的代码,给出修复及优化方案:

修复后的代码

from bs4 import BeautifulSoup
import requests
from urllib.parse import urljoin  # 可选:处理相对链接转绝对链接

toscan = "https://en.wikipedia.org/wiki/Wikipedia:Contents"
url = toscan
source_code = requests.get(url)
plain_text = source_code.text

# 处理文件名,避免特殊字符导致的文件创建失败
toscan_filename = toscan.replace("http://", "").replace("https://", "").replace("/", "_")

soup = BeautifulSoup(plain_text, 'html.parser')

# 把文件打开移到循环外,减少IO操作开销,用with自动管理文件关闭
with open("/home/banana/Desktop/Search engine/data/toscan", "a") as urlwaitinglist:
    for link in soup.find_all('a'):
        href = link.get('href')
        # 只处理非None的链接
        if href is not None:
            print(href)
            # 可选:将维基的相对链接转为绝对链接
            # absolute_href = urljoin(url, href)
            # urlwaitinglist.write(f'\n{absolute_href}')
            urlwaitinglist.write(f'\n{href}')

# 保存页面文本内容
results = soup.get_text().strip()
with open(f"/home/banana/Desktop/Search engine/data/Crawled Data/{toscan_filename}.txt", "w") as f:
    f.write(f'{url}\n{results}')

关键修复与优化点:

  1. 增加None判断:先将href赋值给变量,判断不为None再执行写入操作,从根源避免类型错误
  2. 优化文件操作:用with语句自动管理文件的打开与关闭,同时把文件打开逻辑移到循环外,避免重复IO操作,提升运行效率
  3. 文件名处理:替换URL中的特殊字符为下划线,避免生成无效文件名导致保存失败
  4. 可选功能:通过urljoin将维基百科的相对链接转为绝对链接,方便后续直接爬取

内容的提问来源于stack exchange,提问作者Banana628

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 03:55:14