You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本提取网页链接遇'NoneType'不可下标错误,求协助

解决Python提取维基百科链接时的TypeError: 'NoneType' object is not subscriptable问题

嘿,这个错误原因很清晰:你遍历的<a>标签里,有些标签根本没有href属性——当你用link.get('href')获取属性时,这类标签会返回None,而直接对None执行[0:4]切片操作,自然就触发了TypeError。

下面是修正后的解决方案:

核心改动点

  • 先判断links是否不为None,过滤掉没有href属性的标签
  • 用str.startswith()替代切片判断,代码更简洁,还能避免短链接的潜在问题
  • 修正URL拼接的重复错误(原代码会生成https://en.wikipedia.org/wiki/wiki/xxx这类无效链接)

修正后的代码

from bs4 import BeautifulSoup
import requests

my_url = 'https://en.wikipedia.org/wiki/Kashmir'
response = requests.get(my_url)
page_soup = BeautifulSoup(response.content, "html.parser")

for link in page_soup.find_all('a'):
    links = link.get('href')
    # 先确保links存在,再检查链接格式
    if links is not None and links.startswith('/wiki') and links != '#':
        print(f"https://en.wikipedia.org{links}")

这样运行代码就不会再报错,还能正确提取所有维基百科的站内链接啦。

内容的提问来源于stack exchange,提问作者Bhat Ovas Qayoom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:54:24