You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取文章链接时如何避免AttributeError?

网页爬取问题分析与修复

我正在尝试用BeautifulSoup做网页爬取,之前提问过但需求描述不清,只得到部分解答。我想提取网页内容后再从中提取所有链接,现在代码出错了,帮我分析问题。

现有网页内容获取代码

# 定义要获取的网页URL
quote_page = 'https://bigbangtheory.fandom.com/wiki/Barry_Kripke'

# 请求网页
http = urllib3.PoolManager()
r = http.request('GET', quote_page)

if r.status == 200:
    page = r.data
    print(f'变量"page"的类型: {page.__class__.__name__}')
    print(f'网页获取成功。请求状态: {r.status}, 页面大小:{len(page)}')
else:
    print(f'出现问题。请求状态: {r.status}')

# 将字节流转换为BeautifulSoup对象
soup = BeautifulSoup(page, 'html.parser')
print(f'变量"soup"的类型: {soup.__class__.__name__}')

# 查看部分内容
print(f'{soup.prettify()[:1000]}')

# 检查HTML标题
print(f'Title标签: {soup.title}')
print(f'Title文本: {soup.title.string}')

# 查找主要内容
article_tag = 'p'
articles = soup.find_all(article_tag)
print(f'变量"articles"的类型:{articles.__class__.__name__}')

for p in articles:
    print(p.text)

提取链接的错误代码

# 查找文本中的链接
# 指定要获取的标签类型
tag = 'a'

# 从`<a>`标签创建链接列表
tag_list = [t.get('href') for t in articles.find_all(tag)]
tag_list

错误原因

articles是soup.find_all('p')返回的ResultSet对象(多个标签的集合),它本身没有find_all方法——find_all是单个BeautifulSoup标签对象的方法,不能直接在ResultSet集合上调用。

另外原代码里还有一处小问题:print(f'Type of the variable "article":{article.__class__.__name__}')中的article变量未定义,应该改为articles。

修复方案

需要遍历articles里的每个<p>标签,再在每个标签内部查找<a>标签,代码修改如下:

# 查找文本中的链接
tag = 'a'

# 方式1:常规循环
tag_list = []
for p in articles:
    links_in_p = p.find_all(tag)
    for link in links_in_p:
        href = link.get('href')
        if href:  # 过滤无href属性的空链接
            tag_list.append(href)

# 方式2:列表推导式简化写法
tag_list = [link.get('href') for p in articles for link in p.find_all(tag) if link.get('href')]

print(tag_list)

内容的提问来源于stack exchange,提问作者WMD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 22:01:04