You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup提取作者机构信息遇问题:返回None及元素缺失处理咨询

Python网页爬取问题解决方案

1. 遍历所有span返回None的问题

遍历所有span却拿不到数据,大概率是这几个原因:

  • 属性名错误:你要找的作者/机构属性名可能写错了,比如把data-author写成author,或者目标元素根本没这个属性。先打开网页开发者工具,确认目标span的属性到底叫什么。
  • 元素是动态加载:如果网页用JS渲染内容,requests直接获取的HTML里根本没有这些span,这时候得用Selenium、Playwright这类工具模拟浏览器加载页面。
  • 遍历逻辑低效且不准确:没必要遍历所有span,直接用属性选择器定位更靠谱,比如soup.find_all('span', attrs={'data-author': True}),只筛选有目标属性的span。

示例修正代码:

from bs4 import BeautifulSoup
import requests

url = "你的目标文章URL"
resp = requests.get(url)
soup = BeautifulSoup(resp.text, 'html.parser')

# 直接筛选带目标属性的span
target_spans = soup.find_all('span', attrs={'data-author': True})
for span in target_spans:
    author = span.get('data-author')  # 用get()避免属性不存在报错
    print(author)

2. 清理文本中的多余空格和换行

直接用BeautifulSoup的get_text(strip=True)方法就能自动去掉首尾空格和多余换行,比手动处理方便:

authors = soup.find_all('span', class_='author-name')
institutions = soup.find_all('span', class_='author-affiliation')

for auth, inst in zip(authors, institutions):
    # 清理文本
    clean_author = auth.get_text(strip=True)
    clean_institution = inst.get_text(strip=True)
    # 配对输出
    print(f"作者:{clean_author} | 机构:{clean_institution}")

如果还有顽固的连续空格,也可以用正则替换:

import re
clean_institution = re.sub(r'\s+', ' ', inst.text).strip()

3. 解决元素缺失时.text报错的问题

当用find()找不到元素时会返回None,直接调用.text或.get_text()就会触发AttributeError,给你三个实用解决方法:

方法1:先判断元素是否存在再取值

author_elem = soup.find('span', class_='author-name')
# 三元表达式处理
author = author_elem.get_text(strip=True) if author_elem else "未知作者"

institution_elem = soup.find('span', class_='author-affiliation')
institution = institution_elem.get_text(strip=True) if institution_elem else "未知机构"

方法2:用try-except捕获异常

适合怕漏判断的场景:

try:
    author = soup.find('span', class_='author-name').get_text(strip=True)
except AttributeError:
    author = "未知作者"

try:
    institution = soup.find('span', class_='author-affiliation').get_text(strip=True)
except AttributeError:
    institution = "未知机构"

方法3:封装成工具函数重复使用

如果要多次处理,写个函数更高效:

def get_clean_text(element, default="未知"):
    if element:
        return element.get_text(strip=True)
    return default

author = get_clean_text(soup.find('span', class_='author-name'))
institution = get_clean_text(soup.find('span', class_='author-affiliation'))

内容的提问来源于stack exchange,提问作者Nuno Rodrigues

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 13:07:35