BeautifulSoup提取作者机构信息遇问题:返回None及元素缺失处理咨询
Python网页爬取问题解决方案
1. 遍历所有span返回None的问题
遍历所有span却拿不到数据,大概率是这几个原因:
- 属性名错误:你要找的作者/机构属性名可能写错了,比如把
data-author写成author,或者目标元素根本没这个属性。先打开网页开发者工具,确认目标span的属性到底叫什么。 - 元素是动态加载:如果网页用JS渲染内容,requests直接获取的HTML里根本没有这些span,这时候得用Selenium、Playwright这类工具模拟浏览器加载页面。
- 遍历逻辑低效且不准确:没必要遍历所有span,直接用属性选择器定位更靠谱,比如
soup.find_all('span', attrs={'data-author': True}),只筛选有目标属性的span。
示例修正代码:
from bs4 import BeautifulSoup import requests url = "你的目标文章URL" resp = requests.get(url) soup = BeautifulSoup(resp.text, 'html.parser') # 直接筛选带目标属性的span target_spans = soup.find_all('span', attrs={'data-author': True}) for span in target_spans: author = span.get('data-author') # 用get()避免属性不存在报错 print(author)
2. 清理文本中的多余空格和换行
直接用BeautifulSoup的get_text(strip=True)方法就能自动去掉首尾空格和多余换行,比手动处理方便:
authors = soup.find_all('span', class_='author-name') institutions = soup.find_all('span', class_='author-affiliation') for auth, inst in zip(authors, institutions): # 清理文本 clean_author = auth.get_text(strip=True) clean_institution = inst.get_text(strip=True) # 配对输出 print(f"作者:{clean_author} | 机构:{clean_institution}")
如果还有顽固的连续空格,也可以用正则替换:
import re clean_institution = re.sub(r'\s+', ' ', inst.text).strip()
3. 解决元素缺失时.text报错的问题
当用find()找不到元素时会返回None,直接调用.text或.get_text()就会触发AttributeError,给你三个实用解决方法:
方法1:先判断元素是否存在再取值
author_elem = soup.find('span', class_='author-name') # 三元表达式处理 author = author_elem.get_text(strip=True) if author_elem else "未知作者" institution_elem = soup.find('span', class_='author-affiliation') institution = institution_elem.get_text(strip=True) if institution_elem else "未知机构"
方法2:用try-except捕获异常
适合怕漏判断的场景:
try: author = soup.find('span', class_='author-name').get_text(strip=True) except AttributeError: author = "未知作者" try: institution = soup.find('span', class_='author-affiliation').get_text(strip=True) except AttributeError: institution = "未知机构"
方法3:封装成工具函数重复使用
如果要多次处理,写个函数更高效:
def get_clean_text(element, default="未知"): if element: return element.get_text(strip=True) return default author = get_clean_text(soup.find('span', class_='author-name')) institution = get_clean_text(soup.find('span', class_='author-affiliation'))
内容的提问来源于stack exchange,提问作者Nuno Rodrigues
相关产品推荐
相关产品推荐

