使用BeautifulSoup的soup.find_all时触发AttributeError报错
问题原因
- 核心调用错误
soup.find_all()返回的是ResultSet类型对象,本质是存储所有匹配标签的列表,列表本身没有.find()方法,只有列表内存储的单个标签元素才能调用.find()查找子节点。你的for循环遍历的单个元素变量是list,但循环内部写的是lists.find()——直接对整个元素列表调用查找方法,才触发了属性报错。 - 潜在逻辑问题
- 用
list作为循环变量名,和Python内置的列表类型重名,容易引发不可预期的逻辑错误 - 没有做字段空值判断:部分PubMed文献没有作者地址、引用计数字段,直接访问
.text会触发NoneType错误,导致爬取中途中断 - 没有配置请求头模拟浏览器访问,大概率会被PubMed反爬机制拦截,返回和浏览器端不一致的页面内容,导致元素匹配失败
- 缺少翻页逻辑:当前URL单页最多返回200条结果,不处理翻页的话无法抓取全量符合条件的文献
修复后的可运行代码
from bs4 import BeautifulSoup import requests from csv import writer import time # 加请求头模拟浏览器访问,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } base_url= "https://pubmed.ncbi.nlm.nih.gov/?term=(%22spontaneous%20intracranial%20hypotension%22%5BAll%20Fields%5D%20OR%20%22spontaneous%20cerebrospinal%20fluid%20leak%22%5BAll%20Fields%5D%20OR%20%22cerebrospinal%20fluid%20hypovolemia%22%5BAll%20Fields%5D%20OR%20%22cerebrospinal%20fluid%20hypovolemia%20syndrome%22%5BAll%20Fields%5D%20OR%20%22Hypoliquorrhea%22%5BAll%20Fields%5D%20OR%20%22Spontaneous%20spinal%20cerebrospinal%20fluid%20leak%22%5BAll%20Fields%5D)%20NOT%20%22letter%20to%20the%20editor%22%5BAll%20Fields%5D&filter=dates.1000%2F1%2F1-2022%2F3%2F31&filter=lang.english&ac=no&format=abstract&sort=date&size=200" with open('disstest.csv', 'w', encoding= 'utf8', newline='') as f: thewriter = writer(f) header = ['Herkunftsland', 'Journal', 'Anzahl Zitationen'] thewriter.writerow(header) # 初始页码从1开始,可根据总页数调整循环范围实现全量爬取 for page_num in range(1, 20): url = f"{base_url}&page={page_num}" page = requests.get(url, headers=headers) soup = BeautifulSoup(page.content, 'html.parser') # 变量名改成articles,避免和内置list重名 articles = soup.find_all('article', class_="article-overview") # 如果当前页没有文献条目,终止循环 if not articles: break for article in articles: # 字段提取加空值判断,避免报错中断 country_ele = article.find('ul', class_="item-list") herkunftsland = country_ele.text.replace('\n','').strip() if country_ele else "无相关信息" journal_ele = article.find('div', class_="article-source") journal = journal_ele.text.replace('\n', '').strip() if journal_ele else "无相关信息" cite_ele = article.find('li', class_="references-count") zitationen = cite_ele.text.replace('\n', '').strip() if cite_ele else "0" info = [herkunftsland, journal, zitationen] thewriter.writerow(info) # 加延时,避免请求过快被封 time.sleep(2)
全量爬取注意事项
- 代码示例里的页码循环范围可自行调整,你可以先打开第一页查看符合条件的文献总条数,除以单页200条的数量得到总页数,替换循环范围即可抓完全部数据;代码里已经加了空页判断,即使循环范围写大了也会自动终止。
- 爬取时建议保留1-3秒的请求间隔,避免请求频率过高被PubMed临时封禁IP。
- 如果需要抓取的文献量较大,优先使用PubMed官方提供的E-utilities接口获取结构化XML/JSON数据,比解析前端HTML页面稳定,不会因为PubMed前端改版导致爬虫失效。
内容的提问来源于stack exchange,提问作者clannishduke
相关产品推荐
相关产品推荐

