使用Python爬取ScienceDirect文献机构信息返回空结果求助
问题:无法爬取ScienceDirect文章中的机构文本
我尝试爬取链接中的机构(Affiliation)文本,目标元素为带有class="affiliation"的dl标签,其结构如下:
<dl class="affiliation"><dt><sup>a</sup></dt><dd>Department of Engineering, Università Campus Bio-Medico di Roma, Via Alvaro del Portillo, 21, 00128 Rome, Italy</dd></dl> <dl class="affiliation"><dt><sup>b</sup></dt><dd>Department of Chemical Sciences, University of Naples Federico II, Complesso Universitario di Monte Sant'Angelo, 80126 Napoli, Italy</dd></dl>
目前我能成功获取作者或标题信息,但无法提取机构文本。访问该网站需要使用代理,且网页中的机构文本需点击「Show More」按钮才会显示。我的测试代码如下:
from bs4 import BeautifulSoup import requests response = requests.get('https://www.sciencedirect.com/science/article/abs/pii/S001191642300142X') print('Response Body: ', response) soup = BeautifulSoup(response.content.decode('utf-8'), "html.parser") for aff in soup.find_all('dl', class_='affiliation'): affiliation = aff.get_text() print(affiliation)
预期输出:
Department of Engineering, Università Campus Bio-Medico di Roma, Via Alvaro del Portillo, 21, 00128 Rome, Italy Department of Chemical Sciences, University of Naples Federico II, Complesso Universitario di Monte Sant'Angelo, 80126 Napoli, Italy
问题原因及解决方案
- 静态页面未加载机构内容:机构文本是通过点击「Show More」动态加载的,直接用
requests.get获取的页面源码不包含这部分内容,必须模拟浏览器行为触发加载。 - 缺少代理配置:代码未配置代理,可能导致访问受限或被反爬拦截。
- 文本提取不精准:原代码提取整个
dl标签的文本,会包含sup标签里的标记,应直接提取dd标签的内容。
修正后的代码(使用Selenium模拟浏览器)
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # 配置代理(替换为你的代理地址和端口) chrome_options = webdriver.ChromeOptions() chrome_options.add_argument('--proxy-server=http://your-proxy-ip:port') # 初始化浏览器 driver = webdriver.Chrome(options=chrome_options) driver.get('https://www.sciencedirect.com/science/article/abs/pii/S001191642300142X') try: # 等待并点击「Show More」按钮 show_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Show More")]')) ) show_more_btn.click() time.sleep(2) # 等待内容加载完成 # 解析页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') for aff in soup.find_all('dl', class_='affiliation'): # 提取dd标签的文本并去除多余空格 affiliation_text = aff.find('dd').get_text(strip=True) print(affiliation_text) finally: # 关闭浏览器 driver.quit()
替代方案(直接调用API)
如果不想使用Selenium,可以通过浏览器开发者工具找到加载机构内容的API接口,构造请求获取数据。需要注意携带正确的请求头和Cookie,同时配置代理。
内容的提问来源于stack exchange,提问作者user17356493
相关产品推荐
相关产品推荐

