You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取ScienceDirect文献机构信息返回空结果求助

问题:无法爬取ScienceDirect文章中的机构文本

我尝试爬取链接中的机构(Affiliation)文本,目标元素为带有class="affiliation"的dl标签,其结构如下:

<dl class="affiliation"><dt><sup>a</sup></dt><dd>Department of Engineering, Università Campus Bio-Medico di Roma, Via Alvaro del Portillo, 21, 00128 Rome, Italy</dd></dl>

<dl class="affiliation"><dt><sup>b</sup></dt><dd>Department of Chemical Sciences, University of Naples Federico II, Complesso Universitario di Monte Sant'Angelo, 80126 Napoli, Italy</dd></dl>

目前我能成功获取作者或标题信息,但无法提取机构文本。访问该网站需要使用代理,且网页中的机构文本需点击「Show More」按钮才会显示。我的测试代码如下:

from bs4 import BeautifulSoup
import requests

response = requests.get('https://www.sciencedirect.com/science/article/abs/pii/S001191642300142X')
print('Response Body: ', response)
soup = BeautifulSoup(response.content.decode('utf-8'), "html.parser")

for aff in soup.find_all('dl', class_='affiliation'):
    affiliation = aff.get_text()
    print(affiliation)

预期输出:

Department of Engineering, Università Campus Bio-Medico di Roma, Via Alvaro del Portillo, 21, 00128 Rome, Italy
Department of Chemical Sciences, University of Naples Federico II, Complesso Universitario di Monte Sant'Angelo, 80126 Napoli, Italy

问题原因及解决方案

  1. 静态页面未加载机构内容:机构文本是通过点击「Show More」动态加载的,直接用requests.get获取的页面源码不包含这部分内容,必须模拟浏览器行为触发加载。
  2. 缺少代理配置:代码未配置代理,可能导致访问受限或被反爬拦截。
  3. 文本提取不精准:原代码提取整个dl标签的文本,会包含sup标签里的标记,应直接提取dd标签的内容。

修正后的代码(使用Selenium模拟浏览器)

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# 配置代理(替换为你的代理地址和端口)
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument('--proxy-server=http://your-proxy-ip:port')

# 初始化浏览器
driver = webdriver.Chrome(options=chrome_options)
driver.get('https://www.sciencedirect.com/science/article/abs/pii/S001191642300142X')

try:
    # 等待并点击「Show More」按钮
    show_more_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Show More")]'))
    )
    show_more_btn.click()
    time.sleep(2)  # 等待内容加载完成

    # 解析页面源码
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    for aff in soup.find_all('dl', class_='affiliation'):
        # 提取dd标签的文本并去除多余空格
        affiliation_text = aff.find('dd').get_text(strip=True)
        print(affiliation_text)
finally:
    # 关闭浏览器
    driver.quit()

替代方案(直接调用API)

如果不想使用Selenium,可以通过浏览器开发者工具找到加载机构内容的API接口,构造请求获取数据。需要注意携带正确的请求头和Cookie,同时配置代理。

内容的提问来源于stack exchange,提问作者user17356493

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:22:01