You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup爬取期刊数据时在DataFrame对应行正确存储多作者信息?

期刊爬虫修复:作者与机构信息跨行存储问题

问题说明

爬取期刊数据时,论文标题、关键词可正常存入DataFrame,但作者与机构信息存在异常:单篇论文中除首位作者外,其余作者均被存储到新行,机构信息同理。这导致数据关联性缺失,DataFrame行数混乱,无法正常使用。需要将单篇论文的所有作者以「姓名1, 姓名2, 姓名3...」格式存入对应行,机构信息做相同处理,但不清楚如何调整BS4选择器逻辑。

现有爬虫代码

title = []
authors = []
afiliations = []

for i in urls: 
    page = requests.get(link)
    content = page.text
    soup = BeautifulSoup(content, "html.parser")
    for t in soup.select(".obj_article_details .page_title"):
        title.append(t.get_text(strip=True))
    for au in soup.select(".obj_article_details .authors .name"):
        authors.append(au.get_text(strip=True))
    for af in soup.select(".obj_article_details .item.authors .affiliation"):
        affiliations.append(af.get_text(strip=True))
    time.sleep(3)

目标页面结构

...
<article class="obj_article_details">   
 <h1 class="page_title">
   Lorem ipsum dolor sit amet
 </h1>
  <div class="row">
   <div class="main_entry">
<section class="item authors">
  <ul class="authors">
    <li>
      <span class="name">Brandon Scott </span>
      <span class="affiliation"> Villanova University, Pennsylvania </span>
    </li>
    <li>
      <span class="name">Alvaro Cote </span>
      <span class="affiliation">Carleton College, Minnesota</span>
    </li>
  </ul>
</section>

...

当前与期望的DataFrame格式

当前格式

|Authors       |  Affiliation                       | 
    +--------------+------------------------------------+
    |Brandon Scott | Villanova University, Pennsylvania |
    +--------------+------------------------------------+
    |Alvaro Cote   | Carleton College, Minnesota        |
    +--------------+------------------------------------+
    |...           | ...                                |

期望格式

|Authors                     |  Affiliation                          | 
    +----------------------------+---------------------------------------+
    |Brandon Scott, Alvaro Cote  | Villanova University, Pennsylvania, Carleton College, Minnesota |
    +----------------------------+---------------------------------------+
    |...                         |...                                    |
    +----------------------------+---------------------------------------+

修复后的代码

import requests
from bs4 import BeautifulSoup
import time
import pandas as pd

title = []
authors = []
afiliations = []

for link in urls: 
    page = requests.get(link)
    content = page.text
    soup = BeautifulSoup(content, "html.parser")
    # 遍历单页内的每一篇论文容器
    for article in soup.select(".obj_article_details"):
        # 提取单篇论文标题
        title_elem = article.select_one(".page_title")
        title.append(title_elem.get_text(strip=True) if title_elem else "")
        
        # 提取该论文所有作者并拼接为字符串
        author_names = [au.get_text(strip=True) for au in article.select(".authors .name")]
        authors.append(", ".join(author_names))
        
        # 提取该论文所有机构并拼接为字符串
        aff_names = [af.get_text(strip=True) for af in article.select(".item.authors .affiliation")]
        afiliations.append(", ".join(aff_names))
    
    time.sleep(3)

# 生成结构化DataFrame
df = pd.DataFrame({
    "Title": title,
    "Authors": authors,
    "Affiliations": afiliations
})

修复逻辑说明

  1. 按论文容器遍历:不再单独遍历标题、作者、机构,而是先定位每篇论文的根容器.obj_article_details,确保所有数据都属于当前论文。
  2. 批量提取拼接:对单篇论文内的所有作者/机构,先收集为列表,再用", ".join()拼接成单个字符串后加入列表,保证每个列表的元素数量与论文数量完全匹配。
  3. 容错处理:添加if title_elem else ""避免因页面结构异常导致的报错。

内容的提问来源于stack exchange,提问作者drupaljac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 19:55:13