You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup中按特定字符分割文本提取出版物与大学信息

解决方案

核心思路

问题的关键是识别并使用**长破折号(Em Dash,字符编码U+2014)**进行分割,而非普通短破折号。分割后通过strip()清理前后冗余空格,即可分别提取出版物名和大学名。

示例代码

from bs4 import BeautifulSoup
import requests

# 目标网页
target_url = "http://spajournalism.com/membership/"
response = requests.get(target_url)
soup = BeautifulSoup(response.text, "html.parser")

result_list = []

# 替换为你实际定位出版物链接的选择器
for link in soup.select("a[href]"):
    # 获取父元素文本并清理首尾空格
    full_text = link.parent.text.strip()
    
    # 用长破折号分割文本(注意:这里是Em Dash,不是短破折号-)
    split_parts = full_text.split(" — ")
    
    # 确保分割出两个有效部分
    if len(split_parts) == 2:
        pub_name = split_parts[0].strip()
        uni_name = split_parts[1].strip()
        pub_url = link["href"]
        
        # 组装目标字典并加入列表
        result_list.append({
            "publication_url": pub_url,
            "publication_name": pub_name,
            "uni": uni_name
        })

# 输出结果
for item in result_list:
    print(item)

排查分割失败的情况

如果上述代码无法正常分割,大概率是网页中的长破折号存在字符变体(比如全角破折号、其他编码的特殊破折号)。可以通过以下方式确认字符编码:

# 打印文本中每个字符的编码,定位破折号对应的编码
for char in full_text:
    print(f"字符: '{char}',编码: {ord(char)}")

假设查到破折号编码为8212(标准Em Dash),也可以直接用chr(8212)分割:

split_parts = full_text.split(chr(8212))

内容的提问来源于stack exchange,提问作者shanodin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 12:33:14