You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python正则表达式提取新闻标题中首段连续大写字母公司名

提取新闻标题中的目标公司与次要公司名称

首先明确咱们要遵循的核心规则:

  • 目标公司名称:标题中首段连续以大写字母开头的词汇链
  • 次要公司名称:标题中第二段连续以大写字母开头的词汇链

我用Python写了一个简单高效的实现方案,既能处理常规的公司名称,也能兼容像&、撇号这类特殊字符的情况:

实现代码

import re

def extract_companies(headline):
    # 匹配连续大写开头的单词序列,兼容包含&、撇号的公司名称
    capitalized_word_chains = re.findall(r'([A-Z][a-z\'&]+(?:\s+[A-Z][a-z\'&]+)*)', headline)
    if not capitalized_word_chains:
        return {"target_company": None, "secondary_company": None}
    # 第一组为目标公司,第二组(若存在)为次要公司
    result = {
        "target_company": capitalized_word_chains[0],
        "secondary_company": capitalized_word_chains[1] if len(capitalized_word_chains) >= 2 else None
    }
    return result

# 测试示例标题
headlines = [ 
    "Chicago Policemen's Annuity & Benefit Fund hired Chicago Equity Partners to manage $50 million in active U.S. smidcap value equity.", 
    "Belmont Contributory Retirement System is searching for at least one U.S. small-cap equity manager to run initially up to $5 million.", 
    "Phoenix Employees' Deferred Compensation Board will begin a search for an investment consultant before the end of February." 
]

# 遍历测试并输出结果
for idx, headline in enumerate(headlines, 1):
    companies = extract_companies(headline)
    print(f"标题{idx}:")
    print(f"目标公司: {companies['target_company']}")
    print(f"次要公司: {companies['secondary_company']}\n")

运行结果

标题1:
目标公司: Chicago Policemen's Annuity & Benefit Fund
次要公司: Chicago Equity Partners

标题2:
目标公司: Belmont Contributory Retirement System
次要公司: None

标题3:
目标公司: Phoenix Employees' Deferred Compensation Board
次要公司: None

补充说明

  • 正则表达式([A-Z][a-z\'&]+(?:\s+[A-Z][a-z\'&]+)*)专门用来匹配连续的大写开头单词,同时允许单词内部包含撇号(比如Policemen's)和&这类特殊字符
  • 如果标题里存在像U.S.这类带点的缩写,因为正则里没有匹配.,所以不会被误纳入公司名称链,刚好符合示例里的情况

内容的提问来源于stack exchange,提问作者Merv Merzoug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:19:02