如何用Polars对DataFrame指定列应用命名实体识别并添加Person标签列?
使用Polars结合Spacy实现高效命名实体识别
问题背景
原使用Pandas对DataFrame指定列执行命名实体识别时速度过慢,希望改用Polars实现相同功能,同时保留Spacy的实体识别逻辑。
示例数据与期望结果
原Pandas示例数据:
import pandas as pd df = pd.DataFrame({'source': ['Paul', 'Paul'], 'target': ['GOOGLE', 'Ferrari'], 'edge': ['works at', 'drive'] })
对应表格:
| source | target | edge |
|---|---|---|
| Paul | works at | |
| Paul | Ferrari | drive |
期望Polars输出结果:
| source | target | edge | Entity |
|---|---|---|---|
| Paul | works at | Person | |
| Paul | Ferrari | drive | Person |
已准备的Spacy基础代码:
!python -m spacy download en_core_web_sm
import spacy nlp = spacy.load('en_core_web_sm') # Pandas中低效的写法 df['Entities'] = df['Text'].apply(lambda sent: [(ent.label_) for ent in nlp(sent).ents])
实现步骤
1. 安装依赖并加载Spacy模型
先确保安装好所需库:
pip install polars spacy python -m spacy download en_core_web_sm
2. 创建Polars DataFrame
import polars as pl import spacy # 加载Spacy预训练模型 nlp = spacy.load('en_core_web_sm') # 创建Polars格式的DataFrame df_pl = pl.DataFrame({ 'source': ['Paul', 'Paul'], 'target': ['GOOGLE', 'Ferrari'], 'edge': ['works at', 'drive'] })
3. 定义实体识别函数
针对source列提取PERSON类型的实体标签:
def extract_person_entity(text): doc = nlp(text) # 提取第一个匹配的PERSON标签,无匹配则返回None for ent in doc.ents: if ent.label_ == 'PERSON': return ent.label_ return None
4. 为Polars DataFrame添加Entity列
使用Polars的with_columns结合map_elements实现高效列处理,性能优于Pandas的apply:
df_pl = df_pl.with_columns( pl.col('source').map_elements(extract_person_entity).alias('Entity') ) # 查看最终结果 print(df_pl)
执行后输出:
shape: (2, 4) ┌────────┬─────────┬──────────┬────────┐ │ source ┆ target ┆ edge ┆ Entity │ │ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ str ┆ str ┆ str │ ╞════════╪═════════╪══════════╪════════╡ │ Paul ┆ GOOGLE ┆ works at ┆ Person │ │ Paul ┆ Ferrari ┆ drive ┆ Person │ └────────┴─────────┴──────────┴────────┘
大数据量优化方案
如果处理超大规模数据,可使用Spacy的pipe方法批量处理,进一步提升速度:
def batch_extract_person(texts): docs = nlp.pipe(texts) results = [] for doc in docs: ent_label = None for ent in doc.ents: if ent.label_ == 'PERSON': ent_label = ent.label_ break results.append(ent_label) return results # 批量处理生成Entity列 df_pl = df_pl.with_columns( pl.Series(name='Entity', values=batch_extract_person(df_pl['source'].to_list())) )
内容的提问来源于stack exchange,提问作者Ozioh
相关产品推荐
相关产品推荐

