如何用Polars实现从Stanza Word对象提取属性生成自定义嵌套列表?
问题:用Polars替代Pandas处理Stanza标注结果,提取嵌套Word列表
我之前用Pandas处理Stanza的文本标注结果:通过apply调用带双层循环的函数,把Stanza生成的Document对象转换成外层对应整篇文本、内层对应句子的嵌套Word列表。现在想改用Polars库,想问能不能通过Polars原生API实现,还是必须沿用类似Pandas的实现方式?
附原Pandas示例代码:
import stanza import pandas as pd from typing import NamedTuple nlp = stanza.Pipeline('en') class Word(NamedTuple): id: int head_id: int text: str span: list[int] def get_doc_words(doc: stanza.Document) -> list[list[Word]]: doc_words = [] for sentence in doc.sentences: sentence_words = [] for sent_word in sentence.words: word = Word( id=sent_word.id, head_id=sent_word.head, text=sent_word.text, span=[sent_word.start_char, sent_word.end_char], ) sentence_words.append(word) doc_words.append(sentence_words) return doc_words df=pd.DataFrame( { 'text': [ 'This is some sample text. A second sentence.', 'And a second sample. Having a second sentence as well' ] } ) df['stanza_annotation'] = df['text'].apply(nlp) df['stanza_words'] = df['stanza_annotation'].apply(get_doc_words)
预期输出(单条文本对应的嵌套列表):
[[Word(id=1, head_id=5, text='This', span=[0, 4]), Word(id=2, head_id=5, text='is', span=[5, 7]), Word(id=3, head_id=5, text='some', span=[8, 12]), Word(id=4, head_id=5, text='sample', span=[13, 19]), Word(id=5, head_id=0, text='text', span=[20, 24]), Word(id=6, head_id=5, text='.', span=[24, 25])], [Word(id=1, head_id=3, text='A', span=[26, 27]), Word(id=2, head_id=3, text='second', span=[28, 34]), Word(id=3, head_id=0, text='sentence', span=[35, 43]), Word(id=4, head_id=3, text='.', span=[43, 44])]]
解答
完全可以用Polars实现,核心逻辑可以复用你之前写的get_doc_words函数,搭配Polars的map_elements方法即可,写法贴合Polars API风格,大规模数据场景下性能比Pandas更有优势。
Polars实现代码
import stanza import polars as pl from typing import NamedTuple nlp = stanza.Pipeline('en') class Word(NamedTuple): id: int head_id: int text: str span: list[int] def get_doc_words(doc: stanza.Document) -> list[list[Word]]: doc_words = [] for sentence in doc.sentences: sentence_words = [] for sent_word in sentence.words: word = Word( id=sent_word.id, head_id=sent_word.head, text=sent_word.text, span=[sent_word.start_char, sent_word.end_char], ) sentence_words.append(word) doc_words.append(sentence_words) return doc_words # 创建Polars DataFrame df = pl.DataFrame({ 'text': [ 'This is some sample text. A second sentence.', 'And a second sample. Having a second sentence as well' ] }) # 生成标注结果和嵌套Word列表 df = df.with_columns( # 生成Stanza标注对象 pl.col('text').map_elements(nlp).alias('stanza_annotation'), # 直接调用现有函数生成嵌套列表 pl.col('stanza_annotation').map_elements(get_doc_words).alias('stanza_words') ) # 查看结果 print(df['stanza_words'].to_list())
说明
- Polars的
map_elements和Pandas的apply功能类似,用于对列中每个元素应用自定义函数; - 你之前写的
get_doc_words函数可以直接复用——因为Stanza的Document对象需要遍历句子和单词提取属性,这部分逻辑无法通过纯向量化API替代,遍历是必要的; - 处理大规模数据时,建议用Polars的懒执行模式(
pl.LazyFrame)进一步提升性能:df = pl.LazyFrame({ 'text': [ 'This is some sample text. A second sentence.', 'And a second sample. Having a second sentence as well' ] }) df = df.with_columns( pl.col('text').map_elements(nlp).alias('stanza_annotation'), pl.col('stanza_annotation').map_elements(get_doc_words).alias('stanza_words') ).collect()
内容的提问来源于stack exchange,提问作者exch_cmmnt_memb
相关产品推荐
相关产品推荐

