3D单词列表转2D?Pandas文本列分句分词处理方案问询
解决Pandas文本列的句子拆分与分词问题
没问题,这事儿用spaCy处理起来非常顺手,我给你一步步拆解实现方法:
1. 核心函数实现
首先我们要写一个函数,输入单个文本字符串,输出句子→单词列表的嵌套结构,还要自动过滤掉标点符号(比如你例子里的句号):
import spacy import pandas as pd # 加载spaCy英文模型(注意:新版本spaCy需要用'en_core_web_sm',旧版本可用'en') nlp = spacy.load('en_core_web_sm') def text_to_words(text): """将输入文本拆分为句子,每个句子分词为单词列表(过滤标点)""" doc = nlp(text) sentence_word_lists = [] # 遍历每个句子 for sent in doc.sents: # 提取句子中的单词,跳过标点 words = [token.text for token in sent if not token.is_punct] sentence_word_lists.append(words) return sentence_word_lists
2. 处理Pandas列
用你的测试数据来验证:
# 模拟Pandas列的输入数据 s = ["How are you. Don't wait for me", "this is all fine"] df = pd.DataFrame({'text': s}) # 对列应用函数,得到3D结构(行→句子→单词) df['processed_text'] = df['text'].map(text_to_words)
此时df['processed_text']的输出就是:
0 [['How', 'are', 'you'], ["Don't", 'wait', 'for', 'me']] 1 [['this', 'is', 'all', 'fine']] Name: processed_text, dtype: object
3. 展平为2D列表
如果你想把所有文档的句子合并成一个统一的2D列表(也就是你要的最终结果),用列表推导式就能轻松搞定:
flattened_2d = [sentence for doc in df['processed_text'] for sentence in doc]
最终输出就是:
[["How", "are", "you"], ["Don't", "wait", "for", "me"], ["this", "is", "all", "fine"]]
额外补充:如果需要展平为一维单词列表
要是你之后需要把所有单词放到一个一维列表里,只需要再加一层推导:
flattened_1d = [word for doc in df['processed_text'] for sentence in doc for word in sentence]
输出会是:
["How", "are", "you", "Don't", "wait", "for", "me", "this", "is", "all", "fine"]
内容的提问来源于stack exchange,提问作者Baktaawar
相关产品推荐
相关产品推荐

