如何在Python中定位CSV文件句子里词性标记为NN和VB的单词位置
在Python中定位CSV句子里NN和VB词性单词的位置
嘿,这个需求完全能搞定!我来给你一步步讲清楚怎么实现:
准备工作
首先咱们需要两个实用工具库:pandas用来轻松读取和批量处理CSV文件,nltk用来给单词做词性标注。先安装它们:
pip install pandas nltk
第一次运行时,还得下载NLTK的分词和词性标注模型,你可以在Python里执行以下代码完成下载:
import nltk nltk.download('punkt') nltk.download('averaged_perceptron_tagger')
完整实现代码
用Pandas处理CSV(推荐,适合批量数据)
假设你的CSV文件结构如下(列名为sentence):
sentence "Man walks into a bar." "Cop shoots his gun." "Kid drives into a ditch"
可以用下面的代码快速处理:
import pandas as pd import nltk from nltk.tokenize import word_tokenize # 下载必要的NLTK数据(第一次运行需要) nltk.download('punkt') nltk.download('averaged_perceptron_tagger') def get_nn_vb_positions(sentence): # 把句子拆分成单个单词 tokens = word_tokenize(sentence) # 给每个单词标注词性 tagged_words = nltk.pos_tag(tokens) # 收集符合要求的单词和位置 result = {} for index, (word, pos_tag) in enumerate(tagged_words): # 匹配所有名词类(NN开头,比如NN、NNP、NNS)和动词类(VB开头,比如VB、VBZ、VBD) # 如果你只需要严格的NN和VB标记,就改成 pos_tag == 'NN' or pos_tag == 'VB' if pos_tag.startswith('NN') or pos_tag.startswith('VB'): # 这里位置从1开始计数,要是习惯0索引就直接用index result[word] = (index + 1, pos_tag) return result # 读取CSV文件 df = pd.read_csv('your_sentences.csv') # 对每一行的句子处理,生成新列存储结果 df['nn_vb_info'] = df['sentence'].apply(get_nn_vb_positions) # 打印结果看看 print(df[['sentence', 'nn_vb_info']])
运行后会得到类似这样的输出:
sentence nn_vb_info 0 Man walks into a bar. {'Man': (1, 'NNP'), 'walks': (2, 'VBZ'), 'bar... 1 Cop shoots his gun. {'Cop': (1, 'NNP'), 'shoots': (2, 'VBZ'), 'gu... 2 Kid drives into a ditch {'Kid': (1, 'NN'), 'drives': (2, 'VBZ'), 'dit...
用内置CSV模块处理(轻量选项)
如果不想依赖Pandas,用Python内置的csv模块也能实现:
import csv import nltk from nltk.tokenize import word_tokenize nltk.download('punkt') nltk.download('averaged_perceptron_tagger') def get_pos_details(sentence): tokens = word_tokenize(sentence) tagged_words = nltk.pos_tag(tokens) details = [] for idx, (word, tag) in enumerate(tagged_words): if tag.startswith('NN') or tag.startswith('VB'): details.append(f"单词「{word}」:位置{idx+1},词性{tag}") return '; '.join(details) # 读取CSV并处理 with open('your_sentences.csv', 'r', encoding='utf-8') as file: reader = csv.DictReader(file) for row in reader: print(f"句子:{row['sentence']}") print(f"NN/VB单词详情:{get_pos_details(row['sentence'])}\n")
一些小提示
- 词性标记变体:NLTK的词性标记里,NN开头的都是名词(比如NNP是专有名词,NNS是复数名词),VB开头的都是动词(VBZ是第三人称单数现在时,VBD是过去式),如果你只需要严格的
NN和VB标记,直接修改判断条件即可。 - 标点处理:
word_tokenize会把标点单独拆分,比如句子末尾的.会被算成一个token,如果你想忽略标点,在分词后可以过滤掉非字母的token,比如tokens = [t for t in word_tokenize(sentence) if t.isalpha()],但这样位置计数会跳过标点,要根据你的需求调整。
内容的提问来源于stack exchange,提问作者Beginner
相关产品推荐
相关产品推荐

