You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中定位CSV文件句子里词性标记为NN和VB的单词位置

在Python中定位CSV句子里NN和VB词性单词的位置

嘿,这个需求完全能搞定!我来给你一步步讲清楚怎么实现:

准备工作

首先咱们需要两个实用工具库:pandas用来轻松读取和批量处理CSV文件,nltk用来给单词做词性标注。先安装它们:

pip install pandas nltk

第一次运行时,还得下载NLTK的分词和词性标注模型,你可以在Python里执行以下代码完成下载:

import nltk
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')

完整实现代码

用Pandas处理CSV(推荐,适合批量数据)

假设你的CSV文件结构如下(列名为sentence):

sentence
"Man walks into a bar."
"Cop shoots his gun."
"Kid drives into a ditch"

可以用下面的代码快速处理:

import pandas as pd
import nltk
from nltk.tokenize import word_tokenize

# 下载必要的NLTK数据(第一次运行需要)
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')

def get_nn_vb_positions(sentence):
    # 把句子拆分成单个单词
    tokens = word_tokenize(sentence)
    # 给每个单词标注词性
    tagged_words = nltk.pos_tag(tokens)
    
    # 收集符合要求的单词和位置
    result = {}
    for index, (word, pos_tag) in enumerate(tagged_words):
        # 匹配所有名词类(NN开头,比如NN、NNP、NNS)和动词类(VB开头,比如VB、VBZ、VBD)
        # 如果你只需要严格的NN和VB标记,就改成 pos_tag == 'NN' or pos_tag == 'VB'
        if pos_tag.startswith('NN') or pos_tag.startswith('VB'):
            # 这里位置从1开始计数,要是习惯0索引就直接用index
            result[word] = (index + 1, pos_tag)
    return result

# 读取CSV文件
df = pd.read_csv('your_sentences.csv')
# 对每一行的句子处理,生成新列存储结果
df['nn_vb_info'] = df['sentence'].apply(get_nn_vb_positions)

# 打印结果看看
print(df[['sentence', 'nn_vb_info']])

运行后会得到类似这样的输出:

sentence                                          nn_vb_info
0      Man walks into a bar.  {'Man': (1, 'NNP'), 'walks': (2, 'VBZ'), 'bar...
1       Cop shoots his gun.  {'Cop': (1, 'NNP'), 'shoots': (2, 'VBZ'), 'gu...
2  Kid drives into a ditch   {'Kid': (1, 'NN'), 'drives': (2, 'VBZ'), 'dit...

用内置CSV模块处理(轻量选项)

如果不想依赖Pandas,用Python内置的csv模块也能实现:

import csv
import nltk
from nltk.tokenize import word_tokenize

nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')

def get_pos_details(sentence):
    tokens = word_tokenize(sentence)
    tagged_words = nltk.pos_tag(tokens)
    details = []
    for idx, (word, tag) in enumerate(tagged_words):
        if tag.startswith('NN') or tag.startswith('VB'):
            details.append(f"单词「{word}」:位置{idx+1},词性{tag}")
    return '; '.join(details)

# 读取CSV并处理
with open('your_sentences.csv', 'r', encoding='utf-8') as file:
    reader = csv.DictReader(file)
    for row in reader:
        print(f"句子:{row['sentence']}")
        print(f"NN/VB单词详情:{get_pos_details(row['sentence'])}\n")

一些小提示

  • 词性标记变体:NLTK的词性标记里,NN开头的都是名词(比如NNP是专有名词,NNS是复数名词),VB开头的都是动词(VBZ是第三人称单数现在时,VBD是过去式),如果你只需要严格的NN和VB标记,直接修改判断条件即可。
  • 标点处理:word_tokenize会把标点单独拆分,比如句子末尾的.会被算成一个token,如果你想忽略标点,在分词后可以过滤掉非字母的token,比如tokens = [t for t in word_tokenize(sentence) if t.isalpha()],但这样位置计数会跳过标点,要根据你的需求调整。

内容的提问来源于stack exchange,提问作者Beginner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 15:59:07