You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则匹配美元符号前后5个单词及词频统计方案问询

匹配美元符号前后各5个单词并统计频次

嘿,我来帮你搞定这个需求!你需要匹配$金额前后各5个单词,然后统计它们的出现频次,这里有两种实用的方法,分别适合不同场景:

方法一:正则表达式实现

如果你的文本格式比较规整(单词都是纯字母、标点和单词明显分离),可以用正则快速实现。核心思路是先定位$金额,再提取前后的单词:

import re
from collections import Counter

string = 'I have a sentence with $10.00 within it and this sentence is done. '

# 定义匹配金额及前后单词的正则
pattern = re.compile(r'(?:\b[a-zA-Z]+\b\s+){0,5}\$\d+(?:\.\d+)?\s+(?:\b[a-zA-Z]+\b\s?){0,5}')
match_result = pattern.search(string)

# 拆分前后部分,提取单词并截取目标数量
if match_result:
    # 拆分出$符号前后的文本
    before_part, after_part = match_result.group().split('$')
    # 提取前面最后5个单词
    preceding_words = re.findall(r'\b[a-zA-Z]+\b', before_part)[-5:]
    # 提取后面前5个单词
    following_words = re.findall(r'\b[a-zA-Z]+\b', after_part)[:5]
    # 合并结果
    surrounding = preceding_words + following_words
    print("匹配到的前后单词:", surrounding)
    
    # 统计频次并保持原顺序
    final_return = sorted(Counter(surrounding).items(), key=lambda x: surrounding.index(x[0]))
    print("统计结果:", final_return)

正则说明:

  • \b[a-zA-Z]+\b:匹配一个纯字母组成的单词(\b是单词边界,避免匹配单词片段)
  • (?:\b[a-zA-Z]+\b\s+){0,5}:匹配0到5个单词(非捕获组,避免多余分组)
  • \$\d+(?:\.\d+)?:匹配美元金额(支持整数如$10或小数如$10.00)

方法二:NLTK分词器实现(更稳健)

如果你的文本是真实场景下的复杂文本(比如包含缩写、连字符单词、不规则标点),用NLTK的分词器会更可靠,因为它专门处理自然语言的分词逻辑:

步骤:

  1. 先安装NLTK并下载分词模型(第一次运行需要)
  2. 对整个字符串分词
  3. 定位所有包含$的token位置
  4. 提取该位置前后各5个单词
  5. 过滤非单词token并统计频次
import nltk
from collections import Counter

# 下载分词所需的punkt模型(第一次运行执行)
nltk.download('punkt')

string = 'I have a sentence with $10.00 within it and this sentence is done. '

# 对字符串进行分词
tokens = nltk.word_tokenize(string)

# 找到所有包含$符号的token的索引
dollar_positions = [idx for idx, token in enumerate(tokens) if '$' in token]

surrounding = []
for pos in dollar_positions:
    # 提取当前位置前5个单词(处理边界情况,避免索引越界)
    preceding_tokens = tokens[max(0, pos - 5):pos]
    # 提取当前位置后5个单词
    following_tokens = tokens[pos + 1:pos + 6]
    # 合并到结果列表
    surrounding.extend(preceding_tokens + following_tokens)

# 过滤掉非字母的token(比如标点符号)
surrounding = [token for token in surrounding if token.isalpha()]
print("匹配到的前后单词:", surrounding)

# 统计频次并保持原出现顺序
final_return = sorted(Counter(surrounding).items(), key=lambda x: surrounding.index(x[0]))
print("统计结果:", final_return)

优势:

  • NLTK的word_tokenize能正确处理复杂场景,比如把don't拆成do和n't,把带标点的单词和标点分离(你可以根据需求调整过滤规则)
  • 自动处理边界情况(比如$在字符串开头时,前5个单词不足5个,不会报错)

两种方法对比

  • 正则法:代码简洁,适合简单规整的文本,执行速度快
  • NLTK法:更稳健,适合真实世界的复杂文本,处理边缘情况更靠谱

内容的提问来源于stack exchange,提问作者JRR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:15:42