You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现文本文件单词行号索引及现有代码问题排查

问题修正方案

原代码核心问题

  • 函数build_word_index定义在while循环内部,且全程没有被调用,不会产生任何执行结果
  • 已编译的正则匹配规则未投入使用,直接按空格拆分字符串会导致空值、带特殊字符的无效单词被计入索引
  • 行读取逻辑和函数内的文本拆分逻辑冲突,重复做了行遍历操作

修正后可运行代码

import sys
import re

# 提前编译正则,匹配字母数字组成的单词
pattern = re.compile(r"[a-zA-Z0-9]+")

def build_word_index():
    out = {}
    # 直接逐行遍历标准输入流,行号从1开始计数
    for line_num, line in enumerate(sys.stdin, start=1):
        # 用正则提取所有符合规则的单词,规避空格拆分带来的冗余问题
        words = pattern.findall(line)
        for word in words:
            # 如果不需要区分单词大小写可保留下面这行,需要区分大小写直接删除即可
            word = word.lower()
            if word not in out:
                out[word] = []
            # 避免同一行内重复出现的单词重复记录行号,需要统计出现次数可删除该判断
            if line_num not in out[word]:
                out[word].append(line_num)
    return out

if __name__ == "__main__":
    word_index = build_word_index()
    # 按单词字典序输出结果
    for word, line_nums in sorted(word_index.items()):
        print(f"{word}: {line_nums}")

使用方法

运行时通过标准输入传入目标文本文件即可,示例命令:
python word_index.py < 目标文本文件路径.txt

内容的提问来源于stack exchange,提问作者Feverish123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 21:24:03