如何用Python按列读取文本文件并生成含单词及列位置的列表
竖排文本转单词-列位置列表实现方案
原始文本内容
// H HGD // i eoa // loy // ld // o // _ // W // o // r // l // d
需求说明
需要将上述竖排文本转换为单词+对应列位置的列表,格式要求为:[[<1st_word>,<location of the 1st word>],[<2nd_word>,<location of the 2nd word>],...]
期望输出结果:[["Hi",8],["Hello_World",10],["Good",11],["Day",12]]
现有代码瓶颈
已编写部分代码,但未实现核心逻辑,代码如下:
file = open(text_file, "r") wrd = "" for line in file: #Each line from file for c in ln: #Each character in line if(c != ' ' or c != '\n' or c != '/'): #Ignore slashes (/), white spaces, and newlines #list = put word on a list and count at which column was found #print(list)
具体实现代码与说明
核心思路
原始文本的单词是按列竖排的,因此需要从列的维度遍历文本,将同一列的有效字符按行顺序拼接成单词,同时记录该列的位置(从1开始计数)。
完整实现代码
def parse_vertical_text(text_file): # 读取所有行,保留每行原始长度(去除换行符) lines = [] with open(text_file, 'r') as f: for line in f: lines.append(line.rstrip('\n')) # 确定最大列数(最长行的字符数) max_col = max(len(line) for line in lines) if lines else 0 result = [] # 遍历每一列 for col in range(max_col): current_word = [] # 遍历每行,提取当前列的有效字符 for line in lines: if col < len(line): char = line[col] # 过滤无效字符:斜杠、空格、换行符 if char not in ('/', ' ', '\n'): current_word.append(char) # 仅保留有有效字符的列(即拼接出单词的列) if current_word: # 列位置从1开始计数,对应需求中的数字 result.append([''.join(current_word), col + 1]) return result # 调用示例(替换为你的文件名) final_output = parse_vertical_text('vertical_text.txt') print(final_output)
关键说明
- 按列遍历:打破常规的行遍历逻辑,改为从列维度收集字符,适配竖排单词的结构
- 无效字符过滤:直接判断字符是否为
/或空格,跳过无意义字符 - 列位置计算:需求中的列数从1开始计数,因此用
col + 1转换索引(索引从0开始) - 安全文件操作:使用
with语句自动管理文件句柄,避免资源泄漏
运行代码后,将得到与需求完全匹配的输出结果。
内容的提问来源于stack exchange,提问作者myles_uy
相关产品推荐
相关产品推荐

