使用Python提取文件中指定模式间字符:获取目标驯养动物列表
解决从文件提取指定驯养动物的问题
先帮你梳理下原代码的几个核心问题,这也是导致效果不符合预期的原因:
- 文件指针耗尽问题:外层循环遍历
temp_list的两个元素,内层循环却直接遍历整个文件——第一次循环读完文件后,文件指针已经到末尾了,第二次循环根本不会再读取任何内容,自然只能处理第一个domesticated_0的内容。 - 变量名混乱:代码里混用了
line和li(比如li.endswith("}")但遍历的变量是line),运行时会直接抛出NameError。 - 逻辑偏离需求:你的代码是把匹配行和后续行写入文件,但实际需求是提取指定列表里的动物名称,不是输出整行内容。
下面给你两种贴合需求的可行方案:
方案一:提取指定驯养类别下的所有动物
这个方案直接匹配你指定的domesticated_0和domesticated_1类别,提取对应行大括号内的动物,刚好能得到你要的['sheep','cow','buffalo']:
target_categories = ['domesticated_0', 'domesticated_1'] result = [] # 用with语句自动管理文件,避免手动关闭的麻烦 with open('your_file.txt', 'r') as input_file, open('output.txt', 'w') as output_file: for line in input_file: line = line.strip() # 去除首尾空白(包括换行符) # 判断当前行是否是目标类别的animal定义行 if line.startswith('animal') and any(cat in line for cat in target_categories): # 定位大括号的位置,提取内部内容 brace_start = line.find('{') + 1 brace_end = line.find('}') if brace_start > 0 and brace_end > brace_start: # 拆分动物名称,过滤掉可能的空字符串 animals = [animal.strip() for animal in line[brace_start:brace_end].split() if animal.strip()] result.extend(animals) # 将提取到的动物写入输出文件 output_file.write('\n'.join(animals) + '\n') print("提取到的动物:", result)
方案二:精准提取目标列表中的动物
如果你需要严格只提取['sheep','cow','buffalo']这些动物(不管它们属于哪个类别),可以用这个方案:
target_animals = {'sheep', 'cow', 'buffalo'} # 用集合提高查找效率 result = [] with open('your_file.txt', 'r') as input_file, open('output.txt', 'w') as output_file: for line in input_file: line = line.strip() # 判断是否是包含动物列表的animal行 if line.startswith('animal') and '{' in line and '}' in line: brace_start = line.find('{') + 1 brace_end = line.find('}') animals_in_line = [animal.strip() for animal in line[brace_start:brace_end].split() if animal.strip()] # 筛选出目标动物 matched_animals = [animal for animal in animals_in_line if animal in target_animals] result.extend(matched_animals) output_file.write('\n'.join(matched_animals) + '\n') print("提取到的目标动物:", result)
关键改进说明:
- 用
with语句操作文件,无需手动关闭,更安全简洁。 - 直接从匹配行的大括号内提取内容,避免了错误的跨行读取逻辑(你的示例中所有动物列表都在同一行)。
- 处理空白字符,避免提取到无效的空字符串。
内容的提问来源于stack exchange,提问作者sid
相关产品推荐
相关产品推荐

