You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3读取文本生成数值列表返回双重括号问题排查

解决Python提取数字列表时的双重括号问题

我来帮你搞定这个双重括号的问题!先拆解下问题根源:

你用re.findall(r'(?<=").*?(?=")', line)的时候,这个方法本身返回的就是一个匹配结果的列表——哪怕每行只匹配到一个数字,它也会返回[数字]这样的结构。而你之后用txt_numbers.append(num),相当于把这个小列表整个当成一个元素塞进了txt_numbers里,自然就出现了[[num1], [num2], ...]的双重嵌套结构。

解决方案1:用extend代替append

把添加元素的方式从append改成extend,extend会把传入列表里的每个元素单独加到目标列表中,而不是把整个列表作为一个元素添加:

txt_numbers = []
for file in os.listdir(my_path):
    if file[-3:] == "txt":
        with open(my_path + file, 'r') as txt_file:
            txt = txt_file.readlines()
            for line in txt:
                num = re.findall(r'(?<=").*?(?=")', line)
                txt_numbers.extend(num)  # 这里替换成extend

解决方案2:遍历匹配结果逐个添加

如果你想更清晰地控制每个元素,也可以遍历num里的结果逐个append:

txt_numbers = []
for file in os.listdir(my_path):
    if file[-3:] == "txt":
        with open(my_path + file, 'r') as txt_file:
            txt = txt_file.readlines()
            for line in txt:
                nums = re.findall(r'(?<=").*?(?=")', line)
                for n in nums:
                    txt_numbers.append(n)

额外优化建议

顺便给你两个提升代码健壮性的小技巧:

  • 用file.endswith("txt")代替file[-3:] == "txt":如果遇到文件名长度不足3位的异常情况(比如.txt这种奇怪命名),切片会报错,endswith更安全。
  • 用os.path.join(my_path, file)拼接路径:不同操作系统的路径分隔符不一样(Windows是\,Linux是/),手动拼接容易出问题,os.path.join会自动适配系统规则。

优化后的完整代码:

import os
import re

txt_numbers = []
for file in os.listdir(my_path):
    if file.endswith("txt"):
        with open(os.path.join(my_path, file), 'r') as txt_file:
            # 不需要一次性readlines,直接遍历文件对象更省内存
            for line in txt_file:
                nums = re.findall(r'(?<=").*?(?=")', line)
                txt_numbers.extend(nums)

这样处理后,txt_numbers就会变成[num1, num2, num3, ...]的一维列表,刚好符合你后续检索PDF发票的需求。

内容的提问来源于stack exchange,提问作者user9371615

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:22:42