You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式拆分产品记录异常,求完整提取方法

问题解决:提取单行产品列表中的完整产品记录

你当前用re.split(r',00')分割单行产品列表时,出现记录断行、不完整的情况,核心原因包括:

  • 产品内容中存在类似4,40G的逗号,会被误分割;
  • 部分产品结尾没有,00标识(比如示例中第一个产品仅以重量结尾),导致分割后记录缺失;
  • 分割后未将,00补回原记录,丢失价格结尾标识。

解决方案

通过正则正向先行断言定位产品的起始分隔点(每个0453.开头的位置,除了第一个),精准分割每个完整产品。同时读取整个文件内容确保无遗漏,处理所有产品记录。

修改后的txt_format函数

def txt_format():
    # 读取整个单行文件内容
    with open('catalogo_no_linebreak.txt', mode='r', encoding='utf-8') as arq:
        full_content = arq.read()
    
    # 用正向先行断言分割:匹配"0453."前面的位置(排除字符串开头)
    products = re.split(r'(?=0453\.)', full_content)
    
    # 过滤空字符串,清理无效内容
    products = [p.strip() for p in products if p.strip()]
    
    # 写入结果文件
    with open('saida.txt', mode='w', encoding='utf-8') as out_arq:
        for product in products:
            out_arq.write(product + '\n')

正则逻辑说明

(?=0453\.)是正向先行断言,它仅匹配一个位置——这个位置后面紧跟着0453.,不会消耗任何字符。因为每个产品都以0453.开头,所以用这个断言分割,能精准拆分出每个完整产品,完全不会误分割产品内部的逗号或其他内容。

额外优化点

  1. 修复txt_save函数的重复写入问题:
def txt_save():
    numberOfpages = get_number_of_pages()
    parts.clear()  # 清空历史内容
    for i in range(1, numberOfpages):
        page = pdf_reader.pages[i]
        page.extract_text(visitor_text=visitor_body)
    text_body = "".join(parts)
    with open("catalogo.txt", mode='w', encoding='utf-8') as file:
        file.write(text_body)
  1. 简化remove_line_break函数,避免重复打开文件:
def remove_line_break():
    with open("catalogo.txt", mode="r", encoding="utf-8") as file, \
         open("catalogo_no_linebreak.txt", mode='w', encoding='utf-8') as arq:
        arq.write(file.read().replace('\n', ''))

内容的提问来源于stack exchange,提问作者Lucas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 11:17:11