Python正则表达式拆分产品记录异常,求完整提取方法
问题解决:提取单行产品列表中的完整产品记录
你当前用re.split(r',00')分割单行产品列表时,出现记录断行、不完整的情况,核心原因包括:
- 产品内容中存在类似
4,40G的逗号,会被误分割; - 部分产品结尾没有
,00标识(比如示例中第一个产品仅以重量结尾),导致分割后记录缺失; - 分割后未将
,00补回原记录,丢失价格结尾标识。
解决方案
通过正则正向先行断言定位产品的起始分隔点(每个0453.开头的位置,除了第一个),精准分割每个完整产品。同时读取整个文件内容确保无遗漏,处理所有产品记录。
修改后的txt_format函数
def txt_format(): # 读取整个单行文件内容 with open('catalogo_no_linebreak.txt', mode='r', encoding='utf-8') as arq: full_content = arq.read() # 用正向先行断言分割:匹配"0453."前面的位置(排除字符串开头) products = re.split(r'(?=0453\.)', full_content) # 过滤空字符串,清理无效内容 products = [p.strip() for p in products if p.strip()] # 写入结果文件 with open('saida.txt', mode='w', encoding='utf-8') as out_arq: for product in products: out_arq.write(product + '\n')
正则逻辑说明
(?=0453\.)是正向先行断言,它仅匹配一个位置——这个位置后面紧跟着0453.,不会消耗任何字符。因为每个产品都以0453.开头,所以用这个断言分割,能精准拆分出每个完整产品,完全不会误分割产品内部的逗号或其他内容。
额外优化点
- 修复
txt_save函数的重复写入问题:
def txt_save(): numberOfpages = get_number_of_pages() parts.clear() # 清空历史内容 for i in range(1, numberOfpages): page = pdf_reader.pages[i] page.extract_text(visitor_text=visitor_body) text_body = "".join(parts) with open("catalogo.txt", mode='w', encoding='utf-8') as file: file.write(text_body)
- 简化
remove_line_break函数,避免重复打开文件:
def remove_line_break(): with open("catalogo.txt", mode="r", encoding="utf-8") as file, \ open("catalogo_no_linebreak.txt", mode='w', encoding='utf-8') as arq: arq.write(file.read().replace('\n', ''))
内容的提问来源于stack exchange,提问作者Lucas
相关产品推荐
相关产品推荐

