Python 2.7如何正确拆分含空格的文件路径字符串
问题描述
现有存储路径信息的文本文件,无空格路径的行初始格式如下:
/home/Plugins/file1 e:222 k:dir (327/1) /home/Plugins/file2 e:100 k:dir (326/1)
初始编写的提取代码在路径无空格时可正常运行:
with open('output_file.txt', 'r') as output_file: for line in output_file: file_path = line.split()[0] eId = line.split()[1].split(":")[1] logging.info("file path:"+file_path) logging.info("eId:"+eId)
实际生产环境中文件夹、文件名常包含空格,文件中会出现如下格式的行:
/home/tools/AMS Provider/file3.txt e:224 k:dir (127/1) /home/account validator e:227 k:dir (247/1)
上述示例中AMS Provider是带空格的子文件夹名,account validator是路径末尾带空格的文件名,直接使用默认split()按空白字符拆分,会将完整路径拆分为多个片段导致脚本运行失败,且当前服务器环境仅支持Python 2.7,需要实现准确提取带空格完整文件路径与对应eId的逻辑。
解决方案
核心判断依据:每行结构固定为「文件路径 + 空格分隔的固定属性字段」,第一个属性字段永远是e:数字格式的eId,且合法Linux路径中不可能出现「空格+e:+数字」的片段,因此只要定位到第一个 e:的位置,就能准确切分路径和后续属性,两种完全兼容Python2.7的实现方式如下:
方案1:正则匹配(代码最简洁,推荐)
直接通过正则非贪婪匹配定位路径和eId,不需要多次字符串拆分:
import re import logging # 预编译正则提升遍历性能 path_pattern = re.compile(r'^(.*?) e:(\d+) ') with open('output_file.txt', 'r') as output_file: for line in output_file: line = line.strip() match_res = path_pattern.match(line) if not match_res: logging.warning("invalid line format: %s" % line) continue file_path = match_res.group(1) eId = match_res.group(2) logging.info("file path:" + file_path) logging.info("eId:" + eId)
正则规则说明:
^匹配行首,(.*?)做非贪婪匹配,会匹配到第一个e:片段就停止,刚好捕获完整路径;(\d+)直接捕获e:后的数字作为eId,不需要额外分割字符串。
方案2:字符串定位拆分(无正则依赖,遍历性能更高)
如果不想引入正则模块,可以直接通过字符串查找定位切分点,逻辑直观:
import logging with open('output_file.txt', 'r') as output_file: for line in output_file: line = line.strip() e_flag_pos = line.find(' e:') if e_flag_pos == -1: logging.warning("invalid line format: %s" % line) continue # 切分点前的所有内容就是完整路径 file_path = line[:e_flag_pos] # 从切分点后提取eId e_field = line[e_flag_pos+1:].split()[0] eId = e_field.split(':')[1] logging.info("file path:" + file_path) logging.info("eId:" + eId)
- 两种方案都同时兼容无空格路径、路径任意位置带空格的场景,不需要额外写分支判断
- 代码增加了异常格式行的容错处理,遇到不符合规则的行会打印告警日志,不会直接中断整个文件遍历
- 所有语法完全兼容Python 2.7版本,不需要安装额外第三方依赖
内容的提问来源于stack exchange,提问作者vel
相关产品推荐
相关产品推荐

