求助:如何自动垂直分割多份简历文本文件为左右区块?
批量垂直分割简历文本的优化方案
原方法依赖\s\s\S匹配连续空格+非空格字符,但简历中存在大量非分割线的连续空格(如缩进、内容间隔),导致误匹配过多,无法精准定位左右区块分割线。以下是针对性的优化方案:
核心思路
左右区块的分割线是垂直对齐的空白密集区,大部分行在该位置会出现连续空白。通过统计每列的空白占比,找到跨行一致的高占比空白区间,以此确定分割位置。
代码实现
import re from collections import defaultdict def find_best_split_position(file_path): # 统计每列的空白出现次数 col_space_counts = defaultdict(int) total_lines = 0 with open(file_path, 'r') as f: lines = [line.rstrip('\n') for line in f] for line in lines: total_lines += 1 for idx, char in enumerate(line): if char.isspace(): col_space_counts[idx] += 1 if not col_space_counts: return None # 无空白区域,无法分割 # 计算每列的空白占比 col_space_ratios = {col: cnt / total_lines for col, cnt in col_space_counts.items()} # 筛选高占比空白列,先尝试80%阈值,失败则降级到60% threshold = 0.8 candidate_cols = [col for col, ratio in col_space_ratios.items() if ratio >= threshold] if not candidate_cols: threshold = 0.6 candidate_cols = [col for col, ratio in col_space_ratios.items() if ratio >= threshold] if not candidate_cols: return None # 寻找最长的连续空白列区间 candidate_cols.sort() max_length = 1 current_length = 1 best_start = candidate_cols[0] for i in range(1, len(candidate_cols)): if candidate_cols[i] == candidate_cols[i-1] + 1: current_length += 1 if current_length > max_length: max_length = current_length best_start = candidate_cols[i - current_length + 1] else: current_length = 1 # 取连续区间的中间位置作为分割点 split_pos = best_start + max_length // 2 return split_pos def split_resume_file(file_path): split_pos = find_best_split_position(file_path) if split_pos is None: print(f"无法找到合适的分割位置: {file_path}") return None, None left_block = [] right_block = [] with open(file_path, 'r') as f: for line in f: line = line.rstrip('\n') # 去除区块边缘的多余空白,保证内容整洁 left = line[:split_pos].rstrip() right = line[split_pos:].lstrip() left_block.append(left) right_block.append(right) return '\n'.join(left_block), '\n'.join(right_block) # 使用示例 left, right = split_resume_file('name.txt') if left and right: print("左侧区块:") print(left) print("\n右侧区块:") print(right)
关键说明
- 空白占比统计:通过跨行统计每列的空白出现比例,精准定位垂直分割线的核心区域,避免单个行的偶然匹配干扰。
- 连续区间筛选:取最长的连续高占比空白列区间,确保分割线是垂直对齐的,而非孤立的空白列。
- 阈值可调:根据简历的实际格式,可调整空白占比阈值(比如部分简历分割线空白占比不足80%,可调低至0.6)。
- 内容清洗:分割后自动去除区块边缘的多余空白,直接得到整洁的左右内容。
内容的提问来源于stack exchange,提问作者Nadjib Rahmani
相关产品推荐
相关产品推荐

