You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何自动垂直分割多份简历文本文件为左右区块?

批量垂直分割简历文本的优化方案

原方法依赖\s\s\S匹配连续空格+非空格字符,但简历中存在大量非分割线的连续空格(如缩进、内容间隔),导致误匹配过多,无法精准定位左右区块分割线。以下是针对性的优化方案:

核心思路

左右区块的分割线是垂直对齐的空白密集区,大部分行在该位置会出现连续空白。通过统计每列的空白占比,找到跨行一致的高占比空白区间,以此确定分割位置。

代码实现

import re
from collections import defaultdict

def find_best_split_position(file_path):
    # 统计每列的空白出现次数
    col_space_counts = defaultdict(int)
    total_lines = 0

    with open(file_path, 'r') as f:
        lines = [line.rstrip('\n') for line in f]

    for line in lines:
        total_lines += 1
        for idx, char in enumerate(line):
            if char.isspace():
                col_space_counts[idx] += 1

    if not col_space_counts:
        return None  # 无空白区域,无法分割

    # 计算每列的空白占比
    col_space_ratios = {col: cnt / total_lines for col, cnt in col_space_counts.items()}

    # 筛选高占比空白列,先尝试80%阈值,失败则降级到60%
    threshold = 0.8
    candidate_cols = [col for col, ratio in col_space_ratios.items() if ratio >= threshold]
    if not candidate_cols:
        threshold = 0.6
        candidate_cols = [col for col, ratio in col_space_ratios.items() if ratio >= threshold]
        if not candidate_cols:
            return None

    # 寻找最长的连续空白列区间
    candidate_cols.sort()
    max_length = 1
    current_length = 1
    best_start = candidate_cols[0]

    for i in range(1, len(candidate_cols)):
        if candidate_cols[i] == candidate_cols[i-1] + 1:
            current_length += 1
            if current_length > max_length:
                max_length = current_length
                best_start = candidate_cols[i - current_length + 1]
        else:
            current_length = 1

    # 取连续区间的中间位置作为分割点
    split_pos = best_start + max_length // 2
    return split_pos

def split_resume_file(file_path):
    split_pos = find_best_split_position(file_path)
    if split_pos is None:
        print(f"无法找到合适的分割位置: {file_path}")
        return None, None

    left_block = []
    right_block = []
    with open(file_path, 'r') as f:
        for line in f:
            line = line.rstrip('\n')
            # 去除区块边缘的多余空白,保证内容整洁
            left = line[:split_pos].rstrip()
            right = line[split_pos:].lstrip()
            left_block.append(left)
            right_block.append(right)

    return '\n'.join(left_block), '\n'.join(right_block)

# 使用示例
left, right = split_resume_file('name.txt')
if left and right:
    print("左侧区块:")
    print(left)
    print("\n右侧区块:")
    print(right)

关键说明

  1. 空白占比统计:通过跨行统计每列的空白出现比例,精准定位垂直分割线的核心区域,避免单个行的偶然匹配干扰。
  2. 连续区间筛选:取最长的连续高占比空白列区间,确保分割线是垂直对齐的,而非孤立的空白列。
  3. 阈值可调:根据简历的实际格式,可调整空白占比阈值(比如部分简历分割线空白占比不足80%,可调低至0.6)。
  4. 内容清洗:分割后自动去除区块边缘的多余空白,直接得到整洁的左右内容。

内容的提问来源于stack exchange,提问作者Nadjib Rahmani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 08:36:13