You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于行高自动量化分配文本类别,实现PDF转Markdown结构化处理?

基于行高程序化分配文本样式的启发式方案

核心思路

利用行高的聚类分析自适应区分不同层级的文本样式——由于不同PDF的字号、行高范围差异极大,聚类算法能自动识别行高的自然分组,替代手动阈值适配。

具体实现步骤

1. 预处理行高数据

先过滤行高中的异常值(比如远小于最小正文行高或远大于最大标题行高的数据),避免干扰聚类结果,可采用四分位距法(IQR)筛选:

import numpy as np

def filter_outliers(heights):
    q1 = np.percentile(heights, 25)
    q3 = np.percentile(heights, 75)
    iqr = q3 - q1
    lower_bound = q1 - 1.5 * iqr
    upper_bound = q3 + 1.5 * iqr
    return [h for h in heights if lower_bound <= h <= upper_bound]

2. K-Means聚类分组

因为需要固定映射为5类(对应H1到说明文本的1-5),直接用K-Means将行高分为5组,再按组内行高从大到小映射到对应样式类别:

from sklearn.cluster import KMeans

def cluster_heights(heights, n_clusters=5):
    height_array = np.array(heights).reshape(-1, 1)
    kmeans = KMeans(n_clusters=n_clusters, random_state=42)
    kmeans.fit(height_array)
    
    # 按聚类中心从大到小排序,映射为1-5的样式类别
    sorted_centers = sorted(kmeans.cluster_centers_.flatten(), reverse=True)
    height_to_class = {}
    for idx, center in enumerate(sorted_centers):
        label = np.where(kmeans.cluster_centers_.flatten() == center)[0][0]
        height_to_class[label] = idx + 1
    
    labels = kmeans.predict(height_array)
    return [height_to_class[label] for label in labels]

3. 整合转换流程

将聚类结果与原文本行配对,生成目标格式list[tuple(str, int)]:

def assign_styles(text_height_pairs):
    texts, heights = zip(*text_height_pairs)
    filtered_heights = filter_outliers(heights)
    
    # 若过滤后数据量不足5,跳过过滤避免聚类失败
    if len(filtered_heights) < 5:
        filtered_heights = heights
    
    style_classes = cluster_heights(filtered_heights)
    return list(zip(texts, style_classes))

4. 优化与适配

  • 若聚类结果出现不合理分组(如正文与说明文本混组),可改用层次聚类(AgglomerativeClustering),调整链接方式为ward优化分组效果;
  • 多页PDF建议基于全局行高聚类,避免跨页样式不一致;
  • PyMuPDF中可选择用textline.font_size替代行高(y1-y0),部分PDF的行高包含额外行间距,字号更能准确反映文本层级。

内容的提问来源于stack exchange,提问作者janekb04

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 12:17:30