Swift:PDF生成的NSAttributedString可否解析为字典并提取H1对应内容
问题根因
你调用attributes(at:longestEffectiveRange:in:)时仅查询了字符串起始位置0的属性,且未传入有效指针接收属性的生效范围,因此只会返回当前位置的1组属性,无法遍历全文档的所有属性分段。
实现思路
- 首先固定H1的识别特征:根据你排查的结果,原HTML H1转换后对应34pt的Helvetica字体,以及RGB值为(0.94118, 0.32549, 0.29804)的文字颜色,忽略每次运行会变动的内存指针字段即可。
- 遍历全文档所有属性分段,使用
NSAttributedString自带的enumerateAttributes(in:options:using:)方法,自动遍历所有属性不同的文本段,匹配符合H1特征的段落:
var h1PositionList: [(title: String, range: NSRange)] = [] let fullTextRange = NSMakeRange(0, strPDF.length) strPDF.enumerateAttributes(in: fullTextRange, options: []) { attributes, range, stop in guard let font = attributes[.font] as? NSFont, let color = attributes[.color] as? NSColor else { return } // 特征匹配,可根据实际转换误差调整精度 let isH1Font = font.fontName.contains("Helvetica") && abs(font.pointSize - 34) < 0.5 let isH1Color = abs(color.redComponent - 0.94118) < 0.01 && abs(color.greenComponent - 0.32549) < 0.01 && abs(color.blueComponent - 0.29804) < 0.01 if isH1Font && isH1Color { let h1Title = strPDF.attributedSubstring(from: range).string h1PositionList.append((title: h1Title, range: range)) } }
- 按H1位置拆分模块,提取相邻两个H1之间的文本作为独立模块内容:
var moduleList: [(title: String, content: NSAttributedString)] = [] // 补充文档末尾作为最后一个模块的结束边界 h1PositionList.append((title: "end", range: NSMakeRange(strPDF.length, 0))) for index in 0 ..< h1PositionList.count - 1 { let currentH1 = h1PositionList[index] let nextH1 = h1PositionList[index + 1] // 模块内容范围:当前H1结束位置 到 下一个H1开始位置 let moduleRange = NSMakeRange(currentH1.range.upperBound, nextH1.range.location - currentH1.range.upperBound) let moduleContent = strPDF.attributedSubstring(from: moduleRange) moduleList.append((title: currentH1.title, content: moduleContent)) }
- 可选优化:如果PDF存在多页截断、H1跨页的情况,可以先按页拆分PDF再逐页识别H1,避免大范围遍历的误差。
内容的提问来源于stack exchange,提问作者PruitIgoe
相关产品推荐
相关产品推荐

