使用pdfparser解析PDF后,date_parse提取特定字符串后首个日期年份异常的解决咨询
解决提取特定字符串后首个日期的问题
看来你在处理PDF文本的日期提取时碰到了年份解析乱跳的麻烦,我给你几个实用的方案,针对你给出的示例场景应该都能完美解决:
方案一:用正则表达式精准匹配首个日期
既然你已经能定位到目标搜索字符串,那我们可以先截取该字符串之后的文本,再用正则精准匹配符合格式的第一个日期,完全避开后面其他日期或数字的干扰。
针对你示例里的March 4, 2004这种「英文月份+日期+逗号+四位年份」的格式,我们可以用这个正则规则:/([A-Z][a-z]+) (\d{1,2}), (\d{4})/。
给你写个PHP代码示例:
// 假设这是你从PDF中提取到的完整文本 $fullText = "Satisfaction of the mortgage from Karen Ann Lewis,a single woman to Bank of America, N.A. recorded March 4, 2004 and another date May 10, 2023"; // 你预先定位好的搜索字符串 $searchString = "Satisfaction of the mortgage from Karen Ann Lewis,a single woman to Bank of America, N.A. recorded"; // 截取搜索字符串之后的文本内容 $textAfterSearch = substr($fullText, strpos($fullText, $searchString) + strlen($searchString)); // 匹配首个符合格式的日期 $dateRegex = '/([A-Z][a-z]+) (\d{1,2}), (\d{4})/'; if (preg_match($dateRegex, trim($textAfterSearch), $matches)) { // 提取月、日、年 $monthName = $matches[1]; $day = $matches[2]; $year = $matches[3]; // 可以转换成标准DateTime对象方便后续处理 $targetDate = DateTime::createFromFormat('F j, Y', "$monthName $day, $year"); echo $targetDate->format('Y-m-d'); // 输出:2004-03-04 }
这个方法的优势是精准可控,如果PDF里的日期格式有变化(比如数字月份3/4/2004、无逗号的March 4 2004),只需要调整正则表达式就能适配。
方案二:缩小文本范围后再用date_parse
date_parse出错是因为处理的文本太长,包含了太多干扰项。那我们可以先把搜索字符串之后的文本切分成短片段,只取前几个单词组成小字符串,再交给date_parse处理:
$fullText = "Satisfaction of the mortgage from Karen Ann Lewis,a single woman to Bank of America, N.A. recorded March 4, 2004 and another date May 10, 2023"; $searchString = "Satisfaction of the mortgage from Karen Ann Lewis,a single woman to Bank of America, N.A. recorded"; $textAfterSearch = substr($fullText, strpos($fullText, $searchString) + strlen($searchString)); // 把文本按空格分割,取前5个单词组成短字符串(足够包含完整日期) $shortText = implode(' ', array_slice(explode(' ', trim($textAfterSearch)), 0, 5)); // 现在用date_parse处理短字符串就能得到正确结果 $dateInfo = date_parse($shortText); print_r($dateInfo); // 输出里的year、month、day都会是正确的2004、3、4
推荐优先用方案一,因为正则匹配的稳定性更高,不会因为文本片段里的其他数字出现意外。
内容的提问来源于stack exchange,提问作者nOmi
相关产品推荐
相关产品推荐

