Python Strptime未提取日期却返回整个文本文件的技术问题求助
解决文本日期提取问题
嘿,我懂你现在的烦恼——本来想从文本里抠出日期,结果只拿到了清理掉换行和特殊字符的全文对吧?咱们一步步来搞定这个!
首先先确认下你精简后的目标文本:
03/01/2018 L205: On-site 7:00 AM, no crew on-site On-site 11:30 AM crew has excavated for the vessel pad and is watering rock for placement and compaction. Excavation is measured out and meets the requirement...
问题出在你当前的代码可能只做了文本清理操作,没有针对日期格式做精准匹配提取。针对这种MM/DD/YYYY格式的日期,用正则表达式是最直接高效的方法。
下面是Python的示例代码,你可以直接参考调试:
import re # 模拟读取的精简文本内容 text = "03/01/2018 L205: On-site 7:00 AM, no crew on-site On-site 11:30 AM crew has excavated for the vessel pad and is watering rock for placement and compaction. Excavation is measured out and meets the requirement..." # 匹配MM/DD/YYYY格式的日期正则表达式 date_pattern = r'\d{2}/\d{2}/\d{4}' # 查找第一个匹配的日期(适合单日期场景) match_result = re.search(date_pattern, text) if match_result: extracted_date = match_result.group() print(f"提取到的日期:{extracted_date}") else: print("未找到符合格式的日期") # 如果后续文本有多个日期,用findall获取所有匹配结果 # all_matched_dates = re.findall(date_pattern, text) # print(f"所有提取到的日期:{all_matched_dates}")
代码说明:
- 正则
\d{2}/\d{2}/\d{4}专门定位「两位数字/两位数字/四位数字」的日期格式,完美匹配你的示例里的03/01/2018。 re.search()会返回第一个匹配项,刚好适配你这个单日期的场景;如果之后遇到多日期文本,换成re.findall()就能一次性拿到所有结果。- 加了判断逻辑,避免没有匹配到日期时出现报错。
运行这段代码后,你就能得到想要的03/01/2018,而不是整个文本啦!如果之后遇到其他日期格式(比如DD/MM/YYYY、YYYY-MM-DD),只需要调整正则表达式就能适配,比如匹配YYYY-MM-DD可以用r'\d{4}-\d{2}-\d{2}'。
内容的提问来源于stack exchange,提问作者Rouxgrr
相关产品推荐
相关产品推荐

