Scrapy中日期格式匹配错误排查:参照文档仍报错
解决Scrapy爬虫日期格式不匹配错误
问题分析
你处理后的日期字符串是Sept 23 2022,但Python的datetime.strptime默认使用的locale可能不识别Sept这种4位的月份缩写(%b格式符通常匹配3位缩写,比如Sep),这就导致了格式不匹配的报错。另外,代码中未处理response.css(...).get()返回None的情况,也可能引发额外的属性错误。
修复方案
方案1:统一月份缩写为3位
在字符串处理阶段直接把Sept替换成Sep,适配%b的匹配规则:
from datetime import datetime raw_date = response.css('span[class="published-date-day"] ::text').get() if raw_date: Published_Date = raw_date.strip().replace(",", "").replace(".", "").replace("Sept", "Sep") Published_Date = datetime.strptime(Published_Date, "%b %d %Y").date() else: Published_Date = None # 处理日期不存在的边界情况
方案2:设置英文locale识别完整缩写
通过设置locale为英文,让%b能识别Sept这类4位月份缩写:
import locale from datetime import datetime # 设置英文locale(需确保系统支持对应配置,不同系统可能有差异) try: locale.setlocale(locale.LC_TIME, 'en_US.UTF-8') except locale.Error: # 备用英文locale选项 locale.setlocale(locale.LC_TIME, 'English_US') raw_date = response.css('span[class="published-date-day"] ::text').get() if raw_date: Published_Date = raw_date.strip().replace(",", "").replace(".", "") Published_Date = datetime.strptime(Published_Date, "%b %d %Y").date() else: Published_Date = None
注意事项
- 必须处理
get()返回None的情况,避免调用strip()时触发AttributeError。 - 不同系统支持的locale配置不同,方案2中如果
en_US.UTF-8无效,可以尝试en_GB.UTF-8等其他英文locale。
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

