如何在Python中不受机器locale影响地稳健解析时间戳?
稳健解析时间戳的方案
我偏好使用Python标准库的datetime模块解析时间戳,但部分格式码依赖运行机器的区域设置,导致代码存在脆弱性:
from datetime import datetime string = "08 Aug 2024 01:45 PM" fmt = "%d %b %Y %I:%M %p" datetime.strptime(string, fmt) # 在美国电脑上运行 # datetime(2024, 8, 8, 13, 45) datetime.strptime(string, fmt) # 在法国电脑上运行 # ValueError: time data '08 Aug 2024 01:45 PM' does not match format '%d %b %Y %I:%M %p'
注:对应的法语格式为08 aoû 2024 01:45 (月份名称不同且无AM/PM标识)
是否存在更稳健的系统化时间戳解析方案?
我目前用上下文管理器临时设置locale来控制解析流程,但觉得应该有更适配的工具来完成这项工作:
from contextlib import contextmanager from datetime import datetime @contextmanager def locale_context(locale: str): import locale as l saved_loc = l.getlocale() l.setlocale(l.LC_ALL, locale) try: yield l.getlocale() finally: l.setlocale(l.LC_ALL, saved_loc) def strptime(x: str, fmt: str, *, locale: str) -> datetime: """等价于标准strptime,但可指定locale""" with locale_context(locale): return datetime.strptime(x, fmt) strptime(string, fmt, locale="en_US") # 在美国电脑上运行 # datetime(2024, 8, 8, 13, 45) strptime(string, fmt, locale="en_US") # 在法国电脑上运行 # datetime(2024, 8, 8, 13, 45)
更适配的第三方工具方案
如果不想自己维护locale控制逻辑,推荐使用成熟的第三方库:
- python-dateutil:
parser.parse方法能自动识别多种时间格式,支持通过自定义parserinfo指定本地化规则:from dateutil import parser from dateutil.parser import parserinfo class EnglishParserInfo(parserinfo): MONTHS = [("Jan", "January"), ("Feb", "February"), ("Mar", "March"), ("Apr", "April"), ("May", "May"), ("Jun", "June"), ("Jul", "July"), ("Aug", "August"), ("Sep", "September"), ("Oct", "October"), ("Nov", "November"), ("Dec", "December")] AMPM = [("AM", "am"), ("PM", "pm")] string = "08 Aug 2024 01:45 PM" dt = parser.parse(string, parserinfo=EnglishParserInfo()) # datetime.datetime(2024, 8, 8, 13, 45) - arrow:现代化日期时间库,内置多locale支持,解析时直接指定locale即可:
import arrow string = "08 Aug 2024 01:45 PM" dt = arrow.get(string, "DD MMM YYYY hh:mm A", locale="en_US").datetime # datetime.datetime(2024, 8, 8, 13, 45, tzinfo=datetime.timezone.utc)
内容的提问来源于stack exchange,提问作者Thrastylon
相关产品推荐
相关产品推荐

