使用BeautifulSoup提取HTML中datetime返回None的原因及解决方法
问题原因及解决方法
为什么返回None?
你的代码里用了soup.find("time", class_="datetime"),但看你提供的HTML片段,<time>标签的class是author-details__timestamp formatTimeStampEs6,根本没有datetime这个类,所以BeautifulSoup找不到匹配的元素,自然返回None。另外,你已经从Chrome DevTools复制了正确的CSS选择器,但代码里完全没用到这个选择器。
解决方法
方法1:使用你复制的CSS选择器
直接用BeautifulSoup的select_one()方法传入复制的选择器,精准定位元素:
def extract_time(data): """Extract the time from the HTML of the article page.""" soup = BeautifulSoup(data, 'html.parser') # 使用复制的CSS选择器定位time元素 time_element = soup.select_one('#article > div.mar-article > div > div.mar-article__timestamp > time') if time_element: # 提取datetime属性值 datetime_value = time_element.get('datetime') print(datetime_value) return datetime_value return None
方法2:通过正确的class属性定位
如果不想用CSS选择器,也可以用<time>标签的实际class来查找(多个class选其中一个即可,比如author-details__timestamp):
def extract_time(data): """Extract the time from the HTML of the article page.""" soup = BeautifulSoup(data, 'html.parser') # 使用正确的class查找time元素 time_element = soup.find("time", class_="author-details__timestamp") if time_element: # 提取datetime属性 datetime_value = time_element['datetime'] print(datetime_value) return datetime_value return None
关键注意点
- 提取属性值可以用
element.get('属性名')或者element['属性名'],前者在属性不存在时返回None,后者会抛出KeyError,根据需求选择即可。 - 查找元素时,一定要确保条件和HTML中的实际属性完全匹配,不要凭猜测写属性值。
内容的提问来源于stack exchange,提问作者elksie5000
相关产品推荐
相关产品推荐

