BeautifulSoup爬取类匹配异常及NoneType错误求助
解决TechPowerUp爬虫的两个问题
1. 排除带review类的文章
直接用CSS选择器的:not()伪类可精准筛选仅含newspost类的文章,避免混入newspost review的内容:
articles = soup.select("article.newspost:not(.review)")
也可以用find_all结合lambda判断类名完全匹配:
articles = soup.find_all("article", class_="newspost", attrs={"class": lambda cls: cls == "newspost"})
2. 修复datetime的NoneType错误并优化日期格式
出现NoneType错误是因为部分文章的time标签不存在或无datetime属性,必须先做存在性判断,再处理日期格式:
from datetime import datetime for article in articles: time_tag = article.select_one("time") if time_tag and "datetime" in time_tag.attrs: raw_datetime = time_tag.get("datetime") # 转换为更易读的格式,比如YYYY-MM-DD HH:MM formatted_date = datetime.fromisoformat(raw_datetime).strftime("%Y-%m-%d %H:%M") print(formatted_date) else: # 处理无日期的情况,可设默认值 print("无有效日期")
内容的提问来源于stack exchange,提问作者jamiek
相关产品推荐
相关产品推荐

