如何从BeautifulSoup解析得到的字符串中提取指定字段信息?
问题
需要从以下HTML代码中提取名称、价格、位置和最后发帖时间字段:
<div class="topicicons"><span title="This is a marketplace ad topic." class="icon icon-tag"></span></div> <a data-nologvisit href="/pinball/forum/forum/games-for-sale" rel="7" class="subforum subforum-7" title="Pinball machines for sale">MFS</a> <a class="t" href="/pinball/forum/topic/for-sale-pirates-of-the-caribbean-le-58">FS: Pirates of the Caribbean (LE)<span class="tag tag-price">$ 25,000 </span><span class="tag tag-loc">Whiteland, IN</span></a> <span class="by">By ARW55 (1 year ago)<span class="last"> - Last post 3 days ago</span></span> </div><div rel="319235" data-vu="" class="topic topic-mb0 sf-7 has-new sfbox-1 topic-featured">
当前使用代码:
soup.findAll(True, {'class': ['t', 'by']})
得到输出:
FS: Pirates of the Caribbean (LE)$ 25,000 Whiteland, IN By ARW55 (1 year ago) - Last post 3 days ago
不清楚如何从这些字符串中提取所需信息,且存在大量类似格式的条目(示例):
FS: Teenage Mutant Ninja Turtles (Pro)$ 8,000 (OBO) Downers Grove, IL By Thorn-in-pinball (3 days ago) - Last post 3 days ago
解决方案
直接利用HTML结构中已有的类名定位元素,比拆分字符串更稳定:
提取名称:定位class为
t的a标签,取其第一个子节点文本,去掉前缀FS:并清理空格:title_a = soup.find('a', class_='t') name = title_a.contents[0].replace('FS: ', '').strip() # 示例输出:Pirates of the Caribbean (LE)提取价格:直接定位class为
tag-price的span标签,提取文本并清理多余空格:price = soup.find('span', class_='tag-price').get_text(strip=True) # 示例输出:$25,000提取位置:定位class为
tag-loc的span标签,提取文本:location = soup.find('span', class_='tag-loc').get_text(strip=True) # 示例输出:Whiteland, IN提取最后发帖时间:定位class为
last的span标签,去掉开头的-前缀:last_post_time = soup.find('span', class_='last').get_text(strip=True).replace('- ', '') # 示例输出:Last post 3 days ago
如果要批量处理多个条目,先定位每个topic类的div容器,再在每个容器内执行上述提取操作:
for topic in soup.find_all('div', class_='topic'): # 提取名称 title_a = topic.find('a', class_='t') name = title_a.contents[0].replace('FS: ', '').strip() # 提取价格 price = topic.find('span', class_='tag-price').get_text(strip=True) # 提取位置 location = topic.find('span', class_='tag-loc').get_text(strip=True) # 提取最后发帖时间 last_post_time = topic.find('span', class_='last').get_text(strip=True).replace('- ', '') # 输出或存储结果 print(f"名称:{name},价格:{price},位置:{location},最后发帖时间:{last_post_time}")
内容的提问来源于stack exchange,提问作者Fopoki
相关产品推荐
相关产品推荐

