如何从BeautifulSoup4爬取结果中仅提取并打印数值内容
你当前输出带多余符号有两个原因:
contents属性返回的是子节点列表,直接打印会自带列表的方括号、字符串包裹引号- 爬取的目标债券利率本身自带百分号后缀
你可以直接修改末尾的取值和打印逻辑即可,修改后的完整代码如下:
import requests from bs4 import BeautifulSoup import re urllong = "http://www.worldgovernmentbonds.com/country/russia/" pagelong = requests.get(urllong, timeout=5) souplong = BeautifulSoup(pagelong.content, "html.parser") resultslong = souplong.find(id="page") series_elementlong = resultslong.find("div", class_="post-content box mark-links entry-content") longratelong = series_elementlong.find(class_="w3-center w3-sand") # 直接提取标签内的纯文本,去除首尾空白 longrate_text = longratelong.find("b").text.strip() # 过滤百分号后转成浮点型数值 longrate = float(longrate_text.replace("%", "")) print(longrate)
如果返回的文本里还包含引号、特殊符号等多余字符,可以用正则表达式统一过滤所有非数字和小数点的字符,替换对应取值逻辑即可:
# 匹配所有不是数字、小数点的字符,替换为空 longrate_clean = re.sub(r"[^0-9.]", "", longrate_text) longrate = float(longrate_clean)
内容的提问来源于stack exchange,提问作者JamieC113
相关产品推荐
相关产品推荐

