如何移除爬虫输出中的多余空格及rebtel.bootstrappedData代码段?
解决方案
可以通过字符串分割、正则替换实现你的需求,修改后的代码如下:
import json import requests import re from bs4 import BeautifulSoup url = "https://www.rebtel.com/en/rates/" r = requests.get(url) soup = BeautifulSoup(r.content, "html.parser") script = soup.find_all("script")[4].text.strip()[38:] # 1. 移除rebtel.bootstrappedData代码段 script_cleaned = script.split("; rebtel.bootstrappedData=")[0] # 2. 移除多余空白字符(换行、制表符、连续空格) script_cleaned = re.sub(r'\s+', ' ', script_cleaned).strip() print(script_cleaned)
关键处理说明:
- 移除指定代码段:用
split("; rebtel.bootstrappedData=")从目标代码段起始位置拆分字符串,取第一个元素即可彻底去掉rebtel.bootstrappedData及后续内容。 - 清理多余空格:通过正则表达式
re.sub(r'\s+', ' ', script_cleaned)把所有连续空白字符(换行、制表符、多空格)替换为单个空格,再用strip()去除首尾空格,让输出更整洁。
如果需要将处理后的内容解析为JSON(原内容是JSON结构片段),可添加如下可选步骤:
# 可选:解析为格式化的JSON对象 try: data = json.loads(script_cleaned) print(json.dumps(data, indent=2)) except json.JSONDecodeError as e: print(f"JSON解析错误: {e}")
内容的提问来源于stack exchange,提问作者Demi
相关产品推荐
相关产品推荐

