You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除爬虫输出中的多余空格及rebtel.bootstrappedData代码段?

解决方案

可以通过字符串分割、正则替换实现你的需求,修改后的代码如下:

import json
import requests
import re
from bs4 import BeautifulSoup


url = "https://www.rebtel.com/en/rates/"

r = requests.get(url)

soup = BeautifulSoup(r.content, "html.parser")

script = soup.find_all("script")[4].text.strip()[38:]

# 1. 移除rebtel.bootstrappedData代码段
script_cleaned = script.split("; rebtel.bootstrappedData=")[0]

# 2. 移除多余空白字符(换行、制表符、连续空格)
script_cleaned = re.sub(r'\s+', ' ', script_cleaned).strip()

print(script_cleaned)

关键处理说明:

  • 移除指定代码段:用split("; rebtel.bootstrappedData=")从目标代码段起始位置拆分字符串,取第一个元素即可彻底去掉rebtel.bootstrappedData及后续内容。
  • 清理多余空格:通过正则表达式re.sub(r'\s+', ' ', script_cleaned)把所有连续空白字符(换行、制表符、多空格)替换为单个空格,再用strip()去除首尾空格,让输出更整洁。

如果需要将处理后的内容解析为JSON(原内容是JSON结构片段),可添加如下可选步骤:

# 可选:解析为格式化的JSON对象
try:
    data = json.loads(script_cleaned)
    print(json.dumps(data, indent=2))
except json.JSONDecodeError as e:
    print(f"JSON解析错误: {e}")

内容的提问来源于stack exchange,提问作者Demi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 06:40:51