You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取指定JS链接时如何去除输出中的多余空行?

问题:提取指定JS链接时输出存在大量空行

问题描述

目标是得到简洁的输出:

https://widget.reviews.io/rating-snippet/dist.js

但实际输出时链接前后带有大量空行:

https://widget.reviews.io/rating-snippet/dist.js



已尝试用''.join去除列表符号,但空行问题仍未解决,以下是当前代码:

import requests
import re

from bs4 import BeautifulSoup

html = requests.get("https://www.nutrimuscle.com")

soup = BeautifulSoup(html.text, "html.parser")

# Find all script tags
for n in soup.find_all('script'):

    # Check if the src attribute exists, and if it does grab the source URL
    if 'src' in n.attrs:
        javascript = n['src']

    # Otherwise assume that the javascript is contained within the tags
    else:
        javascript = ''


    kameleoonRegex = re.compile(r'[\w].*rating-snippet/dist.js')
    #Everything I tried :D
    kameleeonScript = kameleoonRegex.findall(javascript)
    text = ''.join(kameleeonScript)
    print(text)

解决方案

问题根源在于遍历所有script标签时,每个标签都会执行print(text),而大部分标签匹配不到目标链接,会输出空字符串,这些空字符串就表现为输出中的空行。

修改方案:

  • 仅在匹配到目标链接时执行打印操作,避免无意义的空输出
  • 找到目标链接后可直接退出循环,减少不必要的遍历

修改后的代码示例:

import requests
import re

from bs4 import BeautifulSoup

html = requests.get("https://www.nutrimuscle.com")

soup = BeautifulSoup(html.text, "html.parser")

# 定义匹配目标链接的正则
target_regex = re.compile(r'.*rating-snippet/dist.js')

# 遍历script标签查找目标
for script_tag in soup.find_all('script'):
    if 'src' in script_tag.attrs:
        src_url = script_tag['src']
        match_result = target_regex.search(src_url)
        if match_result:
            print(match_result.group())
            # 找到目标后直接终止循环
            break

执行上述代码即可得到无多余空行的目标链接输出。

内容的提问来源于stack exchange,提问作者Pryapus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 00:09:20