Python提取指定JS链接时如何去除输出中的多余空行?
问题:提取指定JS链接时输出存在大量空行
问题描述
目标是得到简洁的输出:
https://widget.reviews.io/rating-snippet/dist.js
但实际输出时链接前后带有大量空行:
https://widget.reviews.io/rating-snippet/dist.js
已尝试用''.join去除列表符号,但空行问题仍未解决,以下是当前代码:
import requests import re from bs4 import BeautifulSoup html = requests.get("https://www.nutrimuscle.com") soup = BeautifulSoup(html.text, "html.parser") # Find all script tags for n in soup.find_all('script'): # Check if the src attribute exists, and if it does grab the source URL if 'src' in n.attrs: javascript = n['src'] # Otherwise assume that the javascript is contained within the tags else: javascript = '' kameleoonRegex = re.compile(r'[\w].*rating-snippet/dist.js') #Everything I tried :D kameleeonScript = kameleoonRegex.findall(javascript) text = ''.join(kameleeonScript) print(text)
解决方案
问题根源在于遍历所有script标签时,每个标签都会执行print(text),而大部分标签匹配不到目标链接,会输出空字符串,这些空字符串就表现为输出中的空行。
修改方案:
- 仅在匹配到目标链接时执行打印操作,避免无意义的空输出
- 找到目标链接后可直接退出循环,减少不必要的遍历
修改后的代码示例:
import requests import re from bs4 import BeautifulSoup html = requests.get("https://www.nutrimuscle.com") soup = BeautifulSoup(html.text, "html.parser") # 定义匹配目标链接的正则 target_regex = re.compile(r'.*rating-snippet/dist.js') # 遍历script标签查找目标 for script_tag in soup.find_all('script'): if 'src' in script_tag.attrs: src_url = script_tag['src'] match_result = target_regex.search(src_url) if match_result: print(match_result.group()) # 找到目标后直接终止循环 break
执行上述代码即可得到无多余空行的目标链接输出。
内容的提问来源于stack exchange,提问作者Pryapus
相关产品推荐
相关产品推荐

