如何通过BeautifulSoup提取页面中window.location.href的链接
提取window.location.href对应的链接
要从解析后的soup中提取window.location.href对应的链接,可按以下步骤操作:
- 筛选目标script标签:先获取页面中所有
<script>标签,找出包含window.location.href代码的标签内容。 - 正则匹配提取链接:用正则表达式从script文本中精准捕获赋值给
window.location.href的URL字符串。
具体代码实现:
import re from bs4 import BeautifulSoup # 假设soup为已解析完成的BeautifulSoup对象 script_tags = soup.find_all('script', type='text/javascript') # 遍历script标签定位目标内容 for script in script_tags: # 跳过无文本内容的外部引用script(如jquery.js) if not script.string: continue if 'window.location.href' in script.string: # 正则匹配提取引号内的URL url_match = re.search(r'window\.location\.href\s*=\s*["\'](.*?)["\'];', script.string) if url_match: target_link = url_match.group(1) print("提取到的链接:", target_link) break
代码说明:
soup.find_all('script', type='text/javascript'):获取页面内所有JavaScript类型的script标签。- 正则表达式
r'window\.location\.href\s*=\s*["\'](.*?)["\'];':适配window.location.href = "xxx";或window.location.href = 'xxx';的格式,捕获引号内的完整URL。 url_match.group(1):取出正则表达式捕获组中的URL字符串。
内容的提问来源于stack exchange,提问作者user3595231
相关产品推荐
相关产品推荐

