如何使用Beautiful Soup提取script标签内的邮箱地址?
提取方法
你可以按照以下两步操作:先定位到对应script标签,再从标签文本中提取邮箱,有两种常用实现方案:
方案1:正则匹配(简单快捷)
直接对script标签内的文本做邮箱格式匹配,不需要解析JS结构,适合快速提取:
from bs4 import BeautifulSoup import re # 替换为你自己的HTML文本变量 html_content = "你的HTML内容" soup = BeautifulSoup(html_content, "html.parser") # 定位包含目标JS代码的script标签 target_script = soup.find("script", string=re.compile(r"atomic\(")) script_text = target_script.get_text() # 匹配邮箱 email_match = re.search(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}", script_text) if email_match: email = email_match.group() print(email) # 输出 ename@email.com
方案2:解析JS对象(更稳定)
如果结构固定,建议提取JS参数转JSON解析,避免正则误匹配:
from bs4 import BeautifulSoup import re import json html_content = "你的HTML内容" soup = BeautifulSoup(html_content, "html.parser") target_script = soup.find("script", string=re.compile(r"atomic\(")) script_text = target_script.get_text() # 提取atomic函数的入参部分 param_str = re.search(r"atomic\(([\s\S]*?)\)\.\$mount", script_text).group(1) # 修正原代码里多余的右大括号,转成标准JSON格式 param_str = param_str.strip().rstrip("}") + "}" config = json.loads(param_str) email = config["playlist"][0]["email"] print(email) # 输出 ename@email.com
注意:如果页面内有多个script标签,可以给
find方法补充更多筛选条件(比如根据script的父节点属性、id等),确保定位到正确的标签。
内容的提问来源于stack exchange,提问作者อมรา หน้าหัก
相关产品推荐
相关产品推荐

