You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫无法提取网页script标签中的邮箱链接,求助

解决从script标签中提取邮箱地址的问题

我来帮你搞定这个爬虫提取邮箱的问题!你遇到的TypeError: 'NoneType' object is not subscriptable错误,本质原因是页面里的邮箱并没有以你预期的mailto:链接形式出现在DOM结构中——它藏在页面的<script>标签内部的JavaScript代码里,所以你用soup.select_one("dd a[href^='mailto:']")根本找不到对应的元素,自然返回None,也就没法取['href']了。

下面给你两种可行的解决方案,你可以根据页面实际情况选择:

方案1:用正则表达式直接匹配邮箱格式

这种方法通用性强,只要script里有符合邮箱格式的字符串,就能提取出来:

import requests
import re
from bs4 import BeautifulSoup

url = "http://www.the-ida.com/index.php?option=com_community&view=profile&userid=31033393"
res = requests.get(url)
soup = BeautifulSoup(res.text, "lxml")

# 获取页面所有script标签
script_tags = soup.find_all('script')
# 定义匹配邮箱的正则表达式(基础版,可根据需求调整)
email_regex = r'[\w\.-]+@[\w\.-]+\.[a-zA-Z]{2,}'

for script in script_tags:
    script_content = script.string
    if script_content:  # 跳过空的script标签
        # 查找所有符合格式的邮箱
        matched_emails = re.findall(email_regex, script_content)
        if matched_emails:
            # 去重后打印结果
            unique_emails = list(set(matched_emails))
            print("提取到的邮箱地址:")
            for email in unique_emails:
                print(email)
            break  # 找到目标后退出循环,避免无效遍历

方案2:解析script中的结构化JS对象(更精准)

如果页面的script标签里是结构化的用户数据(比如类似var userInfo = {"email": "xxx@xxx.com", ...}这样的格式),可以直接解析这个对象,避免正则的误匹配:

import requests
import re
import json
from bs4 import BeautifulSoup

url = "http://www.the-ida.com/index.php?option=com_community&view=profile&userid=31033393"
res = requests.get(url)
soup = BeautifulSoup(res.text, "lxml")

script_tags = soup.find_all('script')
for script in script_tags:
    script_content = script.string
    # 先判断当前script是否包含用户信息(这里假设关键字是"email"或特定变量名)
    if script_content and 'email' in script_content:
        # 提取JS对象的JSON部分(需要根据页面实际的变量名调整正则)
        # 比如如果是var profile = {...}; 就把正则里的userInfo换成profile
        json_match = re.search(r'var userInfo = ({.*?});', script_content, re.DOTALL)
        if json_match:
            json_str = json_match.group(1)
            # 把JS对象转为Python字典
            user_data = json.loads(json_str)
            email = user_data.get('email')
            if email:
                print(f"提取到的邮箱地址:{email}")
                break

注意事项

  • 方案2需要你先查看页面的script内容,确认存储邮箱的JS变量名和结构,再调整正则表达式中的变量名(比如userInfo);
  • 如果页面有反爬机制,可能需要给requests.get()添加请求头(比如headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}),避免被拦截。

内容的提问来源于stack exchange,提问作者SIM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:07:45