You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python的BeautifulSoup中解码document.write编码的字符串

如何解码网页中的Base64编码邮箱地址

我卡在这个问题上好几个小时了,找不到相关文档和解决方案。从https://idhsaa.org/directory开始尝试,不管是主页面还是点击学校名称打开的独立页面,都拿不到邮箱ID。

发现网页里的编码格式是这样的:

Email

我提取到的编码代码格式如下:

mailto:

我的问题是:怎么解码这些内容拿到真实邮箱ID?从输出看必须解码才能得到真实邮箱。

以下是我正在编写的代码:

import requests
from bs4 import BeautifulSoup

def url_parser(url):
    headers = {
        "User-Agent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Safari/537.36 Edge/12.246',
    }
    html_doc = requests.get(url, headers=headers).text
    soup = BeautifulSoup(html_doc, 'html.parser')
    return soup


def data_fetch(url):
    soup = url_parser(url)
    table = soup.find('table').find('tbody')
    rows = table.find_all('tr')

    data = []
    for row in rows:
        school_name = row.find_all('a')

        for school in school_name:
            if 'school?' in school.get('href'):
                # 原代码此处school_web_id未定义
                school_website = url.replace('/directory', f'/{school_web_id}')

                school_site = url_parser(school_website)
                principal_email_encoded = school_site.find_all('a')
                for principal_email in principal_email_encoded:
                    email = principal_email.get('href')
                    if 'maito:<script>' in email:
                        print(email.replace('maito:<script>', '').replace(';</script>', ''))


def main():
    url = "https://idhsaa.org/directory"
    data_fetch(url)


if __name__ == "__main__":
    main()

解决方案

这些邮箱地址用Base64编码隐藏在window.atob()的参数里,window.atob()是浏览器原生的Base64解码函数,我们可以用Python实现同样的解码逻辑:

  1. 提取Base64字符串:从href内容里匹配出window.atob('xxx')中的xxx部分
  2. Base64解码:用Python的base64.b64decode()解码得到HTML片段
  3. 提取邮箱:从解码后的HTML里解析出真实的邮箱地址

修改后的完整代码:

import requests
from bs4 import BeautifulSoup
import base64
import re

def url_parser(url):
    headers = {
        "User-Agent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Safari/537.36 Edge/12.246',
    }
    html_doc = requests.get(url, headers=headers).text
    soup = BeautifulSoup(html_doc, 'html.parser')
    return soup


def decode_email(encoded_str):
    # 正则匹配提取Base64编码内容
    match = re.search(r"window\.atob\('(.*?)'\)", encoded_str)
    if not match:
        return None
    b64_str = match.group(1)
    # Base64解码并转成字符串
    decoded_html = base64.b64decode(b64_str).decode('utf-8')
    # 从解码后的HTML中提取邮箱
    soup = BeautifulSoup(decoded_html, 'html.parser')
    email_link = soup.find('a')['href']
    return email_link.replace('mailto:', '')


def data_fetch(url):
    soup = url_parser(url)
    table = soup.find('table').find('tbody')
    rows = table.find_all('tr')

    data = []
    for row in rows:
        school_name = row.find_all('a')

        for school in school_name:
            if 'school?' in school.get('href'):
                # 修复school_web_id未定义问题,从链接参数中提取
                school_web_id = school.get('href').split('?')[1]
                school_website = url.replace('/directory', f'/{school_web_id}')

                school_site = url_parser(school_website)
                principal_email_encoded = school_site.find_all('a')
                for principal_email in principal_email_encoded:
                    email_href = principal_email.get('href')
                    if 'mailto:<script>' in email_href:
                        real_email = decode_email(email_href)
                        if real_email:
                            print(f"学校: {school.text}, 邮箱: {real_email}")
                            data.append({'school': school.text, 'email': real_email})


def main():
    url = "https://idhsaa.org/directory"
    data_fetch(url)


if __name__ == "__main__":
    main()

代码说明

  • decode_email函数专门处理编码内容的提取、解码和邮箱解析
  • 修复了原代码中school_web_id未定义的问题,通过拆分学校链接参数获取ID
  • 用正则匹配替代字符串替换,更精准提取Base64编码内容,避免格式变动导致的错误

内容的提问来源于stack exchange,提问作者theycallmepix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 00:20:51