如何在Python的BeautifulSoup中解码document.write编码的字符串
如何解码网页中的Base64编码邮箱地址
我卡在这个问题上好几个小时了,找不到相关文档和解决方案。从https://idhsaa.org/directory开始尝试,不管是主页面还是点击学校名称打开的独立页面,都拿不到邮箱ID。
发现网页里的编码格式是这样的:
我提取到的编码代码格式如下:
mailto:
我的问题是:怎么解码这些内容拿到真实邮箱ID?从输出看必须解码才能得到真实邮箱。
以下是我正在编写的代码:
import requests from bs4 import BeautifulSoup def url_parser(url): headers = { "User-Agent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Safari/537.36 Edge/12.246', } html_doc = requests.get(url, headers=headers).text soup = BeautifulSoup(html_doc, 'html.parser') return soup def data_fetch(url): soup = url_parser(url) table = soup.find('table').find('tbody') rows = table.find_all('tr') data = [] for row in rows: school_name = row.find_all('a') for school in school_name: if 'school?' in school.get('href'): # 原代码此处school_web_id未定义 school_website = url.replace('/directory', f'/{school_web_id}') school_site = url_parser(school_website) principal_email_encoded = school_site.find_all('a') for principal_email in principal_email_encoded: email = principal_email.get('href') if 'maito:<script>' in email: print(email.replace('maito:<script>', '').replace(';</script>', '')) def main(): url = "https://idhsaa.org/directory" data_fetch(url) if __name__ == "__main__": main()
解决方案
这些邮箱地址用Base64编码隐藏在window.atob()的参数里,window.atob()是浏览器原生的Base64解码函数,我们可以用Python实现同样的解码逻辑:
- 提取Base64字符串:从
href内容里匹配出window.atob('xxx')中的xxx部分 - Base64解码:用Python的
base64.b64decode()解码得到HTML片段 - 提取邮箱:从解码后的HTML里解析出真实的邮箱地址
修改后的完整代码:
import requests from bs4 import BeautifulSoup import base64 import re def url_parser(url): headers = { "User-Agent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Safari/537.36 Edge/12.246', } html_doc = requests.get(url, headers=headers).text soup = BeautifulSoup(html_doc, 'html.parser') return soup def decode_email(encoded_str): # 正则匹配提取Base64编码内容 match = re.search(r"window\.atob\('(.*?)'\)", encoded_str) if not match: return None b64_str = match.group(1) # Base64解码并转成字符串 decoded_html = base64.b64decode(b64_str).decode('utf-8') # 从解码后的HTML中提取邮箱 soup = BeautifulSoup(decoded_html, 'html.parser') email_link = soup.find('a')['href'] return email_link.replace('mailto:', '') def data_fetch(url): soup = url_parser(url) table = soup.find('table').find('tbody') rows = table.find_all('tr') data = [] for row in rows: school_name = row.find_all('a') for school in school_name: if 'school?' in school.get('href'): # 修复school_web_id未定义问题,从链接参数中提取 school_web_id = school.get('href').split('?')[1] school_website = url.replace('/directory', f'/{school_web_id}') school_site = url_parser(school_website) principal_email_encoded = school_site.find_all('a') for principal_email in principal_email_encoded: email_href = principal_email.get('href') if 'mailto:<script>' in email_href: real_email = decode_email(email_href) if real_email: print(f"学校: {school.text}, 邮箱: {real_email}") data.append({'school': school.text, 'email': real_email}) def main(): url = "https://idhsaa.org/directory" data_fetch(url) if __name__ == "__main__": main()
代码说明
decode_email函数专门处理编码内容的提取、解码和邮箱解析- 修复了原代码中
school_web_id未定义的问题,通过拆分学校链接参数获取ID - 用正则匹配替代字符串替换,更精准提取Base64编码内容,避免格式变动导致的错误
内容的提问来源于stack exchange,提问作者theycallmepix
相关产品推荐
相关产品推荐

