Python爬虫无法提取网页script标签中的邮箱链接,求助
解决从script标签中提取邮箱地址的问题
我来帮你搞定这个爬虫提取邮箱的问题!你遇到的TypeError: 'NoneType' object is not subscriptable错误,本质原因是页面里的邮箱并没有以你预期的mailto:链接形式出现在DOM结构中——它藏在页面的<script>标签内部的JavaScript代码里,所以你用soup.select_one("dd a[href^='mailto:']")根本找不到对应的元素,自然返回None,也就没法取['href']了。
下面给你两种可行的解决方案,你可以根据页面实际情况选择:
方案1:用正则表达式直接匹配邮箱格式
这种方法通用性强,只要script里有符合邮箱格式的字符串,就能提取出来:
import requests import re from bs4 import BeautifulSoup url = "http://www.the-ida.com/index.php?option=com_community&view=profile&userid=31033393" res = requests.get(url) soup = BeautifulSoup(res.text, "lxml") # 获取页面所有script标签 script_tags = soup.find_all('script') # 定义匹配邮箱的正则表达式(基础版,可根据需求调整) email_regex = r'[\w\.-]+@[\w\.-]+\.[a-zA-Z]{2,}' for script in script_tags: script_content = script.string if script_content: # 跳过空的script标签 # 查找所有符合格式的邮箱 matched_emails = re.findall(email_regex, script_content) if matched_emails: # 去重后打印结果 unique_emails = list(set(matched_emails)) print("提取到的邮箱地址:") for email in unique_emails: print(email) break # 找到目标后退出循环,避免无效遍历
方案2:解析script中的结构化JS对象(更精准)
如果页面的script标签里是结构化的用户数据(比如类似var userInfo = {"email": "xxx@xxx.com", ...}这样的格式),可以直接解析这个对象,避免正则的误匹配:
import requests import re import json from bs4 import BeautifulSoup url = "http://www.the-ida.com/index.php?option=com_community&view=profile&userid=31033393" res = requests.get(url) soup = BeautifulSoup(res.text, "lxml") script_tags = soup.find_all('script') for script in script_tags: script_content = script.string # 先判断当前script是否包含用户信息(这里假设关键字是"email"或特定变量名) if script_content and 'email' in script_content: # 提取JS对象的JSON部分(需要根据页面实际的变量名调整正则) # 比如如果是var profile = {...}; 就把正则里的userInfo换成profile json_match = re.search(r'var userInfo = ({.*?});', script_content, re.DOTALL) if json_match: json_str = json_match.group(1) # 把JS对象转为Python字典 user_data = json.loads(json_str) email = user_data.get('email') if email: print(f"提取到的邮箱地址:{email}") break
注意事项
- 方案2需要你先查看页面的script内容,确认存储邮箱的JS变量名和结构,再调整正则表达式中的变量名(比如
userInfo); - 如果页面有反爬机制,可能需要给
requests.get()添加请求头(比如headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}),避免被拦截。
内容的提问来源于stack exchange,提问作者SIM
相关产品推荐
相关产品推荐

