You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫无输出文件故障排查:基于邮箱爬取社交资料脚本问题

问题排查与修复方案

核心问题1:代码位置与缩进错误

你的if __name__ == "__main__":代码块被放在了get_social_profiles函数内部,且位于return social_profiles语句之后——函数执行到return就会立即退出,这部分代码永远不会被执行。同时,代码块内的缩进混乱(比如email_list = []和后续文件读取代码缩进不匹配),会直接导致语法错误,脚本根本无法正常运行。

修复方式:
将if __name__ == "__main__":整个代码块移到get_social_profiles函数外部,并修正所有语句的缩进,确保代码块内的逻辑处于正确的层级。

核心问题2:Google搜索链接解析错误

当前代码直接提取的href是Google的跳转链接(格式如/url?q=https://linkedin.com/in/xxx&sa=U...),并非真实的社交平台URL,不仅无法直接访问,还可能导致平台判断失效。

修复方式:
引入urllib.parse模块解析真实目标URL,替换原有的链接处理逻辑:

from urllib.parse import unquote, urlparse

# 替换原循环内的href处理部分
href = link.get("href")
if href and href.startswith("/url?q="):
    # 提取并解码真实链接
    real_url = unquote(href.split("/url?q=")[1].split("&")[0])
    parsed_url = urlparse(real_url)
    # 匹配各社交平台
    if "linkedin.com/in/" in real_url:
        profiles.append({"platform": "LinkedIn", "url": real_url})
    elif parsed_url.netloc.endswith("instagram.com") and len(parsed_url.path.strip("/")) > 0:
        profiles.append({"platform": "Instagram", "url": real_url})
    elif parsed_url.netloc.endswith("twitter.com") and len(parsed_url.path.strip("/")) > 0:
        profiles.append({"platform": "Twitter", "url": real_url})
    elif parsed_url.netloc.endswith("facebook.com") and len(parsed_url.path.strip("/")) > 0:
        profiles.append({"platform": "Facebook", "url": real_url})

其他潜在问题

  • 确认emailList.csv的表头为Email(大小写一致,无额外空格),否则会读取不到邮箱数据,最终无输出内容。
  • Google可能会触发反爬机制,导致请求失败,可适当添加请求延迟(如time.sleep(2))避免被封禁。

修复后的完整代码

import requests
import csv
from bs4 import BeautifulSoup
from urllib.parse import unquote, urlparse
import time

def get_social_profiles(email_list):
    social_profiles = {}

    for email in email_list:
        url = f"https://www.google.com/search?q={email}"
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.3"}

        try:
            response = requests.get(url, headers=headers)
            response.raise_for_status()  # 主动抛出HTTP错误
            soup = BeautifulSoup(response.text, "html.parser")

            profiles = []
            for link in soup.find_all("a"):
                href = link.get("href")
                if href and href.startswith("/url?q="):
                    real_url = unquote(href.split("/url?q=")[1].split("&")[0])
                    parsed_url = urlparse(real_url)
                    if "linkedin.com/in/" in real_url:
                        profiles.append({"platform": "LinkedIn", "url": real_url})
                    elif parsed_url.netloc.endswith("instagram.com") and len(parsed_url.path.strip("/")) > 0:
                        profiles.append({"platform": "Instagram", "url": real_url})
                    elif parsed_url.netloc.endswith("twitter.com") and len(parsed_url.path.strip("/")) > 0:
                        profiles.append({"platform": "Twitter", "url": real_url})
                    elif parsed_url.netloc.endswith("facebook.com") and len(parsed_url.path.strip("/")) > 0:
                        profiles.append({"platform": "Facebook", "url": real_url})

            social_profiles[email] = profiles
            time.sleep(2)  # 添加延迟避免反爬

        except requests.exceptions.RequestException as e:
            print(f"处理邮箱 {email} 时出错: {e}")
            social_profiles[email] = []

    return social_profiles

if __name__ == "__main__":
    # 读取邮箱列表
    email_list = []
    try:
        with open("emailList.csv", newline="", encoding="utf-8") as csvfile:
            reader = csv.DictReader(csvfile)
            if "Email" not in reader.fieldnames:
                print("错误:emailList.csv 中未找到 'Email' 表头")
                exit()
            for row in reader:
                email = row.get("Email", "").strip()
                if email:
                    email_list.append(email)
    except FileNotFoundError:
        print("错误:未找到 emailList.csv 文件")
        exit()

    if not email_list:
        print("错误:未读取到有效邮箱")
        exit()

    social_profiles = get_social_profiles(email_list)

    # 写入输出文件
    with open("social_profiles_output.csv", "w", newline="", encoding="utf-8") as csvfile:
        fieldnames = ["Email", "Platform", "URL"]
        writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
        writer.writeheader()

        for email, profiles in social_profiles.items():
            for profile in profiles:
                writer.writerow({"Email": email, "Platform": profile["platform"], "URL": profile["url"]})

    print("社交资料已保存到 social_profiles_output.csv")

内容的提问来源于stack exchange,提问作者CodeNinjaGrasshopper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 07:05:17