You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理45个LinkedIn URL后遇invalid session id错误及URL格式转换需求

LinkedIn URL转换与Selenium会话错误解决指南

一、解决"invalid session id"错误

错误原因

处理45个URL后出现会话无效,核心原因是LinkedIn反爬机制检测到自动化行为,强制终止浏览器会话;其次是ChromeDriver长时间运行导致连接超时、会话老化。

修复方案

  • 定期重启浏览器会话:每处理15-20个URL就重启一次Driver,避免会话被标记为异常。
  • 替换固定sleep为显式等待:用Selenium的WebDriverWait等待页面关键元素加载完成,既保证页面跳转完成,又避免无意义的等待。
  • 强化反爬规避配置:添加模拟真实用户的Chrome参数,隐藏自动化特征,降低被检测概率。
  • 单URL错误重试:针对单个URL的会话错误单独重试2-3次,避免整批任务中断。
  • 增量保存数据:每处理10个URL就保存一次结果,防止会话崩溃导致数据丢失。

二、自动化转换Sales Navigator URL格式

Sales Navigator返回的linkedin.com/in/ACwAAAJnuBoBdG....是临时跳转链接,会自动301重定向到linkedin.com/in/jan-veleba这类友好URL。提供两种实现方案:

  1. 轻量方案(推荐):用requests库直接获取重定向后的URL,无需启动浏览器,速度快且不易触发反爬限制(需携带LinkedIn登录Cookie)。
  2. Selenium方案:利用现有浏览器会话,页面跳转完成后获取current_url即可。

修改后的完整代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.common.exceptions import WebDriverException, TimeoutException
import time
import random
import requests

# ------------------- URL转换工具函数(轻量方案:requests) -------------------
def get_friendly_linkedin_url(raw_url, cookies):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    try:
        response = requests.get(raw_url, headers=headers, cookies=cookies, allow_redirects=True)
        return response.url if response.status_code == 200 else None
    except Exception as e:
        print(f"URL转换失败 {raw_url}: {str(e)}")
        return None

# ------------------- Selenium相关函数 -------------------
def initialize_driver():
    chrome_options = webdriver.ChromeOptions()
    # 反爬增强配置
    chrome_options.add_argument("--disable-extensions")
    chrome_options.add_argument("--disable-gpu")
    chrome_options.add_argument("--no-sandbox")
    chrome_options.add_argument("--disable-dev-shm-usage")
    chrome_options.add_argument("--disable-blink-features=AutomationControlled")
    chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
    chrome_options.add_experimental_option('useAutomationExtension', False)
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
    
    driver = webdriver.Chrome(options=chrome_options)
    driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")
    return driver

def get_updated_url(driver, url):
    try:
        # 显式等待个人资料标题加载完成
        WebDriverWait(driver, 20).until(
            EC.presence_of_element_located((By.CLASS_NAME, "text-heading-xlarge"))
        )
        return driver.current_url
    except TimeoutException:
        print(f"页面加载超时 {url}")
        return None
    except Exception as e:
        print(f"获取URL失败 {url}: {str(e)}")
        return None

# ------------------- 主逻辑 -------------------
def main():
    df = pd.read_csv('IN.csv', sep='\t')
    url_column_index = 3
    if url_column_index >= len(df.columns):
        print("无效的列索引")
        return
    
    raw_urls = df.iloc[:, url_column_index].tolist()
    out_df = pd.DataFrame(columns=['Skutecne_URL'])
    processed_count = 0
    driver = None

    # 配置requests方案的Cookie(从浏览器导出li_at字段)
    # cookies = {'li_at': '你的LinkedIn登录Cookie值'}
    
    for idx, url in enumerate(raw_urls):
        # 每20个URL重启一次Driver
        if processed_count % 20 == 0:
            if driver:
                driver.quit()
            driver = initialize_driver()
            # 首次启动需手动登录
            if processed_count == 0:
                print("请在浏览器中完成LinkedIn登录,按回车继续...")
                input()
        
        try:
            # 方案1:Selenium获取重定向URL
            driver.get(url)
            updated_url = get_updated_url(driver, url)
            
            # 方案2:requests获取(替换上方代码,需配置cookies)
            # updated_url = get_friendly_linkedin_url(url, cookies)
            
            if updated_url:
                out_df.loc[idx] = [updated_url]
                print(f"已处理 {idx+1}/{len(raw_urls)}: {updated_url}")
                # 每10个URL增量保存
                if idx % 10 == 0:
                    out_df.to_csv('OUT.csv', index=False)
            else:
                out_df.loc[idx] = [None]
            
            processed_count += 1
            time.sleep(random.randint(15, 35))
            
            # 每100个URL休息24小时
            if processed_count % 100 == 0:
                print("已处理100个URL,休息24小时...")
                driver.quit()
                time.sleep(24 * 60 * 60)
                driver = initialize_driver()
                print("重新登录LinkedIn后按回车继续...")
                input()
                processed_count = 0
        
        except WebDriverException as e:
            print(f"会话错误,重试URL {url}: {str(e)}")
            driver.quit()
            driver = initialize_driver()
            driver.get(url)
            updated_url = get_updated_url(driver, url)
            out_df.loc[idx] = [updated_url] if updated_url else [None]
            time.sleep(random.randint(20, 40))
        
        except Exception as e:
            print(f"处理URL失败 {url}: {str(e)}")
            out_df.loc[idx] = [None]
    
    # 最终保存结果
    out_df.to_csv('OUT.csv', index=False)
    df['SLoupec37'] = out_df['Skutecne_URL']
    df.to_csv('IN.csv', sep='\t', index=False)
    
    if driver:
        driver.quit()
    print("任务完成")

if __name__ == "__main__":
    main()

关键修改说明

  • 会话管理:每20个URL重启Driver,避免会话被LinkedIn封禁。
  • 等待优化:用WebDriverWait替代固定sleep,确保页面加载完成再获取URL。
  • 反爬增强:添加多个Chrome参数隐藏自动化特征,模拟真实用户行为。
  • URL转换:提供两种方案,requests更高效,Selenium适合需保持浏览器会话的场景。
  • 数据安全:增量保存结果,防止会话崩溃丢失数据。
  • 错误重试:针对会话错误单独重试当前URL,提升任务完成率。

内容的提问来源于stack exchange,提问作者martas prchal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 15:45:54