You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历50k+URL提取像素高度过慢,求优化方案

批量URL页面区块高度提取性能优化方案

问题描述

遍历50k+的URL列表,提取页面中带有data-analytics-section-engagement属性区块的像素高度时,少量URL运行正常,但数据量增大后循环速度极慢。需要优化循环效率,同时排查并优化JavaScript代码的性能问题。

原代码

# loop through the URLs
for url in tqdm.tqdm(urls):
    # execute the JavaScript code using the webdriver instance
    driver.get(url)
    section_pixels = driver.execute_script("""
        var cumu = 0;
        var secAttr = 'data-analytics-section-engagement';
        sections = document.querySelectorAll("["+secAttr+"]");
        const section_width = [];
        for (i = 0; i < sections.length; i++) {
            if (window.getComputedStyle(sections[i]).display != 'none') {
                cumu = cumu + sections[i].offsetHeight;
                var section_nme = sections[i].getAttribute(secAttr);
                var offsetheight = sections[i].offsetHeight;
                var final_val = section_nme.concat(' *pixels: ',offsetheight);
                section_width.push(final_val);
            }
        }
        return section_width;
    """)
    
    # print the individual heights for the current URL , just put the console log in the list will remove the cumulative    
    for i in section_pixels:
        x=i.split('*')[0]
        y=x.split(':')[1]
        u=i.split('*')[1]
        v=u.split(':')[1]
        final_string=" # ".join([url,y,v])
        cleaned_string_list.append(final_string)
        url_list.append(url)
    df=pd.DataFrame(cleaned_string_list)
# quit the webdriver instance
driver.quit()

输出结果

urlsection_namepixel_height
https://www.apple.comhero-iphone-14-pro704
https://www.apple.comhero-iphone-14748
https://www.apple.comhero-apple-watch-series-8748
https://www.apple.compromo-wwdc23-announce636
https://www.apple.compromo-ipad636

优化方案

一、整体循环效率优化

  1. 启用无头浏览器
    普通浏览器的渲染开销极大,启用无头模式可减少80%以上的资源占用。以Chrome为例,初始化driver时添加参数:

    from selenium.webdriver.chrome.options import Options
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--disable-gpu")
    driver = webdriver.Chrome(options=chrome_options)
    
  2. 多线程/多进程并发处理
    单线程遍历50kURL本质是串行等待,用多线程(适合IO密集型任务)可大幅提升效率。注意每个线程对应独立的driver实例:

    from concurrent.futures import ThreadPoolExecutor
    
    def process_url(url):
        # 每个线程初始化自己的driver
        chrome_options = Options()
        chrome_options.add_argument("--headless=new")
        driver = webdriver.Chrome(options=chrome_options)
        # 执行页面处理逻辑
        driver.get(url)
        section_data = driver.execute_script(优化后的JS代码)
        driver.quit()
        return [(url, item['sectionName'], item['pixelHeight']) for item in section_data]
    
    # 控制并发数(根据机器配置调整,建议10-20)
    with ThreadPoolExecutor(max_workers=15) as executor:
        results = list(tqdm.tqdm(executor.map(process_url, urls), total=len(urls)))
    
  3. 调整页面加载策略
    不需要等待页面完全加载(图片、广告等无关资源),设置eager模式仅等待DOM加载完成:

    chrome_options.page_load_strategy = 'eager'
    

二、JavaScript代码优化

原JS存在冗余的字符串拼接、重复DOM查询问题,优化后直接返回结构化数据,减少Python端的解析开销:

const secAttr = 'data-analytics-section-engagement';
const sections = document.querySelectorAll(`[${secAttr}]`);
const result = [];
const sectionsLen = sections.length; // 缓存长度避免重复计算

for (let i = 0; i < sectionsLen; i++) {
    const section = sections[i];
    const style = window.getComputedStyle(section); // 缓存样式结果
    if (style.display !== 'none') {
        result.push({
            sectionName: section.getAttribute(secAttr),
            pixelHeight: section.offsetHeight
        });
    }
}
return result;

三、Python端数据处理优化

  1. 不要在循环内重复创建DataFrame,待所有数据收集完成后一次性生成:
  2. 直接使用JS返回的结构化数据,避免字符串拆分的冗余操作:
data_list = []
for url in tqdm.tqdm(urls):
    driver.get(url)
    section_data = driver.execute_script(优化后的JS代码)
    for item in section_data:
        data_list.append({
            'url': url,
            'section_name': item['sectionName'],
            'pixel_height': item['pixelHeight']
        })
# 循环结束后生成DataFrame
df = pd.DataFrame(data_list)

四、其他辅助优化

  • 禁用浏览器缓存:添加chrome_options.add_argument("--disable-cache"),避免重复加载无关资源
  • 预过滤无效URL:用requests.head快速判断URL是否可达,跳过无效链接减少无用请求
  • 复用driver实例:若用单线程,每次请求前清理缓存(driver.delete_all_cookies()+window.localStorage.clear()),避免重启driver的开销

内容的提问来源于stack exchange,提问作者user3451371

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:27:19