遍历50k+URL提取像素高度过慢,求优化方案
批量URL页面区块高度提取性能优化方案
问题描述
遍历50k+的URL列表,提取页面中带有data-analytics-section-engagement属性区块的像素高度时,少量URL运行正常,但数据量增大后循环速度极慢。需要优化循环效率,同时排查并优化JavaScript代码的性能问题。
原代码
# loop through the URLs for url in tqdm.tqdm(urls): # execute the JavaScript code using the webdriver instance driver.get(url) section_pixels = driver.execute_script(""" var cumu = 0; var secAttr = 'data-analytics-section-engagement'; sections = document.querySelectorAll("["+secAttr+"]"); const section_width = []; for (i = 0; i < sections.length; i++) { if (window.getComputedStyle(sections[i]).display != 'none') { cumu = cumu + sections[i].offsetHeight; var section_nme = sections[i].getAttribute(secAttr); var offsetheight = sections[i].offsetHeight; var final_val = section_nme.concat(' *pixels: ',offsetheight); section_width.push(final_val); } } return section_width; """) # print the individual heights for the current URL , just put the console log in the list will remove the cumulative for i in section_pixels: x=i.split('*')[0] y=x.split(':')[1] u=i.split('*')[1] v=u.split(':')[1] final_string=" # ".join([url,y,v]) cleaned_string_list.append(final_string) url_list.append(url) df=pd.DataFrame(cleaned_string_list) # quit the webdriver instance driver.quit()
输出结果
| url | section_name | pixel_height |
|---|---|---|
| https://www.apple.com | hero-iphone-14-pro | 704 |
| https://www.apple.com | hero-iphone-14 | 748 |
| https://www.apple.com | hero-apple-watch-series-8 | 748 |
| https://www.apple.com | promo-wwdc23-announce | 636 |
| https://www.apple.com | promo-ipad | 636 |
优化方案
一、整体循环效率优化
启用无头浏览器
普通浏览器的渲染开销极大,启用无头模式可减少80%以上的资源占用。以Chrome为例,初始化driver时添加参数:from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options)多线程/多进程并发处理
单线程遍历50kURL本质是串行等待,用多线程(适合IO密集型任务)可大幅提升效率。注意每个线程对应独立的driver实例:from concurrent.futures import ThreadPoolExecutor def process_url(url): # 每个线程初始化自己的driver chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) # 执行页面处理逻辑 driver.get(url) section_data = driver.execute_script(优化后的JS代码) driver.quit() return [(url, item['sectionName'], item['pixelHeight']) for item in section_data] # 控制并发数(根据机器配置调整,建议10-20) with ThreadPoolExecutor(max_workers=15) as executor: results = list(tqdm.tqdm(executor.map(process_url, urls), total=len(urls)))调整页面加载策略
不需要等待页面完全加载(图片、广告等无关资源),设置eager模式仅等待DOM加载完成:chrome_options.page_load_strategy = 'eager'
二、JavaScript代码优化
原JS存在冗余的字符串拼接、重复DOM查询问题,优化后直接返回结构化数据,减少Python端的解析开销:
const secAttr = 'data-analytics-section-engagement'; const sections = document.querySelectorAll(`[${secAttr}]`); const result = []; const sectionsLen = sections.length; // 缓存长度避免重复计算 for (let i = 0; i < sectionsLen; i++) { const section = sections[i]; const style = window.getComputedStyle(section); // 缓存样式结果 if (style.display !== 'none') { result.push({ sectionName: section.getAttribute(secAttr), pixelHeight: section.offsetHeight }); } } return result;
三、Python端数据处理优化
- 不要在循环内重复创建DataFrame,待所有数据收集完成后一次性生成:
- 直接使用JS返回的结构化数据,避免字符串拆分的冗余操作:
data_list = [] for url in tqdm.tqdm(urls): driver.get(url) section_data = driver.execute_script(优化后的JS代码) for item in section_data: data_list.append({ 'url': url, 'section_name': item['sectionName'], 'pixel_height': item['pixelHeight'] }) # 循环结束后生成DataFrame df = pd.DataFrame(data_list)
四、其他辅助优化
- 禁用浏览器缓存:添加
chrome_options.add_argument("--disable-cache"),避免重复加载无关资源 - 预过滤无效URL:用
requests.head快速判断URL是否可达,跳过无效链接减少无用请求 - 复用driver实例:若用单线程,每次请求前清理缓存(
driver.delete_all_cookies()+window.localStorage.clear()),避免重启driver的开销
内容的提问来源于stack exchange,提问作者user3451371
相关产品推荐
相关产品推荐

