批量爬取URL文本时索引越界问题及提速方案咨询
GW2日志批量爬取问题排查与优化方案
问题描述
现有log_scrape函数可单URL爬取GW2日志页面信息并保存为JSON,但批量读取gw2_urls.txt中的URL调用函数时,data = soup.find_all('script')[8]始终报List out of index错误,单URL运行则正常。
错误原因分析
批量读取URL时,文件中每行URL末尾带有换行符\n,直接传入requests.get会导致请求的URL包含无效换行字符,服务器返回的页面结构与正常请求不同,页面内<script>标签数量不足9个(索引从0开始),因此访问索引8时触发越界错误。而单URL测试时手动输入的URL无换行符,请求正常,页面结构符合预期。
修复步骤
在批量循环中对读取的URL进行首尾空白字符清理:
with open('gw2_urls.txt', 'r') as urls: for url in urls: # 去掉URL首尾的换行、空格等空白字符 cleaned_url = url.strip() if cleaned_url: # 跳过空行 print(log_scrape(cleaned_url))
爬取速度优化方案
1. 并发请求(IO密集型任务优先用多线程)
使用concurrent.futures.ThreadPoolExecutor实现多线程并发请求,大幅提升批量爬取效率:
from concurrent.futures import ThreadPoolExecutor def main(): with open('gw2_urls.txt', 'r') as f: urls = [url.strip() for url in f if url.strip()] # 设定线程数,根据目标网站反爬策略调整,建议5-10 with ThreadPoolExecutor(max_workers=8) as executor: results = executor.map(log_scrape, urls) for result in results: print(result) if __name__ == '__main__': main()
2. 优化页面解析器
将BeautifulSoup的解析器从html.parser替换为lxml(需先安装pip install lxml),解析速度更快:
soup = BeautifulSoup(response.content, 'lxml')
3. 完善请求配置
- 设置请求超时,避免因单个请求卡住拖慢整体速度:
response = requests.get(url=cleaned_url, headers=HEADERS, timeout=10) - 检查响应状态码,仅在状态码为200时继续处理,避免无效页面解析:
if response.status_code != 200: return f"Request failed with status {response.status_code}"
4. 路径处理优化
使用pathlib模块替代字符串拼接路径,避免Windows下的转义问题,同时简化目录创建逻辑(如果目录不存在需先创建):
from pathlib import Path # 示例:替换原路径拼接逻辑 base_path = Path('ETL/EXTRACT_00/Web Scraping/Boss_data') # 根据bossTag设置子路径 if bossTag == 'vg': path_name = base_path / 'Wing_1' / 'Valley_Guardian' # ...其他bossTag逻辑... # 确保目录存在,不存在则创建 path_name.mkdir(parents=True, exist_ok=True) # 写入文件 with open(path_name / f'{bossName}.json', 'w') as f: for line in logData: f.write(line)
5. 减少重复计算
将HEADERS定义移到函数外部,避免每次调用函数都重新创建字典:
HEADERS = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Safari/537.36 Edge/12.246'} def log_scrape(url): response = requests.get(url=url, headers=HEADERS) # ...后续逻辑...
内容的提问来源于stack exchange,提问作者Icharo-tb
相关产品推荐
相关产品推荐

