Python读取Excel/CSV列URL批量爬取网页文本问题求解
问题说明
data.URL返回的是pandas Series结构的整列数据,不是单个URL字符串,而requests.get仅支持传入单个字符串格式的URL地址,因此无法直接传整列数据发起请求,需要逐行遍历URL列逐个发起请求。
完整修改代码
import pandas as pd import requests from bs4 import BeautifulSoup import time # 配置请求头,替换为自己浏览器的User-Agent即可 headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # 读取Excel文件 data = pd.read_excel('/Users/LE/Downloads/url.xlsx') # 新增两列,分别存储解析后的网页纯文本、请求错误信息 data['page_text'] = '' data['error_msg'] = '' # 逐行遍历每个URL发起请求 for idx, url in data['URL'].items(): # 跳过空值/格式无效的URL if pd.isna(url) or not isinstance(url, str): data.loc[idx, 'error_msg'] = 'URL为空或格式无效' continue try: # 发起请求,设置10秒超时避免卡死 res = requests.get(url, headers=headers, timeout=10) # 自动识别网页编码,避免中文乱码 res.encoding = res.apparent_encoding html = res.text soup = BeautifulSoup(html, 'lxml') # 剔除script、style等非内容标签,减少无效文本干扰 for redundant_tag in soup(['script', 'style', 'meta', 'link', 'noscript']): redundant_tag.decompose() # 提取结构化纯文本 page_text = soup.get_text(strip=True, separator='\n') data.loc[idx, 'page_text'] = page_text # 可根据需求调整请求间隔,避免频率过高被反爬拦截 time.sleep(0.5) except Exception as e: # 捕获全流程异常,单个URL失败不中断整体任务 data.loc[idx, 'error_msg'] = str(e) # 处理结果存入新Excel,方便后续做关键词检索 data.to_excel('/Users/LE/Downloads/url_with_text.xlsx', index=False)
关键注意点
- 逐行请求必须加异常捕获,避免单个URL失效、连接超时、反爬拦截等问题直接终止整个程序
- 提取文本前先剔除冗余标签,能大幅减少后续关键词检索的无效干扰内容
- 如果URL数量较多,可适当调大请求间隔,不要短时间发起大量请求,避免被网站封禁IP
- 最终生成的结果文件中,
page_text列存储的就是对应URL的网页纯文本,直接在该列做关键词匹配即可
Excel示例参考

内容的提问来源于stack exchange,提问作者Candice LE
相关产品推荐
相关产品推荐

