Python提取网页唯一URL并与数组对比写入谷歌表格问题求助
问题与解决思路
问题概述
需要提取指定网页(如https://www.ig.com/uk/trading-strategies)中的所有唯一URL,与Google Sheet内已存URL对比后,将全新URL的相关数据写入Sheet。目前在提取唯一URL及新旧URL对比环节受阻,尝试过列表、集合、元组的存储方式但未生效,相关代码如下:
obj = {r[2]: True for r in sh.get_all_values()} ar = [] articles = set() unique = (articles) for url in urls: my_url = requests.get(url) html = my_url.content soup = BeautifulSoup(html, "html.parser") for item in soup.find_all("h3", class_="article-category-section-title"): date = datetime.date.today() title = item.find("a", class_="primary js_target").text.strip() article = item.find("a", class_="primary js_target").get("href") abs = "https://www.ig.com" rel = article pub = rel[-6:] datestring = f"{pub[4:6]} {pub[2:4]} {pub[0:2]}" info = {"date": date, "title": title, "url":urllib.parse.urljoin(abs, rel), "published":datestring} article = str(info["url"].replace("https://","")) articles.add(article) if unique not in obj: ar.append([str(info["date"]), str(info["title"]), article, str(info["published"])]) if ar != []: sh.append_rows(ar, value_input_option="USER_ENTERED")
核心问题分析
- 集合逻辑错误:
unique = (articles)只是给集合起了别名,并非创建新集合;且unique not in obj是判断整个集合是否在字典键中,完全不符合“判断单个URL是否已存在”的需求。 - 去重时机错误:循环中直接添加URL到集合后就判断写入,会导致同一轮循环内重复URL被多次处理,没有真正实现去重对比。
修复步骤
- 简化集合使用:去掉多余的
unique变量,直接用articles集合做去重容器。 - 修正对比逻辑:处理每个URL时,先判断该URL既不在Sheet已存的
obj字典中,也不在当前收集的articles集合中,避免重复添加。 - 统一URL格式:确保用于对比的URL格式完全一致(比如统一去掉
https://),避免因格式差异导致误判。
完整修正代码
obj = {r[2]: True for r in sh.get_all_values()} ar = [] articles = set() for url in urls: my_url = requests.get(url) html = my_url.content soup = BeautifulSoup(html, "html.parser") for item in soup.find_all("h3", class_="article-category-section-title"): date = datetime.date.today() title = item.find("a", class_="primary js_target").text.strip() rel_url = item.find("a", class_="primary js_target").get("href") full_url = urllib.parse.urljoin("https://www.ig.com", rel_url) # 统一URL格式,用于对比 formatted_url = full_url.replace("https://", "") # 提取发布日期逻辑保留 pub = rel_url[-6:] datestring = f"{pub[4:6]} {pub[2:4]} {pub[0:2]}" # 双重判断:不在Sheet已存数据中,也未在本次收集的唯一URL里 if formatted_url not in obj and formatted_url not in articles: articles.add(formatted_url) ar.append([str(date), title, formatted_url, datestring]) if ar: sh.append_rows(ar, value_input_option="USER_ENTERED")
额外优化建议
- 添加异常捕获:给
requests.get增加try-except块,避免网络波动导致程序中断。 - 健壮日期提取:当前通过
rel_url[-6:]取日期的逻辑依赖URL固定格式,建议增加格式校验,防止索引越界。 - 缩小数据范围:如果Sheet数据量大,
sh.get_all_values()可改为只获取URL所在列,减少内存占用。
内容的提问来源于stack exchange,提问作者Mark Leach
相关产品推荐
相关产品推荐

