You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取网页唯一URL并与数组对比写入谷歌表格问题求助

问题与解决思路

问题概述

需要提取指定网页(如https://www.ig.com/uk/trading-strategies)中的所有唯一URL,与Google Sheet内已存URL对比后,将全新URL的相关数据写入Sheet。目前在提取唯一URL及新旧URL对比环节受阻,尝试过列表、集合、元组的存储方式但未生效,相关代码如下:

obj = {r[2]: True for r in sh.get_all_values()}
ar = []

articles = set()
unique = (articles)

for url in urls:
    my_url = requests.get(url)
    html = my_url.content
    soup = BeautifulSoup(html, "html.parser")
    for item in soup.find_all("h3", class_="article-category-section-title"):
        date = datetime.date.today()
        title = item.find("a", class_="primary js_target").text.strip()
        article = item.find("a", class_="primary js_target").get("href")
        abs = "https://www.ig.com"
        rel = article
        pub = rel[-6:]
        datestring = f"{pub[4:6]} {pub[2:4]} {pub[0:2]}"
        info = {"date": date, "title": title, "url":urllib.parse.urljoin(abs, rel), "published":datestring}
        article = str(info["url"].replace("https://",""))
        articles.add(article)
        if unique not in obj:
            ar.append([str(info["date"]), str(info["title"]), article, str(info["published"])])
if ar != []:
    sh.append_rows(ar, value_input_option="USER_ENTERED")

核心问题分析

  1. 集合逻辑错误:unique = (articles)只是给集合起了别名,并非创建新集合;且unique not in obj是判断整个集合是否在字典键中,完全不符合“判断单个URL是否已存在”的需求。
  2. 去重时机错误:循环中直接添加URL到集合后就判断写入,会导致同一轮循环内重复URL被多次处理,没有真正实现去重对比。

修复步骤

  • 简化集合使用:去掉多余的unique变量,直接用articles集合做去重容器。
  • 修正对比逻辑:处理每个URL时,先判断该URL既不在Sheet已存的obj字典中,也不在当前收集的articles集合中,避免重复添加。
  • 统一URL格式:确保用于对比的URL格式完全一致(比如统一去掉https://),避免因格式差异导致误判。

完整修正代码

obj = {r[2]: True for r in sh.get_all_values()}
ar = []
articles = set()

for url in urls:
    my_url = requests.get(url)
    html = my_url.content
    soup = BeautifulSoup(html, "html.parser")
    for item in soup.find_all("h3", class_="article-category-section-title"):
        date = datetime.date.today()
        title = item.find("a", class_="primary js_target").text.strip()
        rel_url = item.find("a", class_="primary js_target").get("href")
        full_url = urllib.parse.urljoin("https://www.ig.com", rel_url)
        # 统一URL格式,用于对比
        formatted_url = full_url.replace("https://", "")
        
        # 提取发布日期逻辑保留
        pub = rel_url[-6:]
        datestring = f"{pub[4:6]} {pub[2:4]} {pub[0:2]}"
        
        # 双重判断:不在Sheet已存数据中,也未在本次收集的唯一URL里
        if formatted_url not in obj and formatted_url not in articles:
            articles.add(formatted_url)
            ar.append([str(date), title, formatted_url, datestring])

if ar:
    sh.append_rows(ar, value_input_option="USER_ENTERED")

额外优化建议

  • 添加异常捕获:给requests.get增加try-except块,避免网络波动导致程序中断。
  • 健壮日期提取:当前通过rel_url[-6:]取日期的逻辑依赖URL固定格式,建议增加格式校验,防止索引越界。
  • 缩小数据范围:如果Sheet数据量大,sh.get_all_values()可改为只获取URL所在列,减少内存占用。

内容的提问来源于stack exchange,提问作者Mark Leach

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 02:30:53