You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup+requests爬取时如何仅返回单个去重Instagram链接

问题说明

现有Python脚本用于读取xlsx表格中存储的URL列表,访问对应站点首页,检测页面内是否存在关联的Instagram账号链接,检测到后将链接写入URL所在单元格的相邻列。
运行过程中发现,部分站点会在页头、页脚、导航栏等多个位置放置Instagram账号入口,或是嵌入Instagram动态feed,导致单元格写入多个重复结果,示例如下:

['https://instagram.com/xxx', 'https://instagram.com/xxx', 'https://instagram.com/xxx']

实际需求为仅返回单个有效结果,不需要全部重复链接。
原实现代码如下:

for cell in sheet[col][1:]:
    try:
        url = cell.value
        r = requests.get(url)
        ig_get = ['instagram.com']
        ig_get_present = []
        soup = BeautifulSoup(r.content, 'html5lib')
        all_links = soup.find_all('a', href=True)
        print(cell.value)
        for ig_get in ig_get:
            for link in all_links:
                if ig_get in link.attrs['href']:
                    ig_get_present.append(link.attrs['href'])
                    ig_got = str(ig_get_present)
                    print(ig_got)
                    sheet.cell(cell.row, col2).value = ig_got
    except requests.exceptions.ConnectionError:
        pass
    except requests.exceptions.TooManyRedirects:
        pass
    except requests.exceptions.MissingSchema:
        pass
解决方法

原代码的重复问题来自两点:

  • 遍历所有链接时只要匹配到Instagram域名就追加到列表,没有做终止判断,会收集页面所有匹配链接
  • 每匹配到一个链接就立刻写入单元格,最终单元格存储的是持续追加的重复列表

两种可直接落地的修改方案:

方案1:匹配到第一个结果直接终止(效率最高)

遍历链接时只要找到第一个包含instagram.com的链接,就立刻终止后续遍历,直接将该链接写入单元格,不需要收集所有结果。修改后代码如下:

for cell in sheet[col][1:]:
    try:
        url = cell.value
        # 建议增加timeout参数,避免请求无响应卡住程序
        r = requests.get(url, timeout=10)
        soup = BeautifulSoup(r.content, 'html5lib')
        all_links = soup.find_all('a', href=True)
        print(cell.value)
        ig_link = None
        for link in all_links:
            href = link.attrs['href']
            if 'instagram.com' in href:
                ig_link = href
                # 找到第一个匹配结果直接跳出循环
                break
        if ig_link:
            print(ig_link)
            sheet.cell(cell.row, col2).value = ig_link
    except requests.exceptions.ConnectionError:
        pass
    except requests.exceptions.TooManyRedirects:
        pass
    except requests.exceptions.MissingSchema:
        pass

方案2:先去重再取值(适配存在多不同链接的场景)

如果页面可能混入非账号主页的Instagram临时链接,可以先把所有匹配到的链接存入集合自动去重,再从去重后的结果里取第一个值即可,核心替换逻辑如下:

ig_links = set()
for link in all_links:
    href = link.attrs['href']
    if 'instagram.com' in href:
        ig_links.add(href)
# 去重后取第一个有效链接写入
if ig_links:
    ig_link = next(iter(ig_links))
    print(ig_link)
    sheet.cell(cell.row, col2).value = ig_link

内容的提问来源于stack exchange,提问作者Kal-Toh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 04:18:30