You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Clearbit API爬取的Logo结果追加至现有Pandas DataFrame

问题:将Clearbit爬取的Logo链接追加到Pandas DataFrame中

我当前正在使用Clearbit API结合Selenium、BeautifulSoup爬取企业Logo,现有代码如下:

import pandas as pd
from selenium import webdriver
from bs4 import BeautifulSoup

data = {'name':  ['tcs', 'orange', 'linkedin'],
        'domain': ["tcs.com",
                    "orange.com",
                    "linkedin.com"]}

df = pd.DataFrame(data)
driver = webdriver.Chrome(r"chromedriver.exe")

for i in df['domain']:
    driver.get("https://logo.clearbit.com/" + str(i))
    clear_api_html = BeautifulSoup(driver.page_source, 'html.parser')
    clear_logo_access = clear_api_html.find_all('img')
    output_dict = {'Logo': clear_logo_access, 'Website': i}
    print(output_dict)

当前输出为字典格式:

{'Logo': [<img src="https://logo.clearbit.com/tcs.com" style="display: block;-webkit-user-select: none;margin: auto;background-color: hsl(0, 0%, 90%);transition: background-color 300ms;"/>], 'Website': 'tcs.com'}
{'Logo': [<img src="https://logo.clearbit.com/orange.com" style="display: block;-webkit-user-select: none;margin: auto;background-color: hsl(0, 0%, 90%);transition: background-color 300ms;"/>], 'Website': 'orange.com'}
{'Logo': [<img src="https://logo.clearbit.com/linkedin.com" style="display: block;-webkit-user-select: none;margin: auto;background-color: hsl(0, 0%, 90%);transition: background-color 300ms;"/>], 'Website': 'linkedin.com'}

希望将这些结果追加到现有的Pandas DataFrame中,得到包含name、domain、Logo列的表格,解决方案如下:


修改后的代码

import pandas as pd
from selenium import webdriver
from bs4 import BeautifulSoup

data = {'name':  ['tcs', 'orange', 'linkedin'],
        'domain': ["tcs.com",
                    "orange.com",
                    "linkedin.com"]}

df = pd.DataFrame(data)
driver = webdriver.Chrome(r"chromedriver.exe")

# 初始化Logo列
df['Logo'] = ''

for idx, domain in enumerate(df['domain']):
    driver.get(f"https://logo.clearbit.com/{domain}")
    clear_api_html = BeautifulSoup(driver.page_source, 'html.parser')
    # 获取单个img标签并提取src属性
    logo_img = clear_api_html.find('img')
    if logo_img:
        df.loc[idx, 'Logo'] = logo_img.get('src')

# 关闭浏览器释放资源
driver.quit()

# 打印最终结果
print(df)

关键修改说明

  • 提前为DataFrame创建空的Logo列,避免动态添加列的额外开销
  • 使用enumerate遍历域名,同时获取行索引,直接定位DataFrame的对应行进行赋值
  • 用find()替代find_all(),因为目标页面仅存在一个Logo图片,无需返回列表对象
  • 提取img标签的src属性值(即Logo的实际URL),而非保存整个BeautifulSoup标签对象
  • 遍历结束后调用driver.quit()关闭浏览器,释放系统资源

内容的提问来源于stack exchange,提问作者Sushmitha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 08:01:26