You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取数据存入pandas DataFrame时Tags列仅显示最后一个标签的原因

问题原因
  • 遍历标签的循环中,你每次都给quote_tag变量赋值,新值会直接覆盖旧值,循环结束后该变量仅保留最后一个标签的内容,因此存入DataFrame的Tags列时就只剩最后一个标签。
  • 补充提示:pd.DataFrame.append方法已在pandas 2.0及以上版本被移除,继续使用会触发报错,建议调整数据构造逻辑。
修正代码
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
import pandas as pd

driver = webdriver.Chrome(service=Service(executable_path=ChromeDriverManager().install()))
driver.maximize_window()
driver.get('https://quotes.toscrape.com/')

# 先通过列表收集所有爬取数据,性能更高也更符合pandas最佳实践
data_list = []
quotes = driver.find_elements(By.CSS_SELECTOR, '.quote')
for quote in quotes:
    text = quote.find_element(By.CSS_SELECTOR, '.text').text
    author = quote.find_element(By.CSS_SELECTOR, '.author').text
    # 收集所有标签文本后用空格拼接,得到完整标签字符串
    tag_texts = [tag.text for tag in quote.find_elements(By.CSS_SELECTOR, '.tag')]
    full_tags = ' '.join(tag_texts)
    data_list.append({
        'Quote': text,
        'Author': author,
        'Tags': full_tags
    })

# 批量生成DataFrame后导出
df = pd.DataFrame(data_list)
df.to_csv('C:/Users/Jay/Downloads/Python/!Learn/practice/scraping/selenium/quotes.csv', index=False)

内容的提问来源于stack exchange,提问作者Jay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 18:36:07