You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:使用XPath获取YouTube视频播放量及上传日期异常

解决Selenium获取YouTube视频信息返回WebElement的问题

你的问题核心是:driver.find_element_by_xpath()方法返回的是WebElement对象,而不是标签里的实际数据。你要提取的name、uploadDate、interactionCount这些信息都存在meta标签的content属性里,必须通过.get_attribute('content')来获取。

另外你的代码还有几个小问题需要修正:

  • 缺少pandas和time的导入语句,代码运行会报错
  • 空的except: pass会掩盖所有错误,不利于调试

修正后的代码:

import requests
from bs4 import BeautifulSoup
from selenium import webdriver
import pandas as pd  # 补充导入pandas
import time  # 补充导入time

driver = webdriver.Chrome(executable_path='C:/XXX/chromedriver.exe')

dataset = pd.read_csv(r"C:\XXXX.csv", skiprows=0)
dataset.head()

for index, row in dataset[0:1].iterrows():
    try:
        links = str(dataset.loc[index,'youtube_link'])
        driver.get(links)
        time.sleep(3)
        # 获取每个meta标签的content属性值
        video_title = driver.find_element_by_xpath("//meta[@itemprop='name']").get_attribute('content')
        upload_date = driver.find_element_by_xpath("//meta[@itemprop='uploadDate']").get_attribute('content')
        view_count = driver.find_element_by_xpath("//meta[@itemprop='interactionCount']").get_attribute('content')
        print(index, ":", links, video_title, upload_date, view_count)
    except Exception as e: 
        print(f"处理链接{links}时出错: {e}")  # 打印错误信息,方便调试

补充说明:

  • get_attribute('content')是Selenium提取元素属性值的标准方法,适用于所有带属性的HTML标签
  • 如果页面加载慢,time.sleep(3)可能不够,你可以改用Selenium的显式等待来替代固定休眠,稳定性更高:
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # 替换time.sleep(3)为:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//meta[@itemprop='name']"))
    )
    

内容的提问来源于stack exchange,提问作者Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 09:40:31