You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Selenium点击按钮后提取网页script中的url属性

问题:TikTok创意中心页面提取URL属性失败

我正在学习Python和网络爬虫,本次项目尝试从TikTok创意中心的页面提取元素的"url"属性。已用Selenium编写脚本实现点击按钮10次的操作,但通过find_elements、类名、XPath、CSS选择器等方法均无法找到目标数据,尝试requests、json、bs4库也出现报错,希望得到解决方法。

现有代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.action_chains import ActionChains


# Set up the Chrome driver and maximize the window
options = Options()
options.add_argument("--start-maximized")

service = Service(r'C:\windows\chromedriver.exe')

driver = webdriver.Chrome(options=options, service=service)
company = 'nvidia-corp'

# Navigate to the URL
url = f"https://ads.tiktok.com/business/creativecenter/hashtag/example/pc/en?countryCode=US&period=7"
driver.get(url)

wait = WebDriverWait(driver, 5)
# Find the button element using the full XPath
button_xpath = '//*[@id="trendHashtagDetail"]/div[2]/div[3]/div[3]/div/button[2]'
button = wait.until(EC.element_to_be_clickable((By.XPATH, button_xpath)))

# Click the button 10 times
for i in range(10):
    actions = ActionChains(driver)
    actions.move_to_element(button).click().perform()

#Find all the url attributes and save/print them

解决方法

1. 检查并切换到iframe

TikTok创意中心的内容可能嵌套在iframe中,先切换到对应iframe再查找元素:

# 点击按钮循环后,切换到iframe(需根据实际页面调整iframe定位方式)
iframe = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "iframe")))
driver.switch_to.frame(iframe)

# 之后再查找目标元素
target_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "[url]")))
urls = [elem.get_attribute("url") for elem in target_elements]
print(urls)

2. 增加元素加载等待

每次点击按钮后,新内容需要时间加载,避免因元素未加载完成导致查找失败:

for i in range(10):
    actions = ActionChains(driver)
    actions.move_to_element(button).click().perform()
    # 等待新内容加载完成(替换为目标元素的选择器)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "[url]")))

3. 优化元素选择器

使用属性选择器直接定位带有url属性的元素,比类名或XPath更可靠:

# 点击完成后提取URL
target_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "[url]")))
urls = []
for elem in target_elements:
    url = elem.get_attribute("url")
    if url:
        urls.append(url)
# 去重(如果有重复)
unique_urls = list(set(urls))
print(unique_urls)

4. 绕过Selenium检测

TikTok有反爬机制,给Chrome添加防检测参数:

options = Options()
options.add_argument("--start-maximized")
# 防检测参数
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

5. 直接调用API接口

如果用requests报错,是因为缺少请求头和Cookie。打开浏览器开发者工具(F12),找到加载数据的接口,复制请求头和参数后调用:

import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Cookie": "从浏览器复制的Cookie内容"
}
# 替换为实际的API接口地址
api_url = "https://ads.tiktok.com/business/creativecenter/api/xxx/xxx"
params = {
    "countryCode": "US",
    "period": "7",
    # 其他必要参数
}
response = requests.get(api_url, headers=headers, params=params)
if response.status_code == 200:
    data = response.json()
    # 根据返回的JSON结构提取url,示例结构
    urls = [item.get("url") for item in data.get("data", {}).get("list", []) if item.get("url")]
    print(urls)
else:
    print(f"请求失败,状态码:{response.status_code}")

内容的提问来源于stack exchange,提问作者Esenomo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 05:47:47