You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合法抓取TMDN网站隐藏iframe内商标数据?直接访问链接受限

合法抓取TMDN tmview隐藏iframe内商标数据的解决方案

我明白你的核心困扰:直接访问iframe的src链接会被网站限制封禁,而你当前基于Selenium的方案已经能进入详情页切换iframe,但需要优化操作模式来避开反爬机制,同时确保行为符合网站规范。下面是针对性的解决思路和优化后的代码:

一、核心优化方向:模拟真实用户行为

网站的反爬系统通常会识别非人类的操作特征,比如固定间隔的点击、无交互的页面跳转、完全禁用资源加载等。我们需要调整代码来模拟真实用户的操作习惯:

  • 替换固定延迟为随机等待:用random.uniform()生成随机间隔,避免每次操作都用相同的time.sleep()
  • 使用显式等待代替强制睡眠:用WebDriverWait等待元素加载完成,既提升效率又让操作更稳定
  • 启用正常资源加载:完全禁用图片容易被识别为爬虫,建议保持默认设置或仅禁用非必要资源
  • 添加真实User-Agent:模拟主流浏览器的请求头,避免被标记为异常请求
  • 增加自然交互:比如页面滚动、鼠标随机移动,让操作轨迹更贴近真实用户

二、优化后的代码实现

import pandas as pd
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException, TimeoutException
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import Select, WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import traceback
import time
import random

def random_sleep(min_sec=2, max_sec=5):
    """生成随机等待时间,模拟用户操作间隔"""
    time.sleep(random.uniform(min_sec, max_sec))

# 配置Chrome选项,模拟真实浏览器
option = webdriver.ChromeOptions()
option.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 可选:启用无头模式(需额外配置避免被检测)
# option.add_argument("--headless=new")
# option.add_argument("--disable-blink-features=AutomationControlled")
# option.add_experimental_option("excludeSwitches", ["enable-automation"])
# option.add_experimental_option('useAutomationExtension', False)

url ="https://www.tmdn.org/tmview/welcome#"
xlsName = 'D:\\test.xlsx'
records = []
start_time = time.time()

driver = webdriver.Chrome(executable_path="D:\\Python\\chromedriver.exe", options=option)
driver.get(url)
random_sleep(8, 12)

# 处理隐私政策弹窗
try:
    WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.XPATH, '//*[@id="buttonBox"]/a'))
    ).click()
except TimeoutException:
    print("未找到同意按钮,跳过")
random_sleep(8, 12)

x=-1
try:
    # 点击高级搜索
    WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.NAME, "lnkAdvancedSearch"))
    ).click()
    random_sleep(3, 6)

    # 选择指定地区:United Kingdom
    driver.find_element(By.ID, 'DesignatedTerritories').click()
    random_sleep(3, 6)
    TerritoryLabelElements = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.optEUGroupContainer label'))
    )
    for elem in TerritoryLabelElements:
        if elem.text == 'United Kingdom':
            elem.click()
            break
    random_sleep(3, 6)
    driver.find_element(By.ID, 'DesignatedTerritories').click()
    random_sleep(3, 6)

    # 选择商标局:GB United Kingdom ( UKIPO )
    driver.find_element(By.ID, 'SelectedOffices').click()
    random_sleep(3, 6)
    TerritoryLabelElements = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.multiSelectOptions label'))
    )
    for elem in TerritoryLabelElements:
        if elem.text == 'GB United Kingdom ( UKIPO )':
            elem.click()
            break
    random_sleep(3, 6)
    driver.find_element(By.ID, 'SelectedOffices').click()
    random_sleep(3, 6)

    # 选择商标状态:Filed和Registered
    driver.find_element(By.ID, 'TradeMarkStatus').click()
    random_sleep(3, 6)
    TerritoryLabelElements = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.multiSelectOptions label'))
    )
    for elem in TerritoryLabelElements:
        if elem.text in ['Filed', 'Registered']:
            elem.click()
    random_sleep(3, 6)
    driver.find_element(By.ID, 'TradeMarkStatus').click()
    random_sleep(3, 6)

    # 设置日期范围
    startdate = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "ApplicationDateFrom"))
    )
    startdate.clear()
    startdate.send_keys('01-10-2018')
    random_sleep(1, 3)

    enddate = driver.find_element(By.ID, "ApplicationDateTo")
    enddate.clear()
    enddate.send_keys('31-10-2018')
    random_sleep(1, 3)

    # 执行搜索
    WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.ID, "SearchCopy"))
    ).click()
    random_sleep(5, 10)

    # 切换到每页100条结果
    WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.LINK_TEXT, '100'))
    ).click()
    random_sleep(5, 10)

    # 循环处理每页数据
    for i in range(1, 73):
        # 等待表格加载完成
        WebDriverWait(driver, 15).until(
            EC.presence_of_element_located((By.ID, "grid"))
        )
        html = driver.page_source
        soup = BeautifulSoup(html, 'html.parser')
        tbl = soup.find("table", id="grid")
        tr_rows = tbl.find_all('tr')[1:]  # 跳过表头

        x = -1
        for tr_row in tr_rows:
            x += 1
            td_cells = tr_row.find_all('td')
            # 提取表格基础数据
            Trade_mark_name = td_cells[4].text.strip()
            Trade_mark_office = td_cells[5].text.strip()
            Designated_territory = td_cells[6].text.strip()
            Application_number = td_cells[7].text.strip()
            Registration_number = td_cells[8].text.strip()
            Trade_mark_status = td_cells[9].text.strip()
            Trade_mark_type = td_cells[13].text.strip()
            Applicant_name = td_cells[11].text.strip()
            Nice_class = td_cells[10].text.strip()
            Application_date = td_cells[12].text.strip()
            Registration_date = td_cells[14].text.strip()

            # 点击进入详情页
            tm_links = WebDriverWait(driver, 10).until(
                EC.presence_of_all_elements_located((By.CLASS_NAME, 'cell_tmName_column'))
            )
            el = tm_links[x]
            action = webdriver.common.action_chains.ActionChains(driver)
            action.move_to_element(el).click().perform()
            random_sleep(3, 6)

            # 切换到详情iframe(通过ID定位更准确)
            Owner_Address = 'No Entry'
            Representative_Name = 'No Entry'
            try:
                iframe = WebDriverWait(driver, 10).until(
                    EC.presence_of_element_located((By.ID, 'iframe_0'))
                )
                driver.switch_to.frame(iframe)
                random_sleep(2, 4)

                # 提取iframe内数据
                html2 = driver.page_source
                soup2 = BeautifulSoup(html2, 'html.parser')

                # 处理所有者地址
                try:
                    tblOwner = soup2.find("div", id="anchorOwner").find_next('table')
                    Owner_Address = tblOwner.find("td", text="Address").find_next('td').text.strip()
                except AttributeError:
                    pass

                # 处理代理人名称
                try:
                    tblRep = soup2.find("div", id="anchorRepresentative").find_next('table')
                    Representative_Name = tblRep.find("td", text="Name").find_next('td').text.strip()
                except AttributeError:
                    pass

                # 切回主页面
                driver.switch_to.default_content()
            except TimeoutException:
                print(f"第{i}页第{x+1}条数据的iframe加载失败,跳过")
                driver.switch_to.default_content()  # 确保切回主页面

            # 添加到记录
            records.append((
                Designated_territory, Applicant_name, Trade_mark_name, Application_date,
                Application_number, Trade_mark_type, Nice_class, Owner_Address,
                Trade_mark_office, Registration_number, Trade_mark_status,
                Registration_date, Representative_Name
            ))

            # 关闭详情标签页
            try:
                WebDriverWait(driver, 10).until(
                    EC.element_to_be_clickable((By.CSS_SELECTOR, 'a.close_tab'))
                ).click()
            except TimeoutException:
                print("关闭标签页失败,尝试刷新页面")
                driver.refresh()
                random_sleep(5, 10)
            random_sleep(1, 3)

        # 跳转到下一页
        try:
            WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.ID, 'next_t_grid_toppager'))
            ).click()
            random_sleep(5, 10)
        except TimeoutException:
            print("已到最后一页,停止循环")
            break

    # 保存数据到Excel
    df = pd.DataFrame(records, columns=[
        'Designated_territory','Applicant_name','Trade_mark_name','Application_date',
        'Application_number','Trade_mark_type', 'Nice_class','Owner_Address',
        'Trade_mark_office', 'Registration_number','Trade_mark_status',
        'Registration_date','Representative_Name'
    ])
    df.to_excel(xlsName, sheet_name='sheet1', index=False, encoding='utf-8')
    print(f"数据抓取完成,共{len(records)}条记录,耗时{time.time()-start_time:.2f}秒")

except Exception as e:
    print(f"抓取过程中出现错误:{str(e)}")
    # 保存已抓取的数据
    df = pd.DataFrame(records, columns=[
        'Designated_territory','Applicant_name','Trade_mark_name','Application_date',
        'Application_number','Trade_mark_type', 'Nice_class','Owner_Address',
        'Trade_mark_office', 'Registration_number','Trade_mark_status',
        'Registration_date','Representative_Name'
    ])
    df.to_excel(xlsName, sheet_name='sheet1', index=False, encoding='utf-8')
    traceback.print_exc()

finally:
    driver.quit()

三、重要注意事项

  1. 遵守网站使用条款:务必先阅读TMDN的使用条款,确保你的抓取行为符合规定,不要用于商业用途或侵犯版权。
  2. 控制抓取频率:即使优化了代码,也不要过于频繁地请求,建议在夜间或低峰时段进行抓取,避免给服务器造成压力。
  3. 避免账号封禁:如果网站需要登录,尽量人工登录后通过保存cookie维持会话,不要用自动化工具登录。
  4. 定期更新代码:网站的HTML结构和反爬机制可能会变化,需要定期检查并调整元素定位方式。

内容的提问来源于stack exchange,提问作者Milah Lalla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:01:51