You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从HTML的tr、td表格提取数据及实现USPTO循环爬虫

解决USPTO批量爬取中的HTML解析问题

Hey Sanjay, let's work through your USPTO scraping issue. I see you're trying to batch fetch data by iterating over reel and frame parameters from a CSV, but hitting snags with HTML parsing. Let's break down what's wrong with your current code and fix it up.

现有代码的核心问题

  • 绝对XPath路径+错误的相对定位:你的循环里用row.xpath('/html/body/...'),但开头的/会让XPath从文档根节点重新查找,完全忽略了当前的row节点。应该用相对路径(以.//开头)来基于当前行定位元素。
  • 冗余且脆弱的路径:全量的绝对路径(比如/html/body/table[3]/tbody/tr/td/table/...)很容易因为页面结构微小变化失效,应该简化定位逻辑,找更稳定的标识。
  • 异常处理不规范:print error没有定义变量,应该明确捕获异常并输出具体的错误信息,方便调试。
  • 编码处理冗余:如果用Python3,字符串默认是UTF-8,不需要手动encode('utf8');如果用Python2,建议升级到Python3来避免编码坑。
  • 库混用但未充分利用:你导入了BeautifulSoup但没使用,其实它的CSS选择器比XPath更直观,适合嵌套表格的解析。

修正后的代码示例

下面是优化后的代码,结合了从CSV读取参数、批量请求、稳定解析和规范的异常处理:

import requests
from bs4 import BeautifulSoup
import csv
from time import sleep

# 基础URL模板,用占位符替换reel和frame
BASE_URL = "http://legacy-assignments.uspto.gov/assignments/q?db=pat&qt=rf&reel={}&frame={}&pat=&pub=&intn=&asnr=&asnri=&asne=&asnei=&asns="

def scrape_uspto(reel, frame):
    try:
        url = BASE_URL.format(reel, frame)
        response = requests.get(url, timeout=10)
        response.raise_for_status()  # 检查请求是否成功
        soup = BeautifulSoup(response.text, 'lxml')
        
        # 定位核心数据表格(避免绝对路径,找最外层的目标表格)
        main_table = soup.find('table', {'border': '0', 'cellpadding': '3', 'cellspacing': '1'})
        if not main_table:
            print(f"Reel {reel}, Frame {frame}: 未找到核心表格")
            return None
        
        data = {}
        
        # 提取第一部分基础信息(Reel/Frame、Recorded Date等)
        info_table = main_table.find('table')
        if info_table:
            # Reel/Frame
            rf_elem = info_table.find('a', href=True)
            data['reel_frame'] = rf_elem.get_text(strip=True) if rf_elem else None
            
            # Recorded Date
            recorded_elem = info_table.find('td', text=lambda t: t and 'Recorded' in t)
            if recorded_elem:
                data['recorded_date'] = recorded_elem.find_next_sibling('td').get_text(strip=True)
            
            # Attorney
            attorney_elem = info_table.find('td', text=lambda t: t and 'Attorney' in t)
            if attorney_elem:
                data['attorney'] = attorney_elem.find_next_sibling('td').get_text(strip=True)
            
            # Conveyance
            conveyance_elem = info_table.find('td', text=lambda t: t and 'Conveyance' in t)
            if conveyance_elem:
                data['conveyance'] = conveyance_elem.find_next_sibling('td').get_text(strip=True)
        
        # 提取第二部分专利信息
        property_table = main_table.find('table', {'class': 'tblBorder'})
        if property_table:
            # Total Properties
            total_elem = property_table.find('div', text=lambda t: t and 'Total properties:' in t)
            if total_elem:
                data['total_properties'] = total_elem.get_text(strip=True).replace('Total properties:', '').strip()
            
            # 专利详情行
            patent_rows = property_table.find_all('tr')[1:]  # 跳过表头行
            for idx, row in enumerate(patent_rows, 1):
                tds = row.find_all('td')
                if len(tds) >= 8:
                    data[f'patent_{idx}_number'] = tds[1].get_text(strip=True) if tds[1].a else tds[1].get_text(strip=True)
                    data[f'patent_{idx}_issue_date'] = tds[3].get_text(strip=True)
                    data[f'patent_{idx}_application'] = tds[5].get_text(strip=True)
                    data[f'patent_{idx}_filing_date'] = tds[7].get_text(strip=True)
                if len(tds) >= 4 and idx == 2:  # 处理publication行
                    data['publication_number'] = tds[1].get_text(strip=True) if tds[1].a else tds[1].get_text(strip=True)
                    data['publication_date'] = tds[3].get_text(strip=True)
                if len(tds) >=2 and idx ==3:  # 处理title行
                    data['title'] = tds[1].get_text(strip=True)
        
        return data
    
    except Exception as e:
        print(f"Reel {reel}, Frame {frame}: 爬取失败 - {str(e)}")
        return None

# 批量从CSV读取参数并爬取
def batch_scrape(csv_path):
    with open(csv_path, 'r', encoding='utf-8') as f:
        reader = csv.DictReader(f)
        # 假设CSV列名为reel和frame
        for row in reader:
            reel = row['reel'].strip()
            frame = row['frame'].strip()
            print(f"正在爬取 Reel: {reel}, Frame: {frame}")
            result = scrape_uspto(reel, frame)
            if result:
                # 这里可以将结果写入新的CSV或者数据库
                print(f"爬取结果: {result}")
            # 加个延迟,避免被封IP
            sleep(2)

# 调用示例(替换为你的CSV路径)
if __name__ == "__main__":
    batch_scrape('your_parameters.csv')

关键改进点说明

  • 相对定位+文本匹配:用find和文本匹配(比如text=lambda t: t and 'Recorded' in t)来定位元素,比绝对路径更稳定,即使页面结构微调也不容易失效。
  • 模块化函数:把单页面爬取逻辑封装成scrape_uspto函数,批量处理封装成batch_scrape,代码更清晰易维护。
  • 异常处理:捕获所有异常并输出具体错误信息,方便定位哪个参数组合出了问题。
  • 请求延迟:加入sleep(2)避免频繁请求被USPTO封禁IP,也可以考虑使用代理池进一步优化。
  • 灵活的专利数据提取:考虑到可能有多条专利记录,用索引区分不同的专利信息,避免数据覆盖。

内容的提问来源于stack exchange,提问作者Sanjay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:31:42