如何从HTML的tr、td表格提取数据及实现USPTO循环爬虫
解决USPTO批量爬取中的HTML解析问题
Hey Sanjay, let's work through your USPTO scraping issue. I see you're trying to batch fetch data by iterating over reel and frame parameters from a CSV, but hitting snags with HTML parsing. Let's break down what's wrong with your current code and fix it up.
现有代码的核心问题
- 绝对XPath路径+错误的相对定位:你的循环里用
row.xpath('/html/body/...'),但开头的/会让XPath从文档根节点重新查找,完全忽略了当前的row节点。应该用相对路径(以.//开头)来基于当前行定位元素。 - 冗余且脆弱的路径:全量的绝对路径(比如
/html/body/table[3]/tbody/tr/td/table/...)很容易因为页面结构微小变化失效,应该简化定位逻辑,找更稳定的标识。 - 异常处理不规范:
print error没有定义变量,应该明确捕获异常并输出具体的错误信息,方便调试。 - 编码处理冗余:如果用Python3,字符串默认是UTF-8,不需要手动
encode('utf8');如果用Python2,建议升级到Python3来避免编码坑。 - 库混用但未充分利用:你导入了
BeautifulSoup但没使用,其实它的CSS选择器比XPath更直观,适合嵌套表格的解析。
修正后的代码示例
下面是优化后的代码,结合了从CSV读取参数、批量请求、稳定解析和规范的异常处理:
import requests from bs4 import BeautifulSoup import csv from time import sleep # 基础URL模板,用占位符替换reel和frame BASE_URL = "http://legacy-assignments.uspto.gov/assignments/q?db=pat&qt=rf&reel={}&frame={}&pat=&pub=&intn=&asnr=&asnri=&asne=&asnei=&asns=" def scrape_uspto(reel, frame): try: url = BASE_URL.format(reel, frame) response = requests.get(url, timeout=10) response.raise_for_status() # 检查请求是否成功 soup = BeautifulSoup(response.text, 'lxml') # 定位核心数据表格(避免绝对路径,找最外层的目标表格) main_table = soup.find('table', {'border': '0', 'cellpadding': '3', 'cellspacing': '1'}) if not main_table: print(f"Reel {reel}, Frame {frame}: 未找到核心表格") return None data = {} # 提取第一部分基础信息(Reel/Frame、Recorded Date等) info_table = main_table.find('table') if info_table: # Reel/Frame rf_elem = info_table.find('a', href=True) data['reel_frame'] = rf_elem.get_text(strip=True) if rf_elem else None # Recorded Date recorded_elem = info_table.find('td', text=lambda t: t and 'Recorded' in t) if recorded_elem: data['recorded_date'] = recorded_elem.find_next_sibling('td').get_text(strip=True) # Attorney attorney_elem = info_table.find('td', text=lambda t: t and 'Attorney' in t) if attorney_elem: data['attorney'] = attorney_elem.find_next_sibling('td').get_text(strip=True) # Conveyance conveyance_elem = info_table.find('td', text=lambda t: t and 'Conveyance' in t) if conveyance_elem: data['conveyance'] = conveyance_elem.find_next_sibling('td').get_text(strip=True) # 提取第二部分专利信息 property_table = main_table.find('table', {'class': 'tblBorder'}) if property_table: # Total Properties total_elem = property_table.find('div', text=lambda t: t and 'Total properties:' in t) if total_elem: data['total_properties'] = total_elem.get_text(strip=True).replace('Total properties:', '').strip() # 专利详情行 patent_rows = property_table.find_all('tr')[1:] # 跳过表头行 for idx, row in enumerate(patent_rows, 1): tds = row.find_all('td') if len(tds) >= 8: data[f'patent_{idx}_number'] = tds[1].get_text(strip=True) if tds[1].a else tds[1].get_text(strip=True) data[f'patent_{idx}_issue_date'] = tds[3].get_text(strip=True) data[f'patent_{idx}_application'] = tds[5].get_text(strip=True) data[f'patent_{idx}_filing_date'] = tds[7].get_text(strip=True) if len(tds) >= 4 and idx == 2: # 处理publication行 data['publication_number'] = tds[1].get_text(strip=True) if tds[1].a else tds[1].get_text(strip=True) data['publication_date'] = tds[3].get_text(strip=True) if len(tds) >=2 and idx ==3: # 处理title行 data['title'] = tds[1].get_text(strip=True) return data except Exception as e: print(f"Reel {reel}, Frame {frame}: 爬取失败 - {str(e)}") return None # 批量从CSV读取参数并爬取 def batch_scrape(csv_path): with open(csv_path, 'r', encoding='utf-8') as f: reader = csv.DictReader(f) # 假设CSV列名为reel和frame for row in reader: reel = row['reel'].strip() frame = row['frame'].strip() print(f"正在爬取 Reel: {reel}, Frame: {frame}") result = scrape_uspto(reel, frame) if result: # 这里可以将结果写入新的CSV或者数据库 print(f"爬取结果: {result}") # 加个延迟,避免被封IP sleep(2) # 调用示例(替换为你的CSV路径) if __name__ == "__main__": batch_scrape('your_parameters.csv')
关键改进点说明
- 相对定位+文本匹配:用
find和文本匹配(比如text=lambda t: t and 'Recorded' in t)来定位元素,比绝对路径更稳定,即使页面结构微调也不容易失效。 - 模块化函数:把单页面爬取逻辑封装成
scrape_uspto函数,批量处理封装成batch_scrape,代码更清晰易维护。 - 异常处理:捕获所有异常并输出具体错误信息,方便定位哪个参数组合出了问题。
- 请求延迟:加入
sleep(2)避免频繁请求被USPTO封禁IP,也可以考虑使用代理池进一步优化。 - 灵活的专利数据提取:考虑到可能有多条专利记录,用索引区分不同的专利信息,避免数据覆盖。
内容的提问来源于stack exchange,提问作者Sanjay
相关产品推荐
相关产品推荐

