如何用Python-Tika从XLS文件中提取独立批注
问题:从XLS文件中提取独立批注(Python-Tika/Java Tika无法区分批注与内容)
尝试用Python-Tika提取XLS文件里的批注,但调用Tika获取文本时,批注会直接夹杂在工作表内容中,无法区分每条批注的起止位置,也没法划分批注和新工作表的边界。尝试解析返回内容也没解决,用Java版Tika同样没法单独提取批注。
以下是Tika返回的示例内容:
Estimates Geography File with Census and FIPS Codes fipscensus Table with column headers in row 7. Area name in column J. Estimates Geography File with Census and FIPS Codes Source: U.S. Census Bureau Internet Release date: January 20, 2004 State Code (FIPS) County Code (FIPS) County Subdivision Code (Census) Place Code (Census) Consolidated City Code (Census) County Subdivision Code (FIPS) Place Code (FIPS) Consolidated City Code (FIPS) Summary Level Area Name (including legal/statistical area description) 01 000 000 0000 0000 00000 00000 00000 040 Alabama 01 000 000 0005 0000 00000 00124 00000 162 Abbeville city 01 000 000 0010 0000 00000 00460 00000 162 12Page &P 01 000 000 0015 0000 00000 00484 00000 162 Addison town 01 000 000 0020 0000 00000 00676 00000 162 Akron town 01 000 000 0025 0000 00000 00820 00000 162 Alabaster city 01 000 000 0030 0000 00000 00988 00000 162 Albertville city 01 000 000 0035 0000 00000 01132 00000 162 Alexander City city 01 000 000 0040 0000 00000 01228 00000 162 Aliceville city 01 000 000 0045 0000 00000 01396 00000 162 Allgood town 01 000 000 0050 0000 00000 01660 00000 162 Altoona town 01 000 000 0055 0000 00000 01708 00000 162 Andalusia city 01 000 000 0057 0000 00000 01756 00000 162 Anderson town 01 000 000 0060 0000 00000 01852 00000 162 Anniston city 01 000 000 0065 0000 00000 02116 00000 162 Arab city 01 000 000 0070 0000 00000 02260 00000 162 Ardmore town 01 000 000 0073 0000 00000 02320 00000 162 Argo town 01 000 000 0075 0000 00000 02428 00000 162 Ariton town 01 000 000 0077 0000 00000 02500 00000 162 Arley town 01 000 000 0080 0000 00000 02836 00000 162 Ashford city 01 000 000 0085 0000 00000 02860 00000 162 Ashland city 01 000 000 0090 0000 00000 02908 00000 162 Ashville town 01 000 000 0095 0000 00000 02956 00000 162 12Page &P 01 000 000 0100 0000 00000 03004 00000 162 Atmore city 01 000 000 0105 0000 00000 03028 00000 162 Attalla city 01 000 000 0110 0000 00000 03076 00000 162 Auburn city 01 000 000 0115 0000 00000 03220 00000 162 Autaugaville town 01 000 000 0120 0000 00000 03364 00000 162 Avon town 01 000 000 0125 0000 00000 03556 00000 162 Babbie town 01 000 000 0127 0000 00000 03676 00000 162 Baileyton town 01 000 000 0128 0000 00000 03724 00000 162 Bakerhill city 01 000 000 0130 0000 00000 03940 00000 162 Banks town &CPage &P of &N Mugdha test comment Meenal: Test Comment to check xls parser Mugdha comment with new line New line comment Sheet2
上述返回内容包含3条独立批注,后续出现的是新工作表名称Sheet2:
- 批注1:
Mugdha comment with new line New line comment - 批注2:
Mugdha test comment - 批注3:
Meenal: Test Comment to check xls parser
解决方案:绕过Tika,直接操作XLS文件提取批注
Tika的文本提取逻辑是整合所有可见文本,不会区分批注和工作表内容,因此无法直接实现精准提取。建议使用Excel专用处理库,直接读取结构化的批注数据:
1. 针对旧版XLS(.xls):使用xlrd3
原版xlrd移除了.xls批注提取功能,可使用分支版本xlrd3:
- 安装:
pip install xlrd3 - 提取代码示例:
import xlrd3 workbook = xlrd3.open_workbook("your_file.xls", formatting_info=True) for sheet_idx in range(workbook.nsheets): sheet = workbook.sheet_by_index(sheet_idx) print(f"工作表:{sheet.name}") # 遍历所有单元格提取批注 for row in range(sheet.nrows): for col in range(sheet.ncols): cell = sheet.cell(row, col) if cell.xf_index is not None: xf = workbook.xf_list[cell.xf_index] if xf.has_comment: comment = workbook.comment_list[xf.comment_index] print(f"单元格({row+1}, {col+1})批注:") print(f"作者:{comment.author}") print(f"内容:{comment.text}\n")
2. 针对新版XLSX(.xlsx):使用openpyxl
openpyxl原生支持.xlsx格式的批注提取:
- 安装:
pip install openpyxl - 提取代码示例:
from openpyxl import load_workbook workbook = load_workbook("your_file.xlsx") for sheet_name in workbook.sheetnames: sheet = workbook[sheet_name] print(f"工作表:{sheet_name}") # 遍历所有批注 for comment in sheet.comments: cell_coord = comment.coordinate print(f"单元格{cell_coord}批注:") print(f"作者:{comment.author}") print(f"内容:{comment.text}\n")
3. 兼容新旧格式:使用pyexcel
pyexcel封装了多格式处理逻辑,可统一处理.xls和.xlsx:
- 安装(需安装对应格式依赖):
pip install pyexcel pyexcel-xls pyexcel-xlsx - 提取代码示例:
import pyexcel as p workbook = p.get_book(file_name="your_file.xls") for sheet in workbook: print(f"工作表:{sheet.name}") # 提取批注(依赖底层库实现) if hasattr(sheet, 'comments'): for coord, comment in sheet.comments.items(): print(f"单元格{coord}批注:{comment}\n")
内容的提问来源于stack exchange,提问作者Aparna
相关产品推荐
相关产品推荐

