使用BeautifulSoup批量提取XML文件文本并关联对应文件名的实现问题
解决方案
基础溯源功能实现
无需额外引入工具库,直接在现有代码基础上调整打印逻辑即可,修改后代码如下:
from bs4 import BeautifulSoup import glob for filename in glob.glob("*.xml"): with open(filename, encoding='utf-8') as open_file: # 补充encoding避免不同系统编码报错 content = open_file.read() soup = BeautifulSoup(content, 'lxml') textExtracted = soup.findAll(text=True) full_text = ' '.join([t.strip() for t in textExtracted if t.strip()]) # 过滤空白字符 if full_text: # 跳过无有效文本的文件 print(f"=== 来源文件:{filename} ===") print(full_text + "\n")
修改说明:
- 输出提取内容前先打印带文件名的标识行,所有对应文件的内容都跟在标识后,可直接溯源
- 新增了空白文本过滤逻辑,避免输出无意义的空内容
- 补充了文件打开的编码参数,避免不同系统默认编码不一致导致的读取报错
进阶优化(直接实现非英语文本筛选)
如果需要直接一步完成非英语文本的提取筛选,可使用langdetect库实现语言自动检测,无需手动判断文本类型。使用前先执行pip install langdetect安装依赖,修改后代码如下:
from bs4 import BeautifulSoup import glob from langdetect import detect, LangDetectException for filename in glob.glob("*.xml"): with open(filename, encoding='utf-8') as open_file: content = open_file.read() soup = BeautifulSoup(content, 'lxml') textExtracted = soup.findAll(text=True) # 逐段检测语言,过滤纯空白和过短的无意义文本 for text in textExtracted: clean_text = text.strip() if len(clean_text) < 2: continue try: lang = detect(clean_text) if lang != 'en': print(f"来源文件:{filename} | 非英语文本:{clean_text}") except LangDetectException: # 无法识别语言的短文本直接跳过 continue
内容的提问来源于stack exchange,提问作者Satanas
相关产品推荐
相关产品推荐

