You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup批量提取XML文件文本并关联对应文件名的实现问题

解决方案

基础溯源功能实现

无需额外引入工具库,直接在现有代码基础上调整打印逻辑即可,修改后代码如下:

from bs4 import BeautifulSoup
import glob

for filename in glob.glob("*.xml"):
    with open(filename, encoding='utf-8') as open_file: # 补充encoding避免不同系统编码报错
        content = open_file.read()
        soup = BeautifulSoup(content, 'lxml')
        textExtracted = soup.findAll(text=True)
        full_text = ' '.join([t.strip() for t in textExtracted if t.strip()]) # 过滤空白字符
        if full_text: # 跳过无有效文本的文件
            print(f"=== 来源文件:{filename} ===")
            print(full_text + "\n")

修改说明:

  • 输出提取内容前先打印带文件名的标识行,所有对应文件的内容都跟在标识后,可直接溯源
  • 新增了空白文本过滤逻辑,避免输出无意义的空内容
  • 补充了文件打开的编码参数,避免不同系统默认编码不一致导致的读取报错

进阶优化(直接实现非英语文本筛选)

如果需要直接一步完成非英语文本的提取筛选,可使用langdetect库实现语言自动检测,无需手动判断文本类型。使用前先执行pip install langdetect安装依赖,修改后代码如下:

from bs4 import BeautifulSoup
import glob
from langdetect import detect, LangDetectException

for filename in glob.glob("*.xml"):
    with open(filename, encoding='utf-8') as open_file:
        content = open_file.read()
        soup = BeautifulSoup(content, 'lxml')
        textExtracted = soup.findAll(text=True)
        # 逐段检测语言,过滤纯空白和过短的无意义文本
        for text in textExtracted:
            clean_text = text.strip()
            if len(clean_text) < 2:
                continue
            try:
                lang = detect(clean_text)
                if lang != 'en':
                    print(f"来源文件:{filename} | 非英语文本:{clean_text}")
            except LangDetectException:
                # 无法识别语言的短文本直接跳过
                continue

内容的提问来源于stack exchange,提问作者Satanas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 09:39:01