如何基于HTML模板生成格式化纯文本文件(Django/Python实现)
从模板生成纯文本文件的Python实现方案(含批量处理)
方案一:直接使用Jinja2编写纯文本模板
这是最高效的方案——跳过HTML模板环节,直接针对纯文本格式编写Jinja2模板,完全匹配输出需求,无需额外格式转换。
纯文本模板示例(book_template.txt)
{{ book.name }} The author of the book is {{ book.author }}. For more description, go to "{{ book.url }}" Some important notions on the book {% for item in points %} - {{ item }} {% endfor %}
Python代码实现(含批量生成)
from jinja2 import Environment, FileSystemLoader import os # 初始化Jinja2环境,指定模板目录 env = Environment(loader=FileSystemLoader('templates')) template = env.get_template('book_template.txt') # 模拟从数据库批量获取的数据集(实际项目中用ORM查询) books_data = [ { "book": {"name": "Python For Beginners", "author": "John Doe", "url": "https://mybooks.example.com/book1"}, "points": ["Variable", "Strings", "File R/W operations", "Loops", "Conditions"] }, # 更多书籍数据... ] # 创建输出目录 os.makedirs('output_books', exist_ok=True) # 批量生成文件 for idx, context in enumerate(books_data, 1): text_content = template.render(context) with open(f'output_books/book_{idx}.txt', 'w', encoding='utf-8') as f: f.write(text_content)
方案二:先渲染HTML模板,再转换为纯文本
如果已有现成的HTML模板,无需重新编写纯文本模板,可先渲染HTML再转成纯文本格式。常用工具包括html2text或BeautifulSoup。
使用html2text实现
from jinja2 import Environment, FileSystemLoader import html2text import os # 加载HTML模板 env = Environment(loader=FileSystemLoader('templates')) html_template = env.get_template('book_template.html') # 批量数据 books_data = [ { "book": {"name": "Python For Beginners", "author": "John Doe", "url": "https://mybooks.example.com/book1"}, "points": ["Variable", "Strings", "File R/W operations", "Loops", "Conditions"] }, # 更多数据... ] # 初始化html2text转换器,配置格式 text_converter = html2text.HTML2Text() text_converter.body_width = 0 # 禁用自动换行 text_converter.ignore_images = True text_converter.ignore_links = False # 保留链接文本 os.makedirs('output_books_html', exist_ok=True) # 批量生成 for idx, context in enumerate(books_data, 1): html_content = html_template.render(context) text_content = text_converter.handle(html_content) # 手动修正列表格式,匹配需求 text_content = text_content.replace(' * ', ' - ') with open(f'output_books_html/book_{idx}.txt', 'w', encoding='utf-8') as f: f.write(text_content)
使用BeautifulSoup自定义转换
如果需要更精细的格式控制,可直接解析HTML标签生成纯文本:
from jinja2 import Environment, FileSystemLoader from bs4 import BeautifulSoup import os env = Environment(loader=FileSystemLoader('templates')) html_template = env.get_template('book_template.html') def html_to_text(html): soup = BeautifulSoup(html, 'html.parser') lines = [] for elem in soup.recursiveChildGenerator(): if isinstance(elem, str): text = elem.strip() if text: lines.append(text) elif elem.name == 'li': lines.append(f' - {elem.get_text(strip=True)}') elif elem.name == 'p': lines.append(elem.get_text(strip=True)) return '\n'.join(lines) # 批量生成逻辑同前 for idx, context in enumerate(books_data, 1): html_content = html_template.render(context) text_content = html_to_text(html_content) with open(f'output_books_custom/book_{idx}.txt', 'w', encoding='utf-8') as f: f.write(text_content)
方案三:Django框架下的实现
Django自带模板系统支持纯文本渲染,也可复用现有HTML模板转文本。
直接渲染纯文本模板
from django.template import loader from myapp.models import Book import os # 批量获取数据库数据(用prefetch_related避免N+1查询) books = Book.objects.prefetch_related('notions').all() # 加载纯文本模板 template = loader.get_template('book_text_template.txt') os.makedirs('django_output', exist_ok=True) for book in books: context = { 'book': book, 'points': [notion.content for notion in book.notions.all()] } text_content = template.render(context) with open(f'django_output/{book.slug}.txt', 'w', encoding='utf-8') as f: f.write(text_content)
Django中HTML转纯文本
可复用html2text或Django自带的strip_tags函数(后者仅去除标签,格式需手动调整):
from django.template import loader from django.utils.html import strip_tags from myapp.models import Book template = loader.get_template('book_html_template.html') books = Book.objects.all() for book in books: html_content = template.render({'book': book, 'points': book.notions.all()}) raw_text = strip_tags(html_content) # 手动调整格式,匹配输出需求 formatted_text = raw_text.replace('\n\n', '\n').replace('Some important notions on the book', '\n Some important notions on the book\n - ') with open(f'django_output/{book.slug}.txt', 'w', encoding='utf-8') as f: f.write(formatted_text)
批量生成的优化建议
- 数据库查询优化:使用批量查询(如Django的
all()、prefetch_related)避免频繁数据库请求,解决N+1查询问题。 - 并行处理:针对超100份文件的生成,可使用
concurrent.futures.ThreadPoolExecutor(IO密集型任务适合线程)提升速度:from concurrent.futures import ThreadPoolExecutor def generate_file(context, idx): text_content = template.render(context) with open(f'output_books/book_{idx}.txt', 'w', encoding='utf-8') as f: f.write(text_content) with ThreadPoolExecutor(max_workers=4) as executor: executor.map(generate_file, books_data, range(1, len(books_data)+1)) - 异步任务:如果是Web应用,可使用Celery将批量生成任务放入异步队列,避免阻塞用户请求。
内容的提问来源于stack exchange,提问作者Steve Yonkeu
相关产品推荐
相关产品推荐

