Python新手求助:如何从output2.txt提取title=后的内容
解决提取HTML中
title=后内容的问题 你当前的代码只收集了title出现的索引位置,并没有实际提取目标内容。要拿到title=之后的文本,你可以按以下方式修改:
- 把查找目标从
"title"改成"title=",直接定位到属性值的起始点 - 找到
title=的位置后,计算属性值的起始索引:start = pos + len("title=") - 从
start位置开始,找到下一个双引号(HTML属性值通常用双引号包裹)的位置作为结束点 - 截取两个位置之间的文本就是你要的内容
修改后的代码示例
with open('output2.txt', 'r', encoding='utf-8') as f: extracted_titles = [] for line in f: # 循环处理一行中多个title=的情况 pos = line.find("title=") while pos != -1: # 计算属性值起始位置 start = pos + len("title=") # 找下一个双引号的位置 end = line.find('"', start) if end != -1: # 截取内容并去除前后空格 title_content = line[start:end].strip() extracted_titles.append(title_content) # 继续查找下一个title=的位置 pos = line.find("title=", end) # 打印提取到的所有内容 for content in extracted_titles: print(content)
补充建议
- 如果你的HTML里属性值用单引号包裹,把
line.find('"', start)改成line.find("'", start)即可 - 更健壮的方式是用HTML解析库(比如BeautifulSoup),能自动处理换行、转义字符等特殊情况:
from bs4 import BeautifulSoup with open('output2.txt', 'r', encoding='utf-8') as f: soup = BeautifulSoup(f.read(), 'html.parser') # 提取所有带title属性的元素的属性值 all_titles = [tag.get('title') for tag in soup.find_all(attrs={"title": True}) if tag.get('title')] for title in all_titles: print(title)
内容的提问来源于stack exchange,提问作者Jakub Sobola
相关产品推荐
相关产品推荐

