运行基于BeautifulSoup的RSS解析代码出现AttributeError,求解决
Hey there! Let's break down why you're getting that AttributeError: 'NavigableString' object has no attribute 'text' error and fix your code step by step.
错误根源
Your code has a small issue with how you're accessing the <link> tags. When you do xml.findAll('link')[3], you're grabbing the 4th <link> tag (since Python uses 0-indexing) from the RSS feed. Then you try to loop over this single tag with for item in xml.findAll('link')[3] — but looping over a BeautifulSoup tag actually iterates over its child nodes, which are NavigableString objects (plain text), not tags. These string objects don't have a .text attribute, hence the error.
Plus, in standard RSS feeds, the news links you want are nested inside <item> tags, not just top-level <link> tags. Your original approach was targeting the wrong elements entirely.
修复后的代码(遍历所有新闻条目)
Here's a corrected version that properly fetches each news link from the RSS feed and reads the content:
from bs4 import BeautifulSoup import urllib.request # Fetch and parse the RSS feed rss_url = "https://www.theguardian.com/international/rss" response = urllib.request.urlopen(rss_url) xml_soup = BeautifulSoup(response, features='xml') # Get all news items (each <item> tag represents one news story) news_items = xml_soup.findAll('item') for item in news_items: # Extract the link from the current news item news_link = item.find('link').text print(f"News Link: {news_link}") # Fetch the news page content and decode it to a readable string news_page_response = urllib.request.urlopen(news_link) news_content = news_page_response.read().decode('utf-8') # Print a preview of the content (adjust the slice length as needed) print(f"Content Preview: {news_content[:500]}...\n")
关键改进点
- We first grab all
<item>tags (these are the individual news entries in the RSS feed) - For each item, we use
item.find('link')to get the specific link tag inside that news entry - We decode the raw response from
read()using.decode('utf-8')to convert it from binary data to a human-readable string
如果只想获取单个特定链接
If you actually wanted to target just the 4th <link> tag (like your original code tried), here's how to do it without looping:
from bs4 import BeautifulSoup import urllib.request rss_url = "https://www.theguardian.com/international/rss" response = urllib.request.urlopen(rss_url) xml_soup = BeautifulSoup(response, features='xml') # Get the 4th <link> tag (index 3) target_link = xml_soup.findAll('link')[3] # Access its text directly (no need to loop) print(f"Target Link: {target_link.text}") # Fetch and read the page content news_page_response = urllib.request.urlopen(target_link.text) news_content = news_page_response.read().decode('utf-8') print(news_content[:500])
内容的提问来源于stack exchange,提问作者Asma singh

