如何将bs4.element.Tag转换为Pandas?解析Blogspot文本遇提取问题
Hey there! Let's work through your Blogspot parsing and Pandas conversion issues together.
Why findAll() returns two objects?
Blogspot pages sometimes have multiple elements matching your selector (like sidebars or related content sections that share the same attributes). Instead of grabbing all matches, use find() to get just the main article content (the first matching element):
Fixing the parsing & converting to Pandas
Here's a revised code snippet that addresses both your issues:
from urllib.request import urlopen from bs4 import BeautifulSoup import pandas as pd # Fetch and parse the page html = urlopen("https://edwvb.blogspot.ru/2018/03/3-tipa-povedeniya-kotorye-opredelyayut-uspeshnyh-prodavcov.html") bsObj = BeautifulSoup(html, "html.parser") # Get the main article content (single Tag object) post_content = bsObj.find("div", {"trbidi":"on"}) # Extract clean text from the Tag clean_article_text = post_content.get_text(strip=True, separator='\n') # Convert to a Pandas DataFrame df = pd.DataFrame({"article_content": [clean_article_text]})
If you need to handle multiple matching Tags
If you intentionally want to keep both objects from findAll(), you can loop through them to extract text, then build the DataFrame:
# Get all matching divs nameList = bsObj.findAll("div", {"dir":"ltr", "style":"text-align: left;", "trbidi":"on"}) # Extract text from each Tag text_list = [tag.get_text(strip=True, separator='\n') for tag in nameList] # Create DataFrame with all entries df = pd.DataFrame({"content": text_list})
Key points to remember:
- Use
find()instead offindAll()when you only need the primary content element to avoid extra matches. get_text()converts abs4.element.Taginto plain text, which is easy to feed into Pandas.- Pandas works best with structured data (lists, dictionaries), so convert your extracted text into a list (even a single-item list) before creating the DataFrame.
内容的提问来源于stack exchange,提问作者Edward
相关产品推荐
相关产品推荐

