使用BeautifulSoup提取Marathi文本时如何移除冗余乱码行?
Got it, let's tackle those annoying garbled Marathi lines like "ल क म ळव ल क म ळव" that keep showing up at the end of your extraction—even after you added script.decompose(). Here's why they're happening and how to get rid of them for good:
Common Sources of the Garbled Text
Those lines are almost always coming from non-content elements you haven't cleaned up yet:
- Unremoved non-text tags: You handled
<script>, but<style>,<noscript>,<iframe>, or even hidden header/footer sections might still be holding junk text. - HTML comments: Some pages hide test characters or leftover code inside
<!-- ... -->comments, which BeautifulSoup will pull into your text if you don't explicitly remove them. - Hidden DOM elements: Elements marked with
display:noneorvisibility:hidden(often ads, popups, or dynamic loader placeholders) might contain repeated, meaningless Marathi characters. - Redundant whitespace/character repeats: The garbled lines you're seeing are just single Marathi characters spaced out—they're not actual content, just noise from page rendering leftovers.
Step-by-Step Fix Code
Here's an updated version of your script that targets all these sources:
from bs4 import BeautifulSoup, Comment import re # Assume your raw HTML is stored in the 'html' variable soup = BeautifulSoup(html, 'html.parser') # 1. Strip out all non-content tags for unwanted_tag in soup(['script', 'style', 'noscript', 'iframe', 'header', 'footer', 'nav']): unwanted_tag.decompose() # 2. Remove HTML comments entirely all_comments = soup.find_all(string=lambda text: isinstance(text, Comment)) for comment in all_comments: comment.extract() # 3. Delete hidden elements (based on style or class attributes) hidden_elements = soup.find_all( lambda tag: # Check for hidden styles (tag.has_attr('style') and ('display:none' in tag['style'].lower() or 'visibility:hidden' in tag['style'].lower())) # Check for hidden classes or (tag.has_attr('class') and any('hidden' in cls.lower() for cls in tag['class'])) ) for elem in hidden_elements: elem.decompose() # 4. Extract and clean the final text raw_text = soup.get_text(separator='\n') # Filter out blank lines and the garbled single-character lines cleaned_marathi_text = '\n'.join( line.strip() for line in raw_text.split('\n') if line.strip() and not re.match(r'^(\s*[\u0900-\u097F]\s*)+$', line.strip()) ) print(cleaned_marathi_text)
What Each Part Does
- Tag removal: We expand beyond
<script>to include other non-content tags that often hold junk. Adjust the list (like adding<aside>if your blog uses that for sidebars) based on the specific page structure. - Comment extraction: Ensures no leftover test text from comments makes it into your output.
- Hidden element cleanup: Gets rid of elements that aren't visible to users but still exist in the HTML.
- Final text filtering: The regex
^(\s*[\u0900-\u097F]\s*)+$specifically targets lines made up of nothing but single Marathi characters (spaced out), which is exactly the garbled text you're seeing. We also strip blank lines to keep the output clean.
Bonus: Target Specific Junk Containers
If you notice the garbled text is coming from a specific part of the page (like a div with an id like loader or ad-container), you can directly remove it with:
junk_container = soup.find(id='loader') if junk_container: junk_container.decompose()
This should eliminate those random garbled lines entirely while keeping your core Marathi content intact.
内容的提问来源于stack exchange,提问作者shantanuo

