Python循环逻辑问题:如何将间接引用内容归类至根引用类别
Hey there! Let's tackle this problem—you want to group not just articles that directly quote A, but also those that quote articles which themselves quote A (like D and E quoting C, which quotes A). Here's how to adjust your code to handle these indirect references:
First, let's recap your existing setup so we're on the same page:
import pandas as pd art1 = ["A",'B','C','D','E','F'] quote = ["",'A','A','C','C','A'] g2 = pd.DataFrame({"Article":art1,"Quote":quote})
This gives us a DataFrame where:
- A has no quotes
- B, C, F directly quote A
- D, E directly quote C (so they indirectly quote A)
The Core Idea: Trace the Reference Chain
We need to build a "root reference" for each article—this is the final article that all its references lead back to. For D, that root is A (D → C → A); for E, it's also A.
Solution 1: Recursive Reference Traversal
This uses a simple recursive function to climb up the reference chain until we hit a root (an article with no quote):
def find_root(article, quote_map): # Stop if the article has no quote, or quotes itself if not quote_map[article] or quote_map[article] == article: return article # Keep climbing up the reference chain return find_root(quote_map[article], quote_map) # Create a dictionary mapping each article to its direct quote quote_map = g2.set_index('Article')['Quote'].to_dict() # Add a new column with the root reference for each article g2['Root_Quote'] = g2['Article'].apply(lambda x: find_root(x, quote_map)) # Now get all articles whose root is A articles_under_a = g2[g2['Root_Quote'] == 'A']['Article'].tolist() print(articles_under_a) # Output: ['A', 'B', 'C', 'D', 'E', 'F']
Solution 2: Iterative Traversal (For Long Reference Chains)
If you're worried about hitting recursion limits with extra-long reference chains, use an iterative approach instead:
quote_map = g2.set_index('Article')['Quote'].to_dict() root_reference = {} for article in quote_map: current_article = article # Traverse up the chain until we find a root or a precomputed root while quote_map[current_article] and quote_map[current_article] != current_article and quote_map[current_article] not in root_reference: current_article = quote_map[current_article] # Assign the root reference for the original article root_reference[article] = root_reference.get(current_article, current_article) # Map the root references back to your DataFrame g2['Root_Quote'] = g2['Article'].map(root_reference) # Get all articles linked to A articles_under_a = g2[g2['Root_Quote'] == 'A']['Article'].tolist() print(articles_under_a) # Same output as above
How to Replace Your Existing For Loop
If your original for loop only checked direct quotes (like q == 'A'), you can replace that logic with the Root_Quote column. Now you'll capture every article that traces back to A, whether directly or indirectly.
内容的提问来源于stack exchange,提问作者judy

