如何用Python消除两个文本块之间的重叠内容
How to Remove Overlapping Text Between Two Strings and Reintegrate Content
Hey there, let's work through this problem of eliminating the overlapping text between your two files and getting the clean merged content you need. You've already taken the key first step by extracting the last sentence of text1 and the first sentence of text2—nice work! Now we just need to pinpoint the overlapping segment and trim text2 appropriately.
Step-by-Step Solution
The core idea is to find the longest overlapping substring that exists at the end of text1_last_sentence and the start of text2_first_sentence. Once we identify that, we can remove the overlapping part from text2 and keep only the unique content.
Here's the complete, working code to achieve this:
text1 = """Some of the major unsolved problems in physics are theoretical, meaning that existing theories seem incapable of explaining a certain observed phenomenon or experimental result. The others are experimental, meaning that there is a difficulty in creating an experiment to test a proposed theory or investigate a phenomenon in""" text2 = """theory or investigate a phenomenon in greater detail.There are still some deficiencies in the Standard Model of physics, such as the origin of mass, the strong CP problem, neutrino mass, matter–antimatter asymmetry, and the nature of dark matter and dark energy.""" # Extract the relevant sentences, strip whitespace for consistency text1_last_sentence = list(filter(None, text1.split(".")))[-1].strip() text2_first_sentence = text2.split(".")[0].strip() # Find the longest overlapping substring between the two sentences max_overlap_length = min(len(text1_last_sentence), len(text2_first_sentence)) overlap_length = 0 # Check from longest possible overlap down to 1 to get the most accurate match for length in range(max_overlap_length, 0, -1): if text1_last_sentence.endswith(text2_first_sentence[:length]): overlap_length = length break # Trim the overlapping part from text2's first sentence and reconstruct text2 trimmed_first_part = text2_first_sentence[overlap_length:] text2_trimmed = f"{trimmed_first_part}.{text2.split('.', 1)[1]}" # Output the results print("Original text1 (unchanged):\n") print(text1, "\n") print("Processed text2 (overlap removed):\n") print(text2_trimmed)
Explanation
- Extract & Clean Sentences: We split the texts by periods, filter out empty strings (to handle trailing periods), and use
.strip()to avoid whitespace-related mismatches. - Detect Longest Overlap: We check from the longest possible overlap (the length of the shorter sentence) down to 1. This ensures we catch the full overlapping segment instead of a partial match.
- Trim & Reconstruct: We slice off the overlapping prefix from
text2's first sentence, then reattach the rest oftext2(usingsplit('.', 1)[1]to grab everything after the first period) to form the cleaned final text.
Expected Output
Original text1 (unchanged): Some of the major unsolved problems in physics are theoretical, meaning that existing theories seem incapable of explaining a certain observed phenomenon or experimental result. The others are experimental, meaning that there is a difficulty in creating an experiment to test a proposed theory or investigate a phenomenon in Processed text2 (overlap removed): greater detail.There are still some deficiencies in the Standard Model of physics, such as the origin of mass, the strong CP problem, neutrino mass, matter–antimatter asymmetry, and the nature of dark matter and dark energy.
内容的提问来源于stack exchange,提问作者henry
相关产品推荐
相关产品推荐

